Method for measuring effluent ammonia nitrogen concentration based on neural network and Bayesian optimization
By applying adaptive bidirectional long and short-term memory neural networks and improving parallel Bayesian optimization algorithms in the field of sewage treatment, the problems of low prediction accuracy and poor generalization ability of effluent ammonia nitrogen concentration in the existing technology are solved, and high-precision and stable prediction effects are achieved, providing strong support for the intelligent regulation of sewage treatment plants.
Patent Information
- Application Number
- CN202510614762.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-14
AI Technical Summary
The existing deep learning methods have poor generalization ability, poor stability and low accuracy when predicting the concentration of ammonia nitrogen in the effluent, making it difficult to provide real-time process control guidance for sewage treatment plants.
The effluent ammonia nitrogen concentration prediction method based on adaptive bidirectional long and short-term memory neural network (ABiLSTM) and improved parallel Bayesian optimization algorithm are used to extract multi-scale features through time-sequence convolution layer and SE attention mechanism, and the input step size is dynamically adjusted to adapt to different working conditions.
The prediction accuracy and generalization ability of the model under different operating conditions is improved, and the water quality change patterns can be fully captured from minute to seasonal time scales, providing reliable support for the intelligent regulation of sewage treatment plants.
Smart Images

Figure CN120144969A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of sewage treatment and artificial intelligence, and specifically relates to a method for measuring the effluent ammonia nitrogen concentration based on neural network and Bayesian optimization. Background Art
[0002] With the increasingly strict environmental protection standards and the complexity of sewage treatment processes, the accurate prediction of the effluent ammonia nitrogen concentration has become a key task to ensure the stable operation of sewage treatment plants and the compliance of water quality. Ammonia nitrogen is one of the core indicators of sewage discharge. Excessive concentration of ammonia nitrogen will not only lead to water eutrophication, but also may cause harm to the ecosystem and human health. The traditional method relying on manual sampling and laboratory detection has a lag, and it is difficult to provide real-time guidance for process control. Therefore, developing an efficient and accurate ammonia nitrogen concentration prediction model has important engineering application value.
[0003] In recent years, deep learning technology has been introduced into the field of sewage treatment due to its powerful feature extraction ability for the prediction of the effluent ammonia nitrogen concentration. However, existing deep learning methods still have problems such as poor generalization ability, weak stability, and low accuracy when predicting the effluent ammonia nitrogen concentration. Summary of the Invention
[0004] Aiming at the problems existing in the prior art, the purpose of the present invention is to propose a method for predicting the effluent ammonia nitrogen concentration based on an adaptive long short-term memory neural network and an improved parallel Bayesian optimization algorithm to improve the prediction accuracy and generalization ability of the model under different working conditions.
[0005] To achieve the above purpose, the technical solution adopted by the present invention is: a method for measuring the effluent ammonia nitrogen concentration based on neural network and Bayesian optimization, including the following steps: Step 1: Construct a basic model with a temporal convolutional layer, an attention mechanism, and a bidirectional long short-term memory neural network, collect feature variables x related to the prediction of the effluent ammonia nitrogen concentration, calculate the Pearson correlation coefficient, retain the feature variables x according to the Pearson correlation coefficient, preprocess the retained feature variables x, and normalize each feature variable x to the interval [0, 1] using the Min-Max method; Step 2: Input the feature variables x into the basic model for feature extraction and enhancement. Extract the local dependence relationship in the time series data through the temporal convolutional layer TCL, and recalibrate the features through the SE attention mechanism to enhance the response of key features and suppress noise; Step 3: Extract the long-term temporal features of the feature sequence through the bidirectional long short-term memory neural network with an adaptive mechanism after feature extraction and enhancement, dynamically adjust the input step size to adapt to the changes in different time periods in the data, capture the bidirectional time dependence relationship in the sequence, and output the predicted value; Step 4: Parallel Bayesian optimization algorithm, dynamically adjusts the acquisition function according to different stages of the optimization process, evaluates multiple candidate points simultaneously, with a parallel adaptive acquisition function, and outputs measured values.
[0006] For the above method for measuring the ammonia nitrogen concentration in the effluent based on neural network and Bayesian optimization, in Step 1, the calculation formula for the Pearson correlation coefficient is: , where r is the Pearson correlation coefficient, x i and y i are the i th observed values of two variables, x m and y m are the means of variables x and y respectively, n is the number of samples, and the characteristic variables include the finally selected dissolved oxygen at the end of aerobic stage, total suspended solids at the end of aerobic stage, effluent pH value, effluent oxidation-reduction potential, and effluent nitrate nitrogen.
[0007] For the above method for measuring the ammonia nitrogen concentration in the effluent based on neural network and Bayesian optimization, the preprocessing includes first cleaning the collected raw data to handle missing values and outliers. For the missing values, different filling strategies are adopted according to the data characteristics. For continuous variables, the time series interpolation method is used to reasonably fill with the values of adjacent time points before and after; for the outliers, combining statistical methods and business experience, the abnormal points are identified by calculating the distribution range of each feature, and corrected by the method of replacing with the median of the sliding window.
[0008] For the above method for measuring the ammonia nitrogen concentration in the effluent based on neural network and Bayesian optimization, Step 2 includes: Step 2-1: Extract the local dependencies of time series data through the temporal convolutional layer TCL. The output x c of TCL is calculated by the formula: , where δ (⋅) is the ReLU activation function, used to introduce non-linearity, W k,c,c′ is the weight matrix of the convolutional kernel, x t+k,c is the value of the input data at time step t + k and input channel c , b c′ is the bias term, b ∈Rc′ ; Step 2-2: Re-calibrate the feature variable x through the SE attention mechanism to enhance the response of key features and suppress noise, perform channel re-calibration, and apply the obtained weight s after excitation to the original input features.
[0009] The above-mentioned method for measuring the concentration of ammonia nitrogen in effluent based on neural network and Bayesian optimization, the step 2-2 includes: Step a: Compress, use global average pooling to generate a channel descriptor z c′′ , used to emphasize the spatial information of the global distribution, and the formula is: , where, x ijc '' represents the eigenvalue of the i th height and j th width at the c ''th channel, H and W represent the height and width of the input channel respectively; Step b: The SE attention mechanism reduces the parameters and complexity by using the compression ratio r, and increases the non-linearity by using the ReLU activation function; Step c: Calculate the weight of each channel through the sigmoid function to achieve effective re-calibration of features, and calculate the output of excitation through the weights M 1 and M 2 of two fully connected layers. The output s is: , where, M 1 ∈R c ′′×c′′ / r , M 2 ∈R c′′ / r×c′′ , z is composed of c '' channel descriptors, z ∈R c′′×1 is the output of the compression step, s ∈R c′′×1 , σ is the sigmoid activation function; Step d: Perform channel re-calibration, apply the obtained weight s after excitation to the original input features. For each channel, its re-calibrated output is: , where, s c′′ is the weight of channel c '', x c′′ is the channelc The original input features of ′′ x s are the outputs after SE recalibration.
[0010] The above method for measuring the ammonia nitrogen concentration in the effluent based on neural network and Bayesian optimization, the step 3 includes: Step 3-1: The bidirectional long short-term memory neural network ABiLSTM processes the input sequence step by step through LSTM units. The LSTM units include the update of the input gate, forget gate, output gate and cell state. The algorithm formula is as follows: , , , , , , Among them, x t ( x t ∈R m×1 ) is the received input vector, h t ( h t ∈R m×1 ) is the hidden state at time t, h t−1 is the hidden state of the previous time step, g t is the new candidate value, f t , i t , o t are the outputs of the forget gate, input gate and output gate at t time respectively, c t is t the cell state at time, W f , W O , W i , W C ( W ∈R h×m ) respectively represent t the weight matrices connecting the input x t to different gates at time, b f, b i , b C , b O ( b ∈R h×1 ) is the bias term; Step 3-2: Define a window with a minimum length of Δ t min . Dynamically adjust the input step size based on the error, and only process the data within the time steps contained in the window each time. The dynamically adjusted window size for the prediction error is: , where α is the adjustment coefficient, and error t is the mean square error at the current time step; Step 3-3: Capture the bidirectional temporal features of the data through forward and backward propagation. For the input x t ∈R m×η , the specific operation process of the ABiLSTM can be expressed as: , , ,
[0011] where H t1 、 H t2 、 H t represent the forward, backward, and hidden states of the ABiLSTM respectively, O t is the output of the ABiLSTM, W f xh , W b xh ∈R h×m , W f hh , W b hh ∈R h×h , W hq ∈R q×2h represents the weight matrix of the model, b f h , b b h ∈Rh×η , b q ∈R q×2η represents the bias term of the model, q is the final output dimension.
[0012] For the above method for measuring the ammonia nitrogen concentration in the effluent based on neural network and Bayesian optimization, the mean square error error at the current time step t is: , where: y t is t the true value at time y p is t the predicted value at time
[0013] For the above method for measuring the ammonia nitrogen concentration in the effluent based on neural network and Bayesian optimization, in step 4, the parameter combination to be optimized by the model Θ is: the number of layers and the number of hidden units of the ABiLSTM layer, the compression ratio of SE r and L 2 regularization parameter, by optimizing the L 2 regularization parameter to optimize the weight matrices W and bias matrices b of the ABiLSTM layer and the TCL layer.
[0014] For the above method for measuring the ammonia nitrogen concentration in the effluent based on neural network and Bayesian optimization, step 4 includes: Step 4-1: Initialize the prior distribution of the Gaussian process and select the initial candidate point set: , where, m ( Θ p ) is the mean function, Θ p represents the p th new input matrix, Θ d represents the set of known input points, k ( Θ p , Θ d ) is the Gaussian kernel function, representing Θ p and Θ d the covariance matrix between each set of data in; Step 4-2: Obtain the observed data by inputting Θ P and combine the prior distribution to calculate Θ PThe formulas for the mean and variance of the posterior distribution are as follows: , , where, μ is the mean of the posterior distribution, V is the variance of the posterior distribution, k ( Θ P ) and ( k ( Θ P ) are the observed value vectors calculated through the kernel function, K is determined by the kernel function K ( Θ P , Θ d )-defined covariance matrix, B is the noise variance, I is the identity matrix, Q is the set of known objective function values; Step 4-3: Select multiple parallel evaluation candidate points by maximizing the acquisition function. The parallel adaptive acquisition function is: , where, D n represents the known observed data set, D n ∈ x , a i represents the i-th acquisition function, i = 1, 2, 3, k = 1, 2, 3,... n ; Step 4-4: Update the posterior distribution according to the formulas for calculating the mean and variance of the posterior distribution of Θ P until the accuracy requirement is met.
[0015] The above method for measuring the effluent ammonia nitrogen concentration based on neural network and Bayesian optimization, the acquisition function includes: In the initial stage of optimization, due to limited information of the objective function, the upper confidence bound UCB function is adopted to enhance the exploration ability, quickly cover the parameter space and identify potential advantageous regions. The UCB function is: where, μ ( Θ P ) is the predicted mean of the objective function at the point Θ P , v ( ΘP ) is the predicted standard deviation at point Θ P , and f is a constant, which is a parameter for controlling the exploration and exploitation balance; In the middle stage of optimization, with the accumulation of information, it turns to the expected improvement EI function to balance exploration and exploitation. The EI function is: , where Θ P + is the input vector corresponding to the currently known optimal parameters; In the later stage of optimization, the probability improvement PI function is adopted, focusing on the identified potential areas for fine search to achieve precise adjustment of the optimal solution. The PI function is: , where P is the probability measure, γ is a non - negative parameter for controlling the exploration intensity.
[0016] The beneficial effects of a method for measuring the ammonia - nitrogen concentration in effluent based on neural network and Bayesian optimization according to the present invention are as follows: By synergistically extracting multi - scale features and strengthening key features through the temporal convolutional layer and SE attention mechanism, the accuracy of the ammonia - nitrogen concentration prediction in effluent is significantly improved; The adaptive adjustment long short - term memory neural network is used to dynamically adjust the input step length to adapt to the temporal patterns of different working conditions; Combining the improved parallel adaptive Bayesian optimization algorithm to improve the parameter search efficiency. At the same time, through ABiLSTM, the complete capture of time scales from minutes to seasons is realized, improving the generalization ability of the model, being able to maintain high prediction accuracy under different working conditions, and providing reliable support for the intelligent regulation of sewage treatment plants. By performing normalization processing, it is ensured that features with different dimensions are comparable, and at the same time, the over - influence of some large - value features on model training is avoided. Through the parallel Bayesian optimization algorithm, the acquisition function can be dynamically adjusted according to different stages of the optimization process, improving the parameter space coverage rate, enhancing the optimization efficiency, and this method speeds up the exploration speed of the parameter space by simultaneously evaluating multiple candidate points, enhancing the generalization ability of the model. Brief Description of the Drawings
[0017] Figure 1 is a schematic structural diagram of the TCL - SE - ABiLSTM model in an embodiment of the present invention; Figure 2 is a schematic structural diagram of ABiLSTM in an embodiment of the present invention; Figure 3 is a schematic diagram of the predicted result of ammonia - nitrogen in effluent in summer in an embodiment of the present invention; Figure 4 is a schematic diagram of the predicted result of ammonia - nitrogen in effluent in winter in an embodiment of the present invention; Detailed Embodiments
[0018] To enable those skilled in the art to better understand the technical solutions of the present invention, the technical solutions of the present invention will be described below in conjunction with specific embodiments and the accompanying drawings.
[0019] Embodiment 1 In the field of predicting the ammonia nitrogen concentration in the effluent of urban sewage treatment plants, accurately capturing the law of water quality change is a key technical problem, which is specifically manifested as follows.
[0020] (1) Existing models are difficult to synchronously capture the ammonia nitrogen concentration fluctuation characteristics from the minute level to the seasonal level, resulting in a decrease in the prediction accuracy of the long-term trend.
[0021] (2) Traditional methods cannot adaptively adjust the weights of input features, affecting the robustness of the model under extreme weather conditions.
[0022] (3) Existing optimization algorithms have a slow convergence speed when dealing with seasonal data distribution offsets and are difficult to meet the requirements of real-time control; (4) The model needs to be retrained under different working conditions due to differences in data distribution, increasing the operation and maintenance costs.
[0023] In view of the above problems, this embodiment proposes an ammonia nitrogen concentration prediction method based on an adaptive bidirectional long short-term memory neural network and an improved parallel Bayesian optimization algorithm. This method synergistically extracts multi-scale features and strengthens key features through a temporal convolutional layer and an SE attention mechanism; uses an adaptive adjustment long short-term memory neural network to dynamically adjust the input step size to adapt to the temporal patterns of different working conditions; and combines an improved parallel adaptive Bayesian optimization algorithm to improve the parameter search efficiency. This embodiment can maintain high prediction accuracy under different working conditions and provide reliable support for the intelligent regulation of sewage treatment plants.
[0024] As Figure 1 - Figure 2 shown, a method for measuring the ammonia nitrogen concentration in the effluent based on a neural network and Bayesian optimization includes the following steps.
[0025] Step 1: Construct a basic model with a temporal convolutional layer, an attention mechanism, and a bidirectional long short-term memory neural network, collect feature variables x related to the prediction of the ammonia nitrogen concentration in the effluent, calculate the Pearson correlation coefficient, retain the feature variables x according to the Pearson correlation coefficient, preprocess the retained feature variables x, and normalize each feature variable x to the interval [0, 1] using the Min-Max method.
[0026] Step 2: Input the feature variables x into the basic model for feature extraction and enhancement. Extract the local dependencies in the time series data through the temporal convolutional layer TCL, and recalibrate the features through the SE attention mechanism to enhance the response of key features and suppress noise.
[0027] Step 3: The feature variables \(x\) after feature extraction and enhancement are used to extract the long-term temporal features of the feature sequence through a bidirectional long short-term memory neural network with an introduced adaptive mechanism, dynamically adjust the input step size to adapt to changes in different time periods in the data, capture the bidirectional time dependencies in the sequence, and output the predicted values.
[0028] Step 4: The parallel Bayesian optimization algorithm dynamically adjusts the acquisition function according to different stages of the optimization process, evaluates multiple candidate points simultaneously, and adaptively acquires the function in parallel to output the measured values.
[0029] Through the feature extraction and enhancement of TCL-SE, the capture of the temporal features of ABiLSTM, and the optimization of parameters by the Bayesian optimization algorithm, the ammonia nitrogen concentration in the effluent can be accurately predicted.
[0030] Example 2 This example is a specific illustration of Example 1.
[0031] As Figure 1 The overall flowchart realizes the real-time prediction of the ammonia nitrogen concentration in the effluent through the combination of TCL-SE, ABiLSTM, and the adaptive Bayesian optimization algorithm and Figure 2 The detailed structure of ABiLSTM shows the working process of ABiLSTM. A method for measuring the ammonia nitrogen concentration in the effluent based on a neural network and Bayesian optimization includes the following steps.
[0032] Step 1: Collect the feature variables related to the prediction of the ammonia nitrogen concentration in the effluent, calculate the Pearson correlation coefficient through formula (1), and retain the feature variables with \(|r|>0.6\).
[0033] (1), where r is the Pearson correlation coefficient, x i and y i are the i th observations of two variables, x m and y m are the means of variables x and y respectively, n is the number of samples. Finally, the dissolved oxygen at the end of aerobic treatment, the total suspended solids at the end of aerobic treatment, the effluent pH value, the effluent redox potential, and the effluent nitrate nitrogen are selected as the input feature variables. The sample data set is x = x 1 , x 2 , …, x n, the collected effluent ammonia nitrogen concentration data is expressed as y = y 1 , y 2 , …, y n , y i is the i-th effluent ammonia nitrogen concentration point, n represents the sample number of easily measurable variables.
[0034] Step 2: First, perform data cleaning on the collected original data to handle missing values and outliers. First, comprehensively clean the original data, focusing on solving the problems of missing values and outliers. For missing values, different filling strategies are adopted according to the data characteristics: for continuous variables such as dissolved oxygen and pH value, the time series interpolation method is used to reasonably fill with the values of adjacent time points before and after; for discrete data, the common values of this feature are used for completion. The detection of outliers combines statistical methods and business experience. By calculating the distribution range (mean ± 3 times the standard deviation) of each feature, the outlier points are identified, and the sliding window median replacement method is used for correction, which not only eliminates the abnormal interference but also maintains the time continuity of the data. After completing the data cleaning, perform normalization processing on all feature variables, and use the Min - Max method to linearly transform each feature to the [0, 1] interval to ensure the comparability of features with different dimensions and avoid the excessive influence of some large - value features on model training.
[0035] Step 3: Feature extraction and enhancement. When the feature variable x is input into the model, the model first extracts the local dependence of the time - series data through the Temporal Convolutional Layer (TCL). The output of the TCL x c can be calculated by the following formula: (2), where, δ (⋅) is the ReLU activation function, used to introduce non - linear characteristics, W k,c,c′ is the weight matrix of the convolution kernel, x t+k,c is the value of the input data at time step t + k and input channel c at, b c′ is the bias term, b ∈R c′ .
[0036] Then, the features are recalibrated through the SE attention mechanism to enhance the response of key features and suppress noise. Specifically, first, compression is performed, and a channel descriptor is generated using global average pooling z c′′ , which is used to emphasize the spatial information of the global distribution. The formula is as follows: (3), where, x ijc '' represents the feature value of the i -th channel at the position with height j and width c '', H and W represent the height and width of the input channel respectively. Secondly, SE reduces the parameters and complexity by using the compression ratio r and increases the non-linearity using the ReLU activation function. Finally, the weights of each channel are calculated through the sigmoid function to achieve effective recalibration of the features. M 1 and M 2 are the weights of two fully connected layers ( M 1 ∈ R c′′×c′′ / r , M 2 ∈ R c′′ / r×c′′ ), then the output M 1 of the excitation is calculated through the weights M 2 of the two fully connected layers s ( s ∈ R c′′×1 ) as: (4), where, z is composed of c '' channel descriptors, z ∈ R c′′×1 is the output of the compression step, s ∈ R c′′×1 , σ is the sigmoid activation function. Finally, channel recalibration is performed, and the weights s obtained after excitation are applied to the original input features. For each channel, its recalibrated output is: (5), where, s c′′ is the weight of channel c '', x c′′Is the channel c The original input feature of ′′ x s Is the output after SE recalibration.
[0037] Step 4: After the data is extracted and enhanced through TCL-SE, the long-term temporal features of the feature sequence are extracted through an adaptive bidirectional long short-term memory neural network (adaptive bidirectional long short-term memory, ABiLSTM), and the predicted value is finally output. Specifically, ABiLSTM first processes the input sequence step by step through the LSTM unit. The core calculations of the LSTM unit include the input gate, forget gate, output gate, and update of the cell state. The algorithm formulas are as follows: (6), (7), (8), (9), (10), (11), Among them, x t ( x t ∈R m×1 ) is the received input vector, h t ( h t ∈R m×1 ) is the hidden state at time t, h t−1 is the hidden state of the previous time step, g t is the new candidate value, f t , i t , o t are the outputs of the forget gate, input gate, and output gate at t time respectively, c t is t the cell state at time, W f , W O , W i , W C ( W ∈Rh×m ) represent respectively t the connection inputs at the moment x t to the weight matrices of different gates, b f , b i , b C , b O ( b ∈R h×1 ) is the bias term, σ and is the sigmoid activation function.
[0038] Then, by defining a window with a minimum length of Δ t min , the input step size is dynamically adjusted based on the error, and only the data within the time steps contained in the window is processed each time. The window size dynamically adjusted according to the prediction error is: (12), where, α is the adjustment coefficient, and error t is the mean square error at the current time step. The calculation formula for error t is: (13), where: y t is t the true value at the moment, y p is t the predicted value at the moment.
[0039] Finally, the bidirectional temporal features of the data are captured through forward and backward propagation. For the input x t ∈R m×η , the specific operation process of ABiLSTM can be expressed as: (14), (15),
[0040] (16), where, H t1 、 H t2 、 H t represent the forward, backward, and hidden states of ABiLSTM respectively, Ot is the output of ABiLSTM, W f xh , W b xh ∈R h×m , W f hh , W b hh ∈R h×h , W hq ∈R q×2h represents the weight matrix of the model, b f h , b b h ∈R h×η , b q ∈R q×2η represents the bias term of the model, q is the final output dimension, and T is the transpose.
[0041] Step 5: Use the improved parallel adaptive Bayesian optimization algorithm for parameter optimization. This embodiment proposes a parallel Bayesian optimization algorithm with an improved adaptive acquisition function, which can dynamically adjust the acquisition function according to different stages of the optimization process, improve the optimization efficiency, and this method can accelerate the exploration speed of the parameter space and enhance the generalization ability of the model by evaluating multiple candidate points simultaneously. The parameter combination to be optimized by the model Θ is: the number of layers and the number of hidden units of the ABiLSTM layer, the compression ratio of SE r and L 2 regularization parameter. By optimizing the L 2 regularization parameter to optimize the weight matrices of the ABiLSTM layer and the TCL layer W and the bias matrix b .
[0042] Step 6: Evaluate the model through two indicators: the normalized root mean square error (NRMSE) and the coefficient of determination (R 2 ): (24), (25), where, y max is the maximum value in the observed data; y min is the minimum value in the observed data; Kis the total number of observed values; y j is the average value of the actual observed values.
[0043] Through the feature extraction and enhancement of TCL-SE, the capture of ABiLSTM time series features, and the optimization of parameters by the Bayesian optimization algorithm, the concentration of ammonia nitrogen in water can be accurately predicted.
[0044] Example 3 As Figure 1 - Figure 4 shown, a method for measuring the concentration of ammonia nitrogen in effluent based on neural network and Bayesian optimization includes the following steps.
[0045] Step 1: Data collection. The data related to the concentration of ammonia nitrogen in the effluent used in this example are all from a sewage treatment plant in Beijing, with time spans from January 16, 2016 to January 22, 2016 and from September 16, 2016 to September 22, 2016 respectively. The data sampling interval is 1h, and there are a total of 1064 samples. Calculate the Pearson correlation coefficient through formula (1), and retain the feature variables with |r|>0.6.
[0046] (1), where r is the Pearson correlation coefficient, x i and y i are the i th observed values of the two variables, x m and y m are the means of the variables x and y respectively, n is the number of samples. Finally, the dissolved oxygen at the end of aerobic stage, the total suspended solids at the end of aerobic stage, the effluent pH value, the effluent redox potential, and the effluent nitrate nitrogen are selected as the input feature variables. The sample data set is x = x 1 , x 2 , …, x n , and the collected data of the ammonia nitrogen concentration in the effluent is expressed as y = y 1 , y 2 , …, y n , y i is the i-th ammonia nitrogen concentration point in the effluent, n represents the number of samples of the easily measurable variables.
[0047] Step 2: Data processing and feature selection. First, clean the collected raw data to handle missing values and outliers. Second, use the Min-Max method to normalize each feature to the interval [0, 1]. Finally,
[0048] Step 3: Feature extraction and enhancement. When the feature variable x is input into the model, the model first extracts the local dependencies in the time series data through the Temporal Convolutional Layer (TCL). The TCL can efficiently capture the mutual influence between each time step in the time series through one-dimensional convolution operations, especially capable of identifying multi-scale features in the sequence. This local feature extraction mechanism enables the model to automatically learn the relationships between each time point when processing complex data with time correlation, thereby better grasping the dynamic changes in the data.
[0049] The output of the TCL x c can be calculated by the following formula: (2), where, δ (⋅) is the ReLU activation function, used to introduce non-linearity, W k,c,c′ is the weight matrix of the convolution kernel, x t+k, c is the value of the input data at time step t + k and input channel c at, b c′ is the bias term, b ∈R C′ .
[0050] Then, the model recalibrates the features through the SE attention mechanism. The SE attention mechanism can effectively enhance the response of key features while suppressing those unimportant or noisy features by learning the non-linear relationships between channels. This process, through the compression of global information and the recalibration of channels, enables the model to focus on the features with high influence on the prediction task while ignoring the redundant information with less impact on the results. Specifically, first, perform compression, and use global average pooling to generate a channel descriptor z c′′ , used to emphasize the spatial information of the global distribution.
[0051] The formula is as follows: (3), where, x ijc'' represents the eigenvalue of the i th channel with height j and width c . H And W represent the height and width of the input channels respectively. Secondly, SE reduces the parameters and complexity by using the compression ratio r and increases the non-linearity by using the ReLU activation function. Finally, the weights of each channel are calculated by the sigmoid function to achieve effective recalibration of the features. M 1 and M 2 are the weights of the two fully connected layers ([[]] M 1 ∈R c′′×c′′ / r , M 2 ∈R c′′ / r×c′′ ). Then, the output of the excitation M 1 and M 2 calculated by the weights of the two fully connected layers s ( s ∈R c′′×1 ) is: (4), where z is composed of c channel descriptors, z ∈R c′′×1 is the output of the compression step, s ∈R c′′×1 , σ is the sigmoid activation function. Finally, the recalibration of the channel is performed, and the weights s obtained after excitation are applied to the original input features. For each channel, its recalibrated output is: (5), where s c′′ is the weight of channel c ( x c′′ ), c is the original input feature of channel x s is the output after SE recalibration.
[0052] Step 4: Prediction of ammonia nitrogen concentration in the effluent in different seasons. After the data is subjected to feature extraction and enhancement through TCL-SE, it then enters the adaptive bidirectional long short-term memory neural network (ABiLSTM) for further processing. When ABiLSTM extracts the long-term temporal features of the feature sequence, it first captures the bidirectional temporal dependencies in the sequence through the BiLSTM structure. Different from the traditional unidirectional LSTM, BiLSTM can take into account the context information at the current moment and its previous and subsequent moments simultaneously, thus effectively capturing the more complex and long-term temporal dependencies in the sequence. This bidirectional feature enables ABiLSTM to more accurately understand the long-term trends and change laws of time series data.
[0053] In addition, ABiLSTM also introduces an adaptive mechanism to adapt to the changes in different time periods in the data by dynamically adjusting the input step size. This feedback mechanism based on the prediction error enables the model to capture features more flexibly by continuously adjusting the input step size. Specifically, ABiLSTM first processes the input sequence time step by time step through the LSTM unit. The core calculations of the LSTM unit include the update of the input gate, forget gate, output gate, and cell state, and its algorithm formulas are as follows: (6), (7), (8), (9), (10), (11), where, x t ( x t ∈ R m×1 ) is the received input vector, h t ( h t ∈ R m×1 ) is the hidden state at time t, h t−1 is the hidden state of the previous time step, g t is the new candidate value, f t , i t , o t are the forget gate, input gate, and output gate at tThe output at a moment c t is t the cell state at a moment, W f , W O , W i , W C ( W ∈R h×m ) respectively represent t the weight matrices connecting the input at a moment x t to different gates, b f , b i , b C , b O ( b ∈R h×1 ) is the bias term, σ and is the sigmoid activation function.
[0054] Then, by defining a window with a minimum length of Δ t min , the input step size is dynamically adjusted based on the error, and only the data within the time steps contained in the window is processed each time. The window size dynamically adjusted according to the prediction error is: (12), where α is the adjustment coefficient, and error t is the mean squared error at the current time step. The calculation formula for error t is: (13), where: y t is t the true value at a moment, y p is t the predicted value at a moment.
[0055] Finally, the bidirectional temporal features of the data are captured through forward and backward propagation. For the input x t ∈R m×η , the specific operation process of ABiLSTM can be expressed as: (14), (15), , (16), Among them, H t1 、 H t2 、 H t represent the forward, backward, and hidden states of ABiLSTM respectively, O t is the output of ABiLSTM, W f xh , W b xh ∈R h×m , W f hh , W b hh ∈R h×h , W hq ∈R q×2h represents the weight matrix of the model, b f h , b b h ∈R h×η , b q ∈R q×2η represents the bias term of the model, q is the final output dimension, Ø is the tanh activation function.
[0056] Step 5: Optimization algorithm. This embodiment proposes an improved parallel Bayesian optimization algorithm with an adaptive acquisition function. This algorithm can dynamically adjust the acquisition function according to different stages of the optimization process, improve the optimization efficiency, and this method speeds up the exploration speed of the parameter space by simultaneously evaluating multiple candidate points and enhances the generalization ability of the model. The parameter combination to be optimized by the model Θ is: the number of layers and hidden units of the ABiLSTM layer, the compression ratio of SE r and L 2 regularization parameter. By optimizing the L 2 regularization parameter to optimize the weight matrix W and bias matrix b of the ABiLSTM layer and the TCL layer.
[0057] Specifically, first initialize the prior distribution of the Gaussian process and select the initial candidate point set, that is: (17), Among them, m ( Θ p ) is the mean function, k ( Θ p , Θ d ) is the Gaussian kernel function, and is the p th new input matrix Θ p ( p <= n , n is Θ 's total number of groups) and the known input point set Θ d ( Θ d ∈ R m×d ) is the covariance matrix between each group of data.
[0058] By inputting Θ P to obtain the observed data, combined with the prior distribution, calculate Θ P 's posterior distribution mean and variance: (18), (19), Among them, μ is the mean of the posterior distribution, V is the variance of the posterior distribution, k ( Θ P )( k ( Θ P ) ∈ R d×1 ) is the observed value vector calculated by the kernel function, K is from the kernel function K ( Θ P , Θ d ) defined covariance matrix ( K ∈ R d×d ), B is the noise variance, used to simulate the random noise in the observed values of the objective function, I is the identity matrix ( I ∈ R d×d ), Q ( Q ∈ R d×1 ) is the known set of objective function values. By inputting Θ P to obtain the observed data, combined with the prior distribution, calculate ΘP The mean and variance of the posterior distribution.
[0059] Then, multiple candidate points for parallel evaluation are selected by maximizing the acquisition function. The parallel adaptive acquisition function is: (20), where D n ∈ x , i = 1, 2, 3, k = 1, 2, 3, … n . By using the acquisition function, the points for which the mean and variance are to be calculated next are selected, and the points that meet the accuracy are finally determined by continuously iterating the acquisition function. Equation (20) means that the left side of the equation represents the entire acquisition function, which is divided into three types, and the right side is divided into three acquisition functions a 1 , a 2 , a 3 Specifically, the acquisition functions in different stages are as follows.
[0060] (1) In the initial stage of optimization, due to limited information about the objective function, the upper confidence bound (UCB) function is adopted to enhance the exploration ability, quickly cover the parameter space, and identify potential advantageous regions. The UCB function is: (21), where μ ( Θ P ) is the predicted mean of the objective function at the point Θ P , v ( Θ P ) is the predicted standard deviation at the point Θ P , f is a constant parameter that controls the balance between exploration and exploitation.
[0061] (2) As information accumulates, the expected improvement (EI) function is adopted to balance exploration and exploitation. The EI function is: (22), where Θ P + is the input vector corresponding to the currently known optimal parameter.
[0062] (3) In the later stage of optimization, the probability of improvement (PI) function is adopted to focus on fine search in the identified potential regions to achieve precise adjustment of the optimal solution. The PI function is: (23), Among them, P is a probability metric, γ is a non - negative parameter for controlling the exploration intensity. Finally, repeat the above steps to obtain a new ([[]] Θ P , f ( Θ P ))). According to equations (18) and (19), update the posterior distribution until the accuracy requirement is met.
[0063] Step 6: Model evaluation. Evaluate the model through two indicators: the normalized root - mean - square error (NRMSE) and the coefficient of determination (R 2 ): (24), (25), where, y max is the maximum value in the observed data; y min is the minimum value in the observed data; K is the total number of observed values; y j is the average value of the actual observed values.
[0064] Table 1 and Table 2 are respectively the comparison of the prediction results of different methods for the effluent ammonia - nitrogen concentration in summer and winter. It can be seen from Table 1 and Table 2 that the R2 and NRMSE in this embodiment are better than other methods in terms of prediction effect both in summer and winter.
[0065] Table 1 Prediction results of different models for the effluent ammonia - nitrogen concentration in summer: .
[0066] Table 2 Prediction results of different models for the effluent ammonia - nitrogen concentration in winter: .
[0067] Through Figure 3 the online prediction result graph of the effluent ammonia - nitrogen concentration in summer and Figure 4 the online prediction result graph of the effluent ammonia - nitrogen concentration in winter, the high precision and generalization ability of this embodiment in predicting the effluent ammonia - nitrogen concentration are demonstrated.
[0068] The above - mentioned embodiments are only used to illustrate the structural concept and characteristics of the present invention, and their purpose is to enable ordinary technicians in the field to understand the content of the present invention and implement it accordingly, and cannot be used to limit the protection scope of the present invention. Any equivalent changes or modifications made according to the essence of the content of the present invention should be covered within the protection scope of the present invention.
Claims
1. A method for measuring effluent ammonia nitrogen concentration based on neural network and Bayesian optimization, characterized in that: The following steps are involved: Step 1: Construct a basic model with a temporal convolutional layer, an attention mechanism, and a bidirectional long short-term memory neural network, collect characteristic variables x related to the prediction of effluent ammonia nitrogen concentration, calculate the Pearson correlation coefficient, retain the characteristic variables x according to the Pearson correlation coefficient, pre-process the retained characteristic variables x, and use the Min-Max method to normalize each characteristic variable x to the interval [0, 1]; Step 2: Input the feature variable x into the basic model for feature extraction and enhancement, extract the local dependencies in the time series data through the temporal convolution layer TCL, recalibrate the features through the SE attention mechanism, enhance the response of key features and suppress noise; Step 3: The feature variable x after feature extraction and enhancement is extracted through a bidirectional long short-term memory neural network with an adaptive mechanism to extract the long-term time series features of the feature sequence, dynamically adjust the input step size, adapt to the changes in different time periods in the data, capture the bidirectional time dependency in the sequence, and output the predicted value; Step 4: Parallel Bayesian optimization algorithm dynamically adjusts the acquisition function according to the different stages of the optimization process, evaluates multiple candidate points at the same time, adaptively acquires the function in parallel, and outputs the measurement value.
2. The method for measuring effluent ammonia nitrogen concentration based on neural network and Bayesian optimization according to claim 1 is characterized in that: In step 1, the calculation formula for calculating the Pearson correlation coefficient is: ,in, r is the Pearson correlation coefficient, x i and y i is the first of two variables i Observations, x m and y m The variables are x and y The mean of n is the number of samples, and the characteristic variables include the final selection of aerobic terminal dissolved oxygen, aerobic terminal total solid suspended matter, effluent pH value, effluent redox potential, and effluent nitrate nitrogen.
3. The method for measuring effluent ammonia nitrogen concentration based on neural network and Bayesian optimization according to claim 1, characterized in that: The preprocessing includes first performing data cleaning on the collected original data to deal with missing values and abnormal values. For the missing values, different filling strategies are adopted according to the data characteristics. The continuous variables are interpolated using the time series method, and the values of the adjacent time points are used for reasonable filling. For the abnormal values, the distribution range of each feature is calculated to identify abnormal points in combination with statistical methods and business experience, and the sliding window median replacement method is used to correct them.
4. The method for measuring effluent ammonia nitrogen concentration based on neural network and Bayesian optimization according to claim 1, characterized in that: The step 2 comprises: Step 2-1: Extract the local dependencies of time series data through the temporal convolution layer TCL. The output of TCL x c The calculation formula is: ,in, δ (⋅) is the ReLU activation function, which is used to introduce nonlinear characteristics. W k,c,c′ is the weight matrix of the convolution kernel, x t+k,c is the input data at time step t + k and input channels c The value at b c′ is the bias term, b ∈R c′ ; Step 2-2: Recalibrate the feature variable x through the SE attention mechanism to enhance the response of key features and suppress noise, perform channel recalibration, and apply the weight s obtained after excitation to the original input features.
5. The method for measuring effluent ammonia nitrogen concentration based on neural network and Bayesian optimization according to claim 4 is characterized in that: The step 2-2 comprises: Step a: compress and generate a channel descriptor using global average pooling z c′′ , used to emphasize the spatial information of global distribution, the formula is: ,in, x ijc '' indicates the height is i Width is j Place c The eigenvalues of ′′ channels, H and W Represent the height and width of the input channel respectively; Step b: The SE attention mechanism reduces parameters and complexity by using the compression ratio r and increases nonlinearity using the ReLU activation function; Step c: Calculate the weight of each channel through the sigmoid function to achieve effective recalibration of features, and use the weights of the two fully connected layers M 1 and M 2 Calculate the output of the stimulus s for: ,in, M 1∈R c′′×c′′ / r , M 2∈R c′′ / r×c′′ , z Depend on c '' channel descriptors, z ∈R c′′×1 is the output of the compression step, s ∈R c′′×1 , σ is the sigmoid activation function; Step d: Perform channel recalibration and convert the weights obtained after excitation s Applied to the original input features, for each channel, the recalibrated output is: ,in, s c′′ It is a channel c The weight of ′′, x c′′ It is a channel c The original input features of ′′, x s is the output of SE after recalibration.
6. The method for measuring effluent ammonia nitrogen concentration based on neural network and Bayesian optimization according to claim 1, characterized in that: The step 3 comprises: Step 3-1: The bidirectional long short-term memory neural network ABiLSTM processes the input sequence time-step by time step through the LSTM unit. The LSTM unit includes an input gate, a forget gate, an output gate, and an update of the cell state. The algorithm formula is as follows: , , , , , , in, x t ( x t ∈R m×1 ) is the input vector received, h t ( h t ∈R m×1 ) is the hidden state at time t, h t−1 is the hidden state of the previous time step, g t is the new candidate value, f t , i t , o t They are the forget gate, input gate and output gate respectively. t Output at the moment, c t for t The cell state at each moment, W f , W O , W i , W C ( W ∈R h×m ) respectively represent t Connect input at this moment x t To the weight matrices of different gates, b f , b i , b C , b O ( b ∈R h×1 ) is the bias term; Step 3-2: Define a minimum length Δ t min The window size of the prediction error is dynamically adjusted based on the input step size, and only the data of the time step contained in the window is processed each time. The window size of the dynamic adjustment of the prediction error is: ,in, α is the adjustment factor, error t is the mean square error of the current time step; Step 3-3: Capture the bidirectional temporal characteristics of the data through forward and backward propagation. x t ∈R m×η , the specific operation process of ABiLSTM can be expressed as: , , , in, H t1 , H t2 , H t Represent the forward, backward and hidden states of ABiLSTM respectively, O t is the ABiLSTM output, W f xh , W b xh ∈R h×m , W f hh , W b hh ∈R h×h , W hq ∈R q×2h represents the weight matrix of the model, b f h , b b h ∈R h×η , b q ∈R q×2η represents the bias term of the model, q is the final output dimension.
7. The method for measuring effluent ammonia nitrogen concentration based on neural network and Bayesian optimization according to claim 6, characterized in that: The mean square error error of the current time step t for: ,in: y t yes t The true value of the moment, y p yes t The predicted value at time.
8. The method for measuring effluent ammonia nitrogen concentration based on neural network and Bayesian optimization according to claim 1, characterized in that: In step 4, the parameter combination to be optimized by the model Θ The number of ABiLSTM layers, the number of hidden units, and the compression ratio of SE r The weight matrix of the ABiLSTM layer and the TCL layer is optimized by optimizing the L2 regularization parameter. W With the bias matrix b .
9. The method for measuring effluent ammonia nitrogen concentration based on neural network and Bayesian optimization according to claim 8, characterized in that: The step 4 comprises: Step 4-1: Initialize the prior distribution of the Gaussian process and select the initial candidate point set: ,in, m ( Θ p ) is the mean function, Θ p Indicates p A new input matrix, Θ d represents a set of known input points, k ( Θ p , Θ d ) is the Gaussian kernel function, indicating Θ p and Θ d The matrix of covariance between each set of data in ; Step 4-2: By input Θ P Get the observed data, combine it with the prior distribution, and calculate Θ P The formulas for the mean and variance of the posterior distribution are: , , in, μ is the mean of the posterior distribution, V is the variance of the posterior distribution, k ( Θ P )and( k ( Θ P ) is the observation value vector calculated by the kernel function, K The kernel function K ( Θ P , Θ d ) defines the covariance matrix, B is the noise variance, I is the identity matrix, Q is a set of known objective function values; Step 4-3: Select multiple parallel evaluation candidate points by maximizing the acquisition function. The parallel adaptive acquisition function is: , in, D n represents a known observation dataset, D n ∈ x , a i represents the ith acquisition function, i =1,2,3, k =1, 2, 3, … n ; Step 4-4: According to the calculation Θ P The formula for the mean and variance of the posterior distribution is used to update the posterior distribution until the accuracy requirement is met.
10. The method for measuring effluent ammonia nitrogen concentration based on neural network and Bayesian optimization according to claim 9, characterized in that: The acquisition function includes: In the early stage of optimization, due to the limited information of the objective function, the upper confidence bound UCB function is used to enhance the exploration ability, quickly cover the parameter space and identify potential advantage areas. The UCB function is: ,in, μ ( Θ P ) is at point Θ P The predicted mean of the objective function at , v ( Θ P ) is at point Θ P The standard deviation of the prediction at f is a constant parameter that controls the balance between exploration and exploitation; In the middle of optimization, as information accumulates, we turn to the expected improvement of the EI function to balance exploration and development. The EI function is: ,in, Θ P + is the input vector corresponding to the currently known optimal parameters; In the later stage of optimization, the probability-improved PI function is used to focus on the identified potential areas for fine search to achieve precise adjustment of the optimal solution. The PI function is: ,in, P is a probability measure, γ is a non-negative parameter that controls the intensity of exploration.
Citation Information
Patent Citations
BiLSTM voltage deviation prediction method based on Bayesian optimization
CN113554148A
Convolutional layer and self-attention mechanism-based effluent ammonia nitrogen concentration measurement method
CN118228766A
Piezoelectric actuator temperature rise prediction method based on Bayesian optimization CNN-LSTM
CN119337082A
Private domain live broadcast peak hot spot prediction and content scheduling method based on deep learning
CN119450099A
Power demand dynamic prediction method based on multi-scale hybrid architecture
CN119990477A
Cited By
Conv-BiLSTM and LS-based simultaneous same-frequency full-duplex digital domain adaptive self-interference elimination method and system
CN120710526A
Dynamic adjusting method for concentration of hydrogen-rich water
CN121433359A
Ammonia nitrogen prediction regulation and control system and method based on conditional diffusion and agent cooperation
CN122470649A
Ammonia nitrogen prediction and regulation system and method based on conditional diffusion and agent cooperation
CN122470649B