Wastewater Treatment Monitoring Method and System Based on Automatic Optimization Algorithm and Deep Learning
By using automatic optimization algorithm and deep learning methods in the anaerobic wastewater treatment system, BiGRU and GPR models are constructed, and the hysteresis and nonlinear problems of monitoring and prediction of anaerobic wastewater treatment system in the prior art are solved, and efficient water quality prediction and monitoring are achieved.
Patent Information
- Application Number
- CN202210508224.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-11
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-05-11
AI Technical Summary
It is difficult for the prior art to realize real-time monitoring and prediction of anaerobic wastewater treatment systems, especially in the determination of indicators such as COD, which leads to difficulty in predicting system abnormalities.
Using an automatic optimization algorithm and deep learning method, a bidirectional gated loop unit (BiGRU) model is constructed, and the tree structure Parzen estimation method (TPE) is combined with the automatic optimization hyperparameter search, and the Gaussian process regression (GPR) model is further constructed for interval and probability prediction.
Accurate prediction of the water quality of anaerobic wastewater treatment system is achieved, point prediction, interval prediction and probability prediction results are provided, real-time and credible monitoring are improved, and cumbersome process of manually adjusting hyperparameters is avoided.
Smart Images

Figure CN114944203B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an intelligent monitoring method for anaerobic wastewater treatment based on an automatic optimization algorithm and deep learning, belonging to the technical field of water body detection. Background Art
[0002] In the food industry, a large amount of organic wastewater is generated in the related processes of food production. These organic wastewaters are characterized by high chemical oxygen demand (COD) and high biodegradability, and anaerobic biochemical methods are usually used to degrade food wastewater.
[0003] The processes of anaerobic biochemical methods in sewage treatment plants include up-flow anaerobic sludge bed and internal circulation anaerobic reactor. These processes have been relatively mature. However, the mechanism of the reaction system is very complex. If not monitored and maintained, phenomena such as acidification, sludge washout, and insufficient gas production will occur in the system, which will further lead to the collapse of the reaction system.
[0004] For the monitoring of the current anaerobic reaction system, operators usually measure the water quality indicators of influent and effluent at regular intervals. However, for some indicators, such as COD, the measurement process is relatively long, and its monitoring data has a certain lag, making it difficult to predict or detect system anomalies in a timely manner; moreover, the water quality is affected by many factors, resulting in highly non-linear, fluctuating, and uncertain characteristics of the water quality, making it difficult to predict system anomalies.
[0005] In recent years, machine learning, especially deep learning, has been developing rapidly. The technology of using neural networks to predict the water quality of influent and effluent in sewage treatment plants has attracted the attention of many scientific researchers. The deep learning method bidirectional gated recurrent units (BiGRU) has a high point prediction accuracy when solving time series prediction problems such as wastewater treatment water quality, but it cannot perform interval prediction and probability prediction, and manually adjusting the model hyperparameters is very time-consuming and laborious. Specifically, the gated recurrent unit GRU is an improved type of recurrent neural network (RNN). It solves the long-term dependence problem existing in RNN when solving time series prediction problems by adding a reset gate and an update gate to the hidden layer of RNN. BiGRU adds a backpropagation mechanism on the basis of the GRU network, so that the output at the current moment can be related to the state at the previous moment and the state at the next moment to improve the prediction accuracy.
[0006] In addition, the process of constructing a traditional neural network inevitably involves the step of manually adjusting the network hyperparameters. R & D personnel need to manually adjust the hyperparameters of the neural network according to the prediction accuracy after one training of the neural network. This not only takes time and effort but may also fall into the local optimal dilemma and fail to obtain the global optimal model. Random search and grid search are also methods for automatic parameter optimization, but their respective combination randomness and global traversal determine their time-consuming disadvantages. The Tree-structured Parzen Estimator (TPE) is a method for automatic parameter optimization based on Bayesian optimization. Compared with random search and grid search, it uses heuristic search with strategies and can find the optimal hyperparameters in a shorter time. How to apply TPE to the hyperparameter adjustment process of BiGRU is one of the research issues.
[0007] Therefore, how to enable BiGRU to perform interval prediction and probability prediction and overcome the deficiencies of manual hyperparameter adjustment is an urgent theoretical and practical engineering problem to be solved. Summary of the Invention
[0008] The present invention provides an intelligent monitoring method for anaerobic wastewater treatment based on an automatic optimization algorithm and deep learning, aiming to solve at least one of the technical problems existing in the prior art.
[0009] The technical solution of the present invention relates to an intelligent monitoring method for anaerobic wastewater treatment based on an automatic optimization algorithm and deep learning. The method includes the following steps:
[0010] S100. Select historical detection data of the anaerobic treatment unit of the sewage treatment plant. The historical detection data includes input variables and output variables, and the historical detection data is recombined into a data set. The data set is divided into a training data set and a test data set, and then the training data set and the test data set are subjected to data preprocessing.
[0011] S200. Construct a bidirectional gated recurrent unit model with automatic optimization. Input the training data set into the bidirectional gated recurrent unit model for training. Use the Tree-structured Parzen Estimator method to automatically optimize the hyperparameters of the bidirectional gated recurrent unit model to obtain the optimal hyperparameters, and then input the optimal hyperparameters into the bidirectional gated recurrent unit model for training to obtain the optimal model.
[0012] S300. Input the test data set into the trained bidirectional gated recurrent unit model to obtain the point prediction result of the output variable.
[0013] S400. Construct a Gaussian process regression (GPR) model, input the point prediction result of the output variable into the trained GPR model, obtain the probability distribution function corresponding to the point prediction result of the output variable, and determine the prediction interval and probability prediction corresponding to each point prediction result of the output variable based on the probability distribution function.
[0014] Further, the input variables include COD, pH value, volatile fatty acid concentration value, organic matter load content, and alkalinity in the water body of the influent wastewater, and the output variables include COD and gas production in the water body of the effluent wastewater.
[0015] Further, step S100 includes:
[0016] S110. Select the historical detection data for a historical period, and reorganize the data set in the form of the following data set matrix:
[0017]
[0018] Among them, the row vector of the data set matrix is the wastewater index data of different categories, the column vector represents the wastewater index data of the same category at different historical moments, m is the data length of the wastewater index data of each category in the data set matrix, and k is the number of wastewater indexes.
[0019] Further, step S100 includes:
[0020] S120. After dividing the data set into a training data set and a test data set according to a ratio of 8:2, construct the training data set and the test data set once.
[0021] The training data set is used to train the bidirectional gated recurrent unit model, and the test data set is used to test the prediction accuracy of the bidirectional gated recurrent unit model.
[0022] S130. Screen and remove the outliers in the data set, and then perform normalization processing on the data set through the following formula:
[0023]
[0024] Among them, i is the dimension of the input variable, X 1i,0 is the original data value of the i-th dimension in the input variable, minX 1i,0 is the minimum value of the original data of the i-th dimension in the input variable, maxX 1i,0 is the maximum value of the original data of the i-th dimension in the input variable, and X 1i is the normalized data value of the i-th dimension in the input variable.
[0025] Further, wherein the step S200 includes:
[0026] S210. Set the neural network hyperparameters of the bidirectional gated recurrent unit model,
[0027] wherein, select some of the neural network hyperparameters as the hyperparameters to be optimized,
[0028] and wherein, the bidirectional gated recurrent unit model includes an input layer, a hidden layer and an output layer, and the hidden layer includes an attention mechanism layer, a bidirectional layer and a GRU layer;
[0029] S220. Set the search space of the hyperparameters to be optimized;
[0030] S230. Combine each hyperparameter in the search space of the hyperparameters;
[0031] S240. Train the bidirectional gated recurrent unit model, automatically optimize the hyperparameters according to the accuracy of multiple trainings, and finally obtain the optimal hyperparameters and the optimal model.
[0032] Further, wherein the hyperparameters include the number of neurons in the input layer, the number of neurons in the hidden layer, the number of neurons in the output layer, the learning rate, the batch size and the number of iterations.
[0033] Further, wherein the step S240 includes:
[0034] S241. According to the set optimized hyperparameters and the corresponding hyperparameter space, in the first training, use the tree-structured Parzen estimation method to randomly search and combine each hyperparameter in the hyperparameter space, and create a set of error observation values {x (i) ,y (i) ,i = 1,2,...,N init},
[0035] wherein, x is the combination of hyperparameters to be optimized, and y is the error value obtained by training with the corresponding hyperparameter combination;
[0036] S242. According to the results and training accuracy of the first N trainings of the bidirectional gated recurrent unit model, use the tree-structured Parzen estimation method to set an error quantile value y * , and divide the set of error observation values into two parts, and calculate its probability value through the following formula:
[0037]
[0038] wherein, l(x) is the probability density function of the set of hyperparameters with error values less than y*, and g(x) is the probability density function of the set of hyperparameters with error values less than y*;
[0039] S243. Calculate the EI value using the following formula:
[0040]
[0041] S244. Select the next set of hyperparameter values by maximizing the EI value;
[0042] S245. Repeat steps S241 to S244 until the bidirectional gated recurrent unit model reaches the set number of training times and then ends.
[0043] Furthermore, wherein, the step S400 includes the following steps:
[0044] S410. Use the input variables in the previous training set and test set as the input variables in the current training set and test set, and use the point prediction results of the output variables to construct the output variables for the next training set and test set.
[0045] S420. Input the next training set and test set into the Gaussian process regression GPR model to obtain a probability distribution function corresponding to the point prediction results of the output variables, and the probability distribution function follows a Gaussian distribution.
[0046] S430. Based on the mean, standard deviation, and preset confidence level of the probability distribution function, determine the prediction interval of each point prediction result at the preset confidence level.
[0047] Furthermore, wherein, compare the calculated metrics of point prediction, interval prediction, and probability prediction with the BiGRU-GPR model, BiLSTM-GPR model, and Gaussian process regression GPR model to obtain a better model. The metric of point prediction is calculated according to the following formula:
[0048]
[0049]
[0050]
[0051] wherein, y i is the i-th observation value, is the average value of the i-th observation value, is the i-th predicted value, and n is the number of prediction samples;
[0052] The metric of interval prediction is calculated according to the following formula:
[0053]
[0054]
[0055] MC = MWP / CP
[0056] Wherein, y upper,i is the upper limit of the prediction interval of the predicted value of the i-th point, and y lower,i is the lower limit of the prediction interval of the predicted value of the i-th point, and n j is the number of observed values within the prediction interval.
[0057] The index of the probability prediction is calculated according to the following formula:
[0058]
[0059]
[0060]
[0061] Wherein, F(y i ) is the cumulative distribution function of the predicted value, is the unit step function.
[0062] The technical solution of the present invention also relates to a computer device, including a memory and a processor. When the processor executes the computer program stored in the memory, the above method is implemented.
[0063] The beneficial effects of the present invention are as follows.
[0064] 1. The intelligent monitoring method for anaerobic wastewater treatment based on the automatic optimization algorithm and deep learning can accurately predict the water quality of the wastewater body. According to the changes of various indicators in the anaerobic reaction process, the above method predicts the COD in the water and the gas production of the water body, and at the same time gives the interval prediction and probability prediction corresponding to the point prediction result, and the output result has good credibility.
[0065] 2. Aiming at the pain point that researchers need to manually adjust the hyperparameters of the neural network, the present invention introduces the tree-structured Parzen estimation method to automatically optimize the hyperparameters of the neural network, so that the neural network model reaches the optimal state. This method not only saves time for researchers to manually adjust the hyperparameters, but also can effectively improve the prediction accuracy of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Figure 1 is the overall flowchart of the method according to the present invention.
[0067] Figure 2 is the schematic diagram of the GRU hidden layer structure in the embodiment of the present invention.
[0068] Figure 3 is the implementation schematic diagram of the division and prediction of the data set in the embodiment of the present invention.
[0069] Figure 4 It is the BiGRU neural network structure diagram according to the embodiment of the present invention.
[0070] Figure 5 It is the comparison chart of COD point prediction results according to the embodiment of the present invention.
[0071] Figure 6 It is the COD interval prediction result graph according to the embodiment of the present invention.
[0072] Figure 7 It is the biogas production point prediction result graph according to the embodiment of the present invention.
[0073] Figure 8 It is the biogas production interval prediction result graph according to the embodiment of the present invention.
[0074] Figure 9 It is the comparison table of point prediction results of three models according to the embodiment of the present invention.
[0075] Figure 10 It is the comparison table of interval prediction results of three models according to the embodiment of the present invention.
[0076] Figure 11 It is the comparison table of probability prediction results of three models according to the embodiment of the present invention. Detailed implementation manners
[0077] The concept, specific structure and technical effects of the present invention will be clearly and completely described below in combination with the embodiments and the drawings to fully understand the purpose, solution and effects of the present invention.
[0078] The nouns and terms involved in the embodiments of the present invention are explained as follows:
[0079] Tree-structured Parzen Estimator, TPE: Tree-based Parzen estimation method
[0080] Bidirectional Gated Recurrent Units, BiGRU: Bidirectional gated recurrent unit
[0081] Gaussian Progress Regression, GPR: Gaussian process regression
[0082] COD: Chemical oxygen demand
[0083] VFA: Volatile fatty acid concentration value
[0084] OLR: Organic matter load content
[0085] ALK: Alkalinity
[0086] CP: Interval Coverage Probability
[0087] MWP: Mean Width of Interval
[0088] CRPS: Continuous Ranked Probability Score
[0089] Reference Figures 1 to 11 , in some embodiments, the present invention discloses an intelligent monitoring method for anaerobic wastewater treatment based on an automatic optimization algorithm and deep learning. The method includes the following steps. Reference Figure 1 to the flowchart of
[0090] S100. Select historical detection data of the anaerobic treatment unit of the sewage treatment plant. The historical detection data includes input variables and output variables, and reorganize the historical detection data into a data set. The data set is divided into a training data set and a test data set, and then data preprocessing is performed on the training data set and the test data set;
[0091] S200. Construct an automatically optimized bidirectional gated recurrent unit model, input the training data set into the bidirectional gated recurrent unit model for training, use the tree-structured Parzen estimation method to automatically optimize the hyperparameters of the bidirectional gated recurrent unit model to obtain the optimal hyperparameters, and then use the optimal hyperparameters to input the bidirectional gated recurrent unit model for training to obtain the optimal model;
[0092] S300. Input the test data set into the trained bidirectional gated recurrent unit model to obtain the point prediction result of the output variable;
[0093] S400. Construct a Gaussian process regression (GPR) model, input the point prediction result of the output variable into the trained Gaussian process regression (GPR) model to obtain the probability distribution function corresponding to the point prediction result of the output variable, and determine the prediction interval and probability prediction corresponding to each point prediction result of the output variable based on the probability distribution function.
[0094] For step S100
[0095] S100. Select historical detection data of the anaerobic treatment unit of the sewage treatment plant. The historical detection data includes input variables and output variables, and reorganize the historical detection data into a data set. This data set is used as the data set for deep learning. The data set is divided into a training data set and a test data set, and then data preprocessing is performed on the training data set and the test data set.
[0096] Determined according to historical detection data, the input variables include COD, pH value, volatile fatty acid concentration value, organic matter load content, and alkalinity in the water body entering the wastewater, and the output variables include COD and gas production in the water body flowing out of the wastewater.
[0097] S110. Select the historical detection data for a historical period, and reorganize the data set in the form of the following data set matrix:
[0098]
[0099] Among them, the row vectors of the data set matrix are wastewater index data of different categories, the column vectors represent wastewater index data of the same category at different historical moments, m is the data length of the wastewater index data of each category in the data set matrix, and k is the number of wastewater indexes.
[0100] S120. After dividing the data set into a training data set and a test data set according to a ratio of 8:2, construct the training data set and the test data set once. Among them, the input variables of the training data set once are The output variables are The input variables of the test data set once are The output variables are The training data set once is used to train the bidirectional gated recurrent unit model, and the test data set once is used to test the prediction accuracy of the bidirectional gated recurrent unit model.
[0101] S130. Screen and remove the outliers in the data set, and then perform normalization processing on the data set through the following formula:
[0102]
[0103] Among them, i is the dimension of the input variable, X 1i,0 is the original data value of the i-th dimension in the input variable, minX 1i,0 is the minimum value of the i-th dimension of the original data in the input variable, maxX 1i,0 is the maximum value of the i-th dimension of the original data in the input variable, X 1i is the normalized data value of the i-th dimension in the input variable.
[0104] For step S200
[0105] S200. Construct a bidirectional gated recurrent unit model with automatic optimization. Input the training data set into the bidirectional gated recurrent unit model for training. Use the tree-structured Parzen estimation method to automatically optimize the hyperparameters of the bidirectional gated recurrent unit model to obtain the optimal hyperparameters. Then, use the optimal hyperparameters to input into the bidirectional gated recurrent unit model for training to obtain the optimal model.
[0106] S210. Refer to Figure 4 , set the neural network hyperparameters of the bidirectional gated recurrent unit model, and select some of the neural network hyperparameters as the hyperparameters to be optimized.
[0107] Construct a bidirectional gated recurrent unit (BiGRU) model. The bidirectional gated recurrent unit model includes an input layer, a hidden layer, and an output layer. The hidden layer includes an attention mechanism layer (Attention layer), a bidirectional layer, and a GRU layer.
[0108] S220. Set the search space of the hyperparameters to be optimized. The hyperparameters include the number of neurons n i in the input layer, the number of neurons n h in the hidden layer, the number of neurons n o in the output layer, the learning rate L, the batch size B, and the number of iterations E. Among them, n i and n o are the numbers of input variables and output variables respectively, that is, 5 and 1.
[0109] S230. Combine the hyperparameters in the search space of the hyperparameters. n o , L, B, and E are designated as the hyperparameters to be optimized by TPE, and their search ranges are set respectively before training. The optimizer of the bidirectional gated recurrent unit (BiGRU) model is Adam, and the loss function is the mean square error MSE.
[0110] S240. Train the bidirectional gated recurrent unit model, automatically optimize the hyperparameters according to the accuracy of multiple trainings, and finally obtain the optimal hyperparameters and the optimal model. Among them, the calculation formula for the forward propagation of information at time t is as follows, refer to Figure 2 as shown.
[0111] r t = σ(W r · h t-1 + W r · x t )
[0112] z t = σ(W z · h t-1 + W z · x t )
[0113]
[0114]
[0115] y t = σ(W o ·h t )
[0116] where z t is the update gate, r t is the reset gate, x t is the input information, h t is the current state, h t-1 is the information of the previous state, is the candidate set, y t is the output information, W r 、W z 、 and W O represent the corresponding weights respectively, σ(·) and tanh(·) are activation functions, and ⊙ is the matrix product.
[0117] where W r 、W z and are all concatenated, and they are split by the following formula:
[0118] W r = W rx + W rh
[0119] W z = W zx + W zh
[0120]
[0121] The calculation formula for error backpropagation at time t is as follows:
[0122] First, calculate the loss function transmitted by the network at time t:
[0123]
[0124] Then the loss function at all times is:
[0125]
[0126] Next, calculate the error of the output layer and find the partial derivatives of the loss function with respect to each parameter:
[0127]
[0128]
[0129]
[0130]
[0131]
[0132]
[0133]
[0134] After calculating the partial derivatives with respect to each parameter, the parameters can be updated, and the above process is iterated successively until the loss converges, which means the training is completed.
[0135] After completing one training of the bidirectional gated recurrent unit model, the hyperparameters will be adjusted according to the output result and the output accuracy. Until the set number of training times is reached, the automatic optimization ends, and the optimal model is output. The bidirectional gated recurrent unit model (TPE) follows the principle of the following steps, and the step S240 includes:
[0136] S241. According to the set optimized hyperparameters and the corresponding hyperparameter space, in the first training, use the tree-structured Parzen estimation method to randomly search for combinations of hyperparameters within the hyperparameter space and create a set of error observation values {x (i) , y (i) , i = 1, 2,..., N init},
[0137] where x is the combination of hyperparameters to be optimized, and y is the error value obtained by training with the corresponding hyperparameter combination;
[0138] S242. According to the results and training accuracy of the bidirectional gated recurrent unit model in the previous N trainings, use the tree-structured Parzen estimation method to set an error quantile value y * , and divide the set of error observation values into two parts. Calculate its probability value through the following formula:
[0139]
[0140] where l(x) is the probability density function of the set of hyperparameters with error values less than y*, and g(x) is the probability density function of the set of hyperparameters with error values less than y*;
[0141] S243. Calculate the EI value through the following formula:
[0142]
[0143] S244. Select the next set of hyperparameter values by maximizing the EI value;
[0144] S245. Repeat steps S241 to S244 until the bidirectional gated recurrent unit model reaches the set number of training times and then ends.
[0145] For step S300
[0146] S300. Input the test data set into the trained bidirectional gated recurrent unit model to obtain the point prediction result of the output variable.
[0147] For step S400
[0148] S400. Construct a Gaussian process regression (GPR) model, input the point prediction result of the output variable into the trained Gaussian process regression (GPR) model to obtain the probability distribution function corresponding to the point prediction result of the output variable, and determine the prediction interval and probability prediction corresponding to each point prediction result of the output variable based on the probability distribution function.
[0149] Step S400 includes the following steps:
[0150] S410. Use the input variables in the previous training set and test set as the input variables in the current training set and test set, and use the point prediction result of the output variable to construct the output variable for the next training set and test set.
[0151] S420. Input the current training set and test set into the Gaussian process regression (GPR) model to obtain the probability distribution function corresponding to the point prediction result of the output variable, and the probability distribution function follows a Gaussian distribution. Among them, the probability distribution function of the i-th sample point in the current test set Among them, is the point prediction result of the i-th sample, and the point prediction, interval prediction, and probability prediction results of the i-th sample point are obtained.
[0152] S430. Based on the mean, standard deviation, and preset confidence level of the probability distribution function, determine the prediction interval of each point prediction result under the preset confidence level.
[0153] Combined with Figure 3 , specifically, in the process of calculating the probability distribution function, for the sake of simplicity of explanation, is set as x, is set as y, is set as x * , is set as f(x * ) = y * ;
[0154] In the case where there is noise at the observation point, it is assumed that the noise follows a normal distribution with a mean of 0 and a variance of For a normal distribution, we have:
[0155]
[0156] The Gaussian process regression assumes that the sample function values follow a joint normal distribution, that is:
[0157]
[0158] where K is the kernel function, * and ** are symbols to distinguish different parameters K, and I N is the identity matrix.
[0159] Calculate the posterior probability of y * according to the following Bayesian regression formula:
[0160]
[0161]
[0162]
[0163] where, now let That is, the required probability distribution function is obtained.
[0164] Then, use the following formula to calculate the upper and lower limits of the 95% prediction interval for the i-th sample point:
[0165]
[0166]
[0167] Among them, the explanations of the terms in the calculation of point prediction indicators are as follows:
[0168] (1) Coefficient of determination R 2 : It is used to measure the deviation degree between the predicted value and the true value. It is between 0 and 1. The closer it is to 1, the more consistent the predicted value is with the true value.
[0169] (2) Root mean square error RMSE: It is used to calculate the square root of the ratio of the sum of the squares of the deviations between the predicted value and the observed value to the number of observations. The larger the RMSE, the greater the error of the predicted value.
[0170] (3) Mean absolute percentage error MAPE: It is used to calculate the percentage of the mean absolute error between the predicted value and the observed value. The smaller the MAPE, the more perfect the prediction model. If it is greater than 1, the prediction model is a poor model.
[0171] The indicators of the said point prediction are calculated according to the following formula:
[0172]
[0173]
[0174]
[0175] where y i is the i-th observed value, is the average value of the i-th observed value, is the i-th predicted value, and n is the number of prediction samples;
[0176] Among them, the glossary in calculating the interval prediction:
[0177] (1) Interval coverage rate CP: used to calculate the percentage of the prediction interval covering the observed value. The closer CP is to 1, the more observed values the interval covers;
[0178] (2) Mean width of prediction interval MWP: used to calculate the average width of the prediction interval. The smaller MWP is, the higher the reliability of the interval prediction;
[0179] (3) Comprehensive index MC of interval prediction: an index that combines MWP and CP. The smaller the MC value, the better the interval prediction effect.
[0180] The indexes of the interval prediction are calculated according to the following formula:
[0181]
[0182]
[0183] MC = MWP / CP
[0184] where y upper,i is the upper limit of the prediction interval of the i-th point prediction value, y lower,i is the lower limit of the prediction interval of the i-th point prediction value, n j is the number of observed values within the prediction interval,
[0185] Among them, the glossary in calculating the probability prediction index:
[0186] (1) Continuous ranked probability score CRPS: used to measure the difference between the prediction distribution and the true distribution. The closer CRPS is to 0, the higher the consistency between the prediction distribution and the true distribution.
[0187] The indexes of the probability prediction are calculated according to the following formula:
[0188]
[0189]
[0190]
[0191] Among them, F(y i ) is the cumulative distribution function of the predicted value, is the unit step function.
[0192] Compare the calculated indicators of point prediction, interval prediction and probability prediction with the BiGRU-GPR model, BiLSTM-GPR model and Gaussian process regression GPR model to obtain a better model, combined with Figures 5 to 8 .
[0193] Specifically, referring to Figure 8 the point prediction results in the table: In point prediction, for the prediction of COD and biogas production, the prediction accuracies of the above three models are very high, and the prediction curves basically coincide with the observed value curves to a very high degree, indicating that the BiGRU-GPR model and BiLSTM-GPR model deep learning methods that automatically optimize using the tree-structured Parzen estimation method (TPE) not only eliminate the complexity of manually adjusting hyperparameters but also can obtain very high prediction accuracies. Among them, the three indicators of the BiGRU-GPR model are better than those of the BiLSTM-GPR model and the Gaussian process regression GPR model, indicating that the point prediction results obtained by the method of the present invention have the highest accuracy.
[0194] Referring to Figure 9 the interval prediction results in the table: In interval prediction, for the prediction of COD and biogas production, the comparison of the above three models shows the same trend: they all have a relatively high CP, and the BiLSTM-GPR model is slightly better; for MWP, the prediction interval of the BiGRU-GPR model is the narrowest; for the comprehensive index MC, the BiGRU-GPR is the lowest, indicating that the interval prediction results obtained by the method of the present invention have the highest accuracy and reliability.
[0195] Referring to Figure 10 the probability prediction results in the table: In probability prediction, for COD, the CRPS value of the BiGRU-GPR model is the smallest, which is 0.0329; for biogas production, the CRPS value of the BiLSTM-GPR model is the smallest, which is 0.1297. It shows that the distribution of the predicted values obtained by the method of the present invention has a high consistency with the true value distribution.
[0196] It should be recognized that the method steps in the embodiments of the present invention can be implemented or carried out by computer hardware, a combination of hardware and software, or computer instructions stored in a non-transitory computer-readable memory. The method can use standard programming techniques. Each program can be implemented in a high-level procedural or object-oriented programming language to communicate with the computer system. However, if desired, the program can be implemented in assembly or machine language. In any case, the language can be a compiled or interpreted language. In addition, for this purpose, the program can run on a dedicated integrated circuit programmed for this purpose.
[0197] In addition, the operations of the processes described herein can be performed in any suitable order, unless otherwise indicated herein or otherwise clearly contradicted by the context. The processes described herein (or variations and / or combinations thereof) can be executed under the control of one or more computer systems configured with executable instructions, and can be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) executed jointly on one or more processors, by hardware, or a combination thereof. The computer program includes a plurality of instructions executable by one or more processors.
[0198] Furthermore, the method can be implemented in any type of computing platform operably connected, including but not limited to personal computers, minicomputers, mainframes, workstations, network or distributed computing environments, separate or integrated computer platforms, or communicating with charged particle tools or other imaging devices, etc. Aspects of the present invention can be implemented in machine-readable code stored on a non-transitory storage medium or device, whether removable or integrated into the computing platform, such as a hard disk, optical read and / or write storage medium, RAM, ROM, etc., such that it can be read by a programmable computer and can be used to configure and operate the computer to execute the processes described herein when the storage medium or device is read by the computer. In addition, the machine-readable code, or portions thereof, can be transmitted via a wired or wireless network. When such media includes instructions or programs that implement the above-described steps in combination with a microprocessor or other data processor, the inventions described herein include these and other different types of non-transitory computer-readable storage media. When programmed according to the methods and techniques of the present invention, the present invention can also include the computer itself.
[0199] The computer program can be applied to the input data to perform the functions described herein, thereby transforming the input data to generate output data stored in a non-volatile memory. The output information can also be applied to one or more output devices such as a display. In a preferred embodiment of the present invention, the transformed data represents physical and tangible objects, including a specific visual depiction of the physical and tangible objects generated on the display.
[0200] As described above, these are only the preferred embodiments of the present invention. The present invention is not limited to the above-mentioned embodiments. As long as the same means are used to achieve the technical effects of the present invention, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the scope of protection of the present invention. Within the scope of protection of the present invention, various different modifications and changes can be made to its technical solutions and / or implementation manners.
Claims
1. An intelligent monitoring method for anaerobic wastewater treatment based on an automatic optimization algorithm and deep learning, characterized in that, the method comprises the following steps: S100. Select the historical detection data of the anaerobic treatment unit of the sewage treatment plant. The historical detection data includes input variables and output variables, and reorganize the historical detection data into a data set. The data set is divided into a training data set and a test data set, and then data preprocessing is performed on the training data set and the test data set; S200. Construct an automatically optimized bidirectional gated recurrent unit model, input the training data set into the bidirectional gated recurrent unit model for training, use the tree-structured Parzen estimation method to automatically optimize the hyperparameters of the bidirectional gated recurrent unit model to obtain the optimal hyperparameters, and then input the optimal hyperparameters into the bidirectional gated recurrent unit model for training to obtain the optimal model; S300. Input the test data set into the trained bidirectional gated recurrent unit model to obtain the point prediction result of the output variable; S400. Construct a Gaussian process regression GPR model, input the point prediction result of the output variable into the trained Gaussian process regression GPR model to obtain the probability distribution function corresponding to the point prediction result of the output variable, and determine the prediction interval and probability prediction corresponding to each point prediction result of the output variable based on the probability distribution function; wherein, the input variables include COD, pH value, volatile fatty acid concentration value, organic matter load content and alkalinity in the water body of the incoming wastewater, and the output variables include COD and gas production in the water body of the outgoing wastewater.
2. The method according to claim 1, wherein, the step S100 includes: S110. Select the historical detection data of a historical period, and reorganize the data set in the form of the following data set matrix: wherein, the row vectors of the data set matrix are wastewater index data of different categories, and the column vectors represent wastewater index data of the same category at different historical moments. m is the data length of the wastewater index data of each category in the data set matrix, and k is the number of wastewater indexes.
3. The method according to claim 2, wherein, the step S100 includes: S120. After dividing the data set into a training data set and a test data set according to a ratio of 8:2, construct the training data set and the test data set once. The training data set is used to train the bidirectional gated recurrent unit model, and the test data set is used to test the prediction accuracy of the bidirectional gated recurrent unit model; S130. Screen and remove the outliers in the data set, and then perform normalization processing on the data set through the following formula: where i is the dimension of the input variable, X 1i,0 is the original data value of the i-th dimension in the input variable, min X 1i,0 is the minimum value of the original data of the i-th dimension in the input variable, max X 1i,0 is the maximum value of the original data of the i-th dimension in the input variable, X 1i is the normalized data value of the i-th dimension in the input variable.
4. The method according to claim 1, wherein, the step S200 includes: S210. Set the neural network hyperparameters of the bidirectional gated recurrent unit model. Among them, select some of the neural network hyperparameters as the hyperparameters to be optimized. And among them, the bidirectional gated recurrent unit model includes an input layer, a hidden layer and an output layer, and the hidden layer includes an attention mechanism layer, a bidirectional layer and a GRU layer; S220. Set the search space of the hyperparameters to be optimized; S230. Combine the hyperparameters in the search space of the hyperparameters; S240. Train the bidirectional gated recurrent unit model, automatically optimize the hyperparameters according to the accuracy of multiple trainings, and finally obtain the optimal hyperparameters and the optimal model.
5. The method according to claim 4, wherein, the hyperparameters include the number of neurons in the input layer, the number of neurons in the hidden layer, the number of neurons in the output layer, the learning rate, the batch size, and the number of iterations.
6. The method according to claim 4, wherein, the step S240 includes: S241. According to the set optimization hyperparameters and the corresponding hyperparameter space, in the first training, use the tree-structured Parzen estimation method to randomly search for combinations of hyperparameters within the hyperparameter space and create a set of error observation values {x (i) , y (i) , i = 1, 2,..., N init}, where x is a combination of hyperparameters to be optimized, and y is the error value trained using the corresponding hyperparameter combination; S242. Set an error quantile value y using the tree-structured Parzen estimation method based on the results and training accuracy of the previous N training sessions of the bidirectional gated recurrent unit model * , and divide the error observation value set into two parts, and calculate its probability value through the following formula: where \(l(x)\) is the probability density function of the set of hyperparameters with error values less than \(y\), and \(g(x)\) is the probability density function of the set of hyperparameters with error values less than \(y\); * and \(g(x)\) is the probability density function of the set of hyperparameters with error values less than \(y\); * ; S243. Calculate the EI value through the following formula: S244. Select the next set of hyperparameter values by maximizing the EI value; S245. Repeat steps S241 to S244 until the bidirectional gated recurrent unit model reaches the set number of training times and then ends.
7. The method according to claim 1, wherein, the step S400 includes the following steps: S410. Use the input variables in the previous training set and test set as the input variables in the current training set and test set, and use the point prediction results of the output variables to construct the output variables of the next training set and test set; S420. Input the current training set and test set into the Gaussian process regression GPR model to obtain a probability distribution function corresponding to the point prediction result of the output variable, and the probability distribution function follows a Gaussian distribution; S430. Based on the mean, standard deviation and preset confidence level of the probability distribution function, determine the prediction interval of each point prediction result under the preset confidence level.
8. The method according to claim 1, wherein, Compare the calculated indicators of point prediction, interval prediction and probability prediction with the BiLSTM GPR model and the Gaussian process regression GPR model to obtain a better model, the indicator of the point prediction is calculated according to the following formula: where y i is the i-th observation value, is the average value of the i-th observation value, is the i-th predicted value, and n is the number of prediction samples; the indicator of the interval prediction is calculated according to the following formula: MC = MWP / CP where y upper,i is the upper limit of the prediction interval of the predicted value of the i-th point, and y lower,i is the lower limit of the prediction interval of the predicted value of the i-th point, and n j is the number of observed values within the prediction interval. the indicator of the probability prediction is calculated according to the following formula: where F(y i ) is the cumulative distribution function of the predicted value, is the unit step function.
9. A computer device includes a memory and a processor, characterized in that, when the processor executes the computer program stored in the memory, the method described in any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Ship motion forecasting method based on long-short-term memory network and Gaussian process regression
CN112684701A
River total nitrogen concentration prediction method based on hybrid neural network
CN113408799A