A sewage water quality prediction method based on a pyramid-shaped three-dimensional model
By constructing a pyramid-type three-dimensional model with multi-model fusion stacking, the problem that existing sewage water quality prediction methods are difficult to reflect real-time changes is solved, and high-accuracy and real-time sewage water quality prediction is achieved, and the intelligence and automation of sewage treatment is supported.
Patent Information
- Application Number
- CN202510475162.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-04-16
AI Technical Summary
The existing sewage water quality prediction methods are difficult to accurately reflect the real-time changes in sewage water quality, and traditional chemical analysis methods have problems such as high cost, environmental pollution and real-time monitoring.
A pyramid-type three-dimensional model based on multi-model fusion stacking is adopted to deform and expand by obtaining wastewater water quality index monitoring data, and a stereo prediction model including RFR, CNN, GBDT, LSTM and BPNN models is constructed to achieve accurate prediction of COD and NH3-N.
It improves the accuracy and real-time nature of sewage water quality prediction, reduces overfitting, enhances the prediction performance of sewage water quality, and supports the intelligence and automation of sewage treatment processes.
Smart Images

Figure CN119990481B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of predictive analysis in the water treatment process, and particularly relates to a sewage water quality prediction method based on a pyramid-shaped three-dimensional model with multi-model fusion and stacking. Background Art
[0002] Importance of sewage treatment and water quality index determination: In modern sewage treatment processes, the Sequencing Batch Reactor (SBR) is widely used due to its strong adaptability and flexible operation. Among them, Chemical Oxygen Demand (COD), Ammonia Nitrogen (NH 3 -N), Total Nitrogen (TN), and Total Phosphorus (TP) are key indicators for measuring the content of pollutants in water. Accurately measuring their values is of irreplaceable significance for optimizing sewage treatment plans and ensuring that sewage discharges meet strict regulatory standards. They are not only directly related to the efficiency of sewage treatment, but also play a key role in the assessment of environmental impacts and the regulation of processes.
[0003] Limitations of traditional COD and NH 3 -N determination methods: However, traditional chemical analysis methods for COD and NH 3 -N have many drawbacks. This method consumes a large amount of chemical reagents during the determination process, which not only increases the detection cost, but may also cause secondary pollution to the environment. At the same time, the time required for water sample digestion and determination is long, and it is impossible to obtain monitoring data in a timely and online manner, resulting in difficulties in adjusting treatment strategies in a timely manner during sewage treatment and being unable to meet the requirements of modern sewage treatment for real-time monitoring and precise control.
[0004] Rise and problems of soft sensing methods: To overcome the deficiencies of traditional chemical determination methods, a soft sensing model based on methods such as machine learning, namely a water quality soft sensor, has emerged. This method predicts the content of COD and NH 3 -N in water quality by establishing a mathematical model, and is expected to achieve rapid and online monitoring. However, due to the complex kinetic characteristics during the sewage treatment process and the complexity and discontinuity of COD and NH 3 -N measurement itself, the existing soft sensing methods still need to improve in terms of prediction accuracy and are difficult to accurately reflect the real-time changes in sewage water quality. Summary of the Invention
[0005] The purpose of the present invention is to solve the problem that existing predictive analysis methods in the water treatment process are difficult to accurately reflect the real-time changes in sewage water quality, and a sewage water quality prediction method based on a pyramid-shaped three-dimensional model with multi-model fusion and stacking is proposed.
[0006] The technical solution of the present invention is as follows: A sewage water quality prediction method based on a pyramid-shaped three-dimensional model with multi-model fusion and stacking includes the following steps:
[0007] S1. Obtain the monitoring data of wastewater quality indicators, and perform deformation and expansion on the monitoring data of wastewater quality indicators to obtain deformed and expanded data;
[0008] S2. Construct a three-dimensional prediction model based on multi-model fusion and stacking;
[0009] S3. Input the wastewater quality monitoring data and the deformed and expanded data into the three-dimensional prediction model, and output the final COD prediction value and NH 3 -N prediction value to complete the prediction of sewage quality.
[0010] Preferably, the monitoring data of the wastewater quality indicators in step S1 includes the conductivity EC, pH value, oxidation-reduction potential ORP, reaction temperature T, dissolved oxygen DO, and turbidity NTU of the wastewater.
[0011] Preferably, the deformed and expanded data includes the conductivity change rate △EC, pH change rate △pH, oxidation-reduction potential change rate △ORP, dissolved oxygen change rate △DO, reaction temperature change rate △T, turbidity change rate △NTU, conductivity cumulative value EC Cum , pH cumulative value pH Cum , oxidation-reduction potential cumulative value ORP Cum , dissolved oxygen cumulative value DO Cum , reaction temperature cumulative value T Cum , turbidity cumulative value NTU Cum , the product of each wastewater quality indicator monitoring data and the quotient of each wastewater quality indicator monitoring data.
[0012] Preferably, the three-dimensional prediction model in step S2 includes a wastewater quality indicator prediction model and a meta-model;
[0013] The wastewater quality indicator prediction model is used to receive the monitoring data of the wastewater quality indicators and output the preliminary COD prediction value and NH 3 -N preliminary prediction value;
[0014] The meta-model is used to receive the preliminary COD prediction value and NH 3 -N preliminary prediction value and output the final COD prediction value and NH 3 -N prediction value.
[0015] Preferably, the wastewater quality index prediction model includes a parallel RFR model, a CNN model, a GBDT model, and an LSTM model; the input ends of the RFR model, the CNN model, the GBDT model, and the LSTM model are all the input ends of the entire wastewater quality index prediction model; the output ends of the RFR model, the CNN model, the GBDT model, and the LSTM model are all connected to the input end of the meta-model;
[0016] The meta-model is a BPNN model.
[0017] Preferably, the RFR model constructs multiple decision trees and averages the prediction results of the multiple decision trees to obtain each preliminary prediction value; the calculation formula for each preliminary prediction value is:
[0018]
[0019] where, represents each preliminary prediction value output by the RFR model, represents the total number of decision trees, represents the monitored data of the wastewater quality index input, represents the th decision tree's prediction result for COD and NH 3 -N.
[0020] Preferably, the GBDT model captures the non-linear characteristics of the monitored data of the wastewater quality index by constructing a series of decision trees, specifically:
[0021] Calculate the negative gradient of the current decision tree, and its calculation formula is:
[0022]
[0023] where, represents the negative gradient of the j th wastewater quality index in the mth iteration, represents the partial derivative, represents the loss function, represents the true label of the j th instance, represents the prediction value of the decision tree in the (m - 1)th iteration for the j th wastewater quality index, represents the j th input parameter of the wastewater quality index, represents the derivative of the loss function with respect to the prediction value of the boosting model;
[0024] Take the negative gradient of the current decision tree as the objective of the loss function, and calculate the next decision tree. Its calculation formula is:
[0025]
[0026] Wherein, represents the predicted value of the decision tree in the m-th iteration, represents the predicted value of the decision tree in the (m - 1)-th stage, represents the learning rate of the th decision tree, represents the output of the th decision tree,
[0027] Repeat the above steps until the preset number of iterations is reached to obtain the non-linear characteristics of the wastewater quality index monitoring data.
[0028] Preferably, the LSTM model captures the long-term dependence relationship between the wastewater quality monitoring data and the preliminary predicted values of COD and NH 3 -N through including an input gate, a forget gate and an output gate;
[0029] The input gate is used to determine how much of the current input information is retained;
[0030] The forget gate is used to determine the information that needs to be forgotten in the previous state;
[0031] The output gate is used to determine the influence of the current cell state on the output;
[0032] The LSTM model specifically includes the following formulas:
[0033]
[0034] Wherein, represents the output parameter of the input gate, represents the output parameter of the forget gate, represents the output parameter of the output gate, represents the cell state of the LSTM model at time, represents the sigmoid activation function, represents the weight of the input gate, represents the hidden state at time, represents the hidden state at time, represents the bias of the input gate, represents the weight of the forget gate, represents the bias of the forget gate, represents an adjustment factor dynamically calculated according to the current environmental conditions, represents the weight of the output gate, represents the bias of the output gate, represents the LSTM model at the cell state at the moment, represents the memory representation corresponding to the input information at the moment, represents the hyperbolic tangent activation function, represents the weight corresponding to the cell state update, represents the bias corresponding to the cell state update.
[0035] Preferably, the output of the BPNN neural network is:
[0036]
[0037] wherein, represents the final COD prediction value and NH 3 -N prediction value output by the BPNN neural network, represents the activation function, represents the weight of the output layer, represents the weight of the hidden layer, represents the fusion feature of the preliminary prediction values obtained by the parallel fusion of the RFR model, CNN model, GBDT model and LSTM model, represents the bias term of the hidden layer, represents the bias term of the output layer.
[0038] Preferably, the step S2 specifically includes the following sub-steps:
[0039] Perform expansion processing on the wastewater quality index monitoring data to obtain deformed expansion data;
[0040] Use the equal difference interpolation method to fill the wastewater quality index monitoring data to obtain complete wastewater quality index monitoring data;
[0041] Construct a variable data set according to the complete wastewater quality index monitoring data and the deformed expansion data;
[0042] Input the variable data set into the wastewater quality index prediction model for training to obtain a trained wastewater quality index prediction model, a preliminary COD prediction value and NH 3 -N preliminary prediction value;
[0043] The preliminary COD prediction value, NH 3-N preliminary predicted value, measured COD value, and measured NH 3 -N measured values are input into the meta-model for training to obtain a three-dimensional prediction model.
[0044] The beneficial effects of the present invention are as follows:
[0045] 1. By performing deformation and expansion processing on the monitoring data of wastewater quality indicators, the obtained deformed and expanded data can more comprehensively reflect the dynamic change characteristics of water quality data and the correlations between data, enhancing the amount of information input into the model.
[0046] 2. The present invention designs a three-dimensional architecture for multi-model fusion, combining the BPNN model, LSTM model, RNN model, RFR model, and GBDT model to construct a pyramid-shaped three-dimensional prediction model with multi-model fusion stacking, which can fully combine the advantages of each model, better capture the temporal characteristics and non-linear relationships in the data, solve the limitations of a single model in complex water quality prediction, and improve prediction accuracy and reduce overfitting in real-time prediction, maximizing the prediction performance for sewage water quality.
[0047] 3. The present invention proposes a dynamic meta-model optimization mechanism, which can improve prediction accuracy and adaptability.
[0048] 4. Through in-depth research on model stacking, it can effectively promote the intelligence and automation of water quality monitoring. Description of the Drawings
[0049] Figure 1 The figure shows a flowchart of a sewage water quality prediction method based on a pyramid-shaped three-dimensional model with multi-model fusion stacking provided in an embodiment of the present invention.
[0050] Figure 2 The figure shows a flowchart block diagram of a sewage water quality prediction method based on a pyramid-shaped three-dimensional model with multi-model fusion stacking provided in an embodiment of the present invention.
[0051] Figure 3 The figure shows a flowchart of independent variables of an iterative prediction meta-model provided in an embodiment of the present invention.
[0052] Figure 4 The figure shows a flowchart of meta-model testing and data processing provided in an embodiment of the present invention.
[0053] Figure 5 The figure shows a comparison chart of the predicted value and the actual test value of ammonia nitrogen NH3-N predicted by using the solution provided by the present invention in an embodiment of the present invention.
[0054] Figure 6 The figure shows a comparison chart of the predicted value and the actual test value of COD predicted by using the solution provided by the present invention in an embodiment of the present invention. Detailed implementation manners
[0055] The exemplary implementation manners of the present invention will now be described in detail with reference to the accompanying drawings. It should be understood that the implementation manners shown and described in the drawings are merely exemplary, intended to illustrate the principles and spirit of the present invention, rather than limiting the scope of the present invention.
[0056] Example:
[0057] As Figure 1 and Figure 2 shown, a sewage water quality prediction method based on a pyramid-shaped three-dimensional model with multi-model fusion stacking includes the following steps:
[0058] S1. Obtain the monitoring data of wastewater water quality indicators, and perform deformation and expansion on the monitoring data of wastewater water quality indicators to obtain deformed and expanded data; the monitoring data of wastewater water quality indicators includes the conductivity EC, pH value, oxidation-reduction potential ORP, reaction temperature T, dissolved oxygen DO, and turbidity NTU of the wastewater; the deformed and expanded data includes the conductivity change rate △EC, pH change rate △pH, oxidation-reduction potential change rate △ORP, dissolved oxygen change rate △DO, reaction temperature change rate △T, turbidity change rate △NTU, conductivity cumulative value EC Cum , pH cumulative value pH Cum , oxidation-reduction potential cumulative value ORP Cum , dissolved oxygen cumulative value DO Cum , reaction temperature cumulative value T Cum , turbidity cumulative value NTU Cum , the product of each wastewater water quality indicator monitoring data and the quotient of each wastewater water quality indicator monitoring data.
[0059] S2. Construct a pyramid-shaped three-dimensional prediction model based on multi-model fusion stacking;
[0060] The three-dimensional prediction model includes a wastewater water quality indicator prediction model and a meta-model;
[0061] The wastewater water quality indicator prediction model is used to receive the monitoring data of the wastewater water quality indicators and output the preliminary predicted values of COD and NH 3 -N preliminary predicted values; it includes a parallel random forest (RFR) model, a convolutional neural network (CNN) model, a gradient boosting tree (GBDT) model, and a long short-term memory network (LSTM) model; the input ends of the RFR model, the CNN model, the GBDT model, and the LSTM model are all the input ends of the entire wastewater water quality indicator prediction model; the output ends of the RFR model, the CNN model, the GBDT model, and the LSTM model are all connected to the input end of the meta-model.
[0062] The RFR model obtains each preliminary prediction value by constructing multiple decision trees and averaging the prediction results of the multiple decision trees; the calculation formula for each preliminary prediction value is:
[0063]
[0064] where represents each preliminary prediction value output by the RFR model, represents the total number of decision trees, represents the monitoring data of the wastewater quality index input, represents the th decision tree's prediction result for COD and NH 3 -N.
[0065] The GBDT model captures the non-linear characteristics of the monitoring data of the wastewater quality index by constructing a series of decision trees, specifically:
[0066] Calculate the negative gradient of the current decision tree, and its calculation formula is:
[0067]
[0068] where represents the negative gradient of the j th wastewater quality index in the mth iteration, represents the partial derivative, represents the loss function, represents the true label of the j th instance, represents the prediction value of the decision tree in the (m - 1)th iteration for the j th wastewater quality index, represents the j th input parameter of the wastewater quality index, represents the derivative of the loss function with respect to the prediction value of the boosting model;
[0069] Take the negative gradient of the current decision tree as the target of the loss function, and calculate to obtain the next decision tree, and its calculation formula is:
[0070]
[0071] where represents the prediction value of the decision tree in the mth iteration, represents the prediction value of the decision tree in the (m - 1)th stage, represents the th learning rate of the decision tree, represents the th output of the decision tree, represents the monitoring data of the wastewater quality index;
[0072] Repeat the above steps until the preset number of iterations is reached to obtain the non-linear characteristics of the wastewater quality index monitoring data.
[0073] The LSTM (Long Short-Term Memory) model is a special type of Recurrent Neural Network (RNN) that can effectively capture long-term dependencies and context information in time series data, aiming to solve the problems of vanishing and exploding gradients faced by traditional RNNs in long sequence learning. LSTM enables the network to effectively capture long-term dependencies and context information by introducing a gating mechanism. Its structure consists of multiple LSTM units, and the storage and update of information are mainly achieved through the following steps.
[0074] Input gate: The input gate determines how much of the current input information should be retained and combines the hidden state from the previous time step. It generates a weight between 0 and 1 through a sigmoid activation function to control the degree of information inflow.
[0075] Forget gate: The forget gate determines which information in the previous state should be forgotten. Similarly, the forget gate uses the sigmoid activation function to generate weights.
[0076] Cell state update: Combining the outputs of the input gate and the forget gate, the LSTM updates the cell state. First, a new candidate cell state is generated and the value range is restricted using the tanh activation function, and then the current cell state is updated.
[0077] Output gate: The output gate determines the influence of the current cell state on the output. After being activated by sigmoid, it combines the cell state activated by tanh to generate the current hidden state.
[0078] Such a design enables the LSTM to effectively learn and remember important information in long sequence data and perform well in various tasks, such as sequence prediction, natural language processing, etc. Through the gating mechanism, the LSTM overcomes the limitations of traditional RNNs and can adapt to complex time series data. The specific LSTM model includes the following formulas:
[0079]
[0080] Where, represents the output parameter of the input gate, represents the output parameter of the forget gate, represents the output parameter of the output gate, represents the LSTM model at the cell state at time step, represents the sigmoid activation function, represents the weight of the input gate, represents The hidden state at a moment, denotes the hidden state at a moment, denotes the monitoring data of wastewater quality indicators input into the LSTM model at a moment, denotes the bias of the input gate, denotes the weight of the forget gate, denotes the bias of the forget gate, denotes the adjustment factor dynamically calculated according to the current environmental conditions, denotes the weight of the output gate, denotes the bias of the output gate, denotes that in the LSTM model at a moment, the cell state, denotes the memory representation corresponding to the input information at a moment, denotes the hyperbolic tangent activation function, denotes the weight corresponding to the cell state update, denotes the bias corresponding to the cell state update.
[0081] The meta-model is used to receive the preliminary COD prediction value and NH 3 -N preliminary prediction value, and output the final COD prediction value and NH 3 -N prediction value; the meta-model is a backpropagation neural network (BPNN) neural network. The BPNN neural network further integrates the features extracted from the wastewater quality indicator prediction model through the connection weight optimization and error backpropagation mechanism for intelligent prediction. Its output is:
[0082]
[0083] Among them, denotes the final COD prediction value and NH 3 -N prediction value output by the BPNN neural network, denotes the activation function, denotes the weight of the output layer, denotes the weight of the hidden layer, denotes the bias term of the hidden layer, denotes the bias term of the output layer, denotes the fusion feature of each preliminary prediction value output by the wastewater quality indicator prediction model, which is fused in parallel through the RFR model, CNN model, GBDT model, and LSTM model, that is, the RFR model, CNN model, GBDT model, and LSTM model are used to predict the wastewater quality indicators respectively to obtain the preliminary COD prediction value and NH 3-N preliminary prediction value, and then use the preliminary COD prediction value and the NH 3 -N preliminary prediction value as the input of the meta-model for final prediction. Compared with the simple processing methods such as weighted averaging or median method for the prediction values of the parallel models, the root mean square error (RMSE) and mean absolute error (MAE) of the prediction values obtained by the method of the present invention are better, and its prediction values are closer to the measured values (true values).
[0084] Step S2 specifically includes the following sub-steps:
[0085] Perform expansion processing on the monitoring data of wastewater quality indicators to obtain deformed expansion data;
[0086] The SBR operating parameters are as follows: the influent time of the biochemical pool is 30 minutes - stirring for 30 minutes - aeration for 30 minutes -... - stirring for 30 minutes - aeration for 30 minutes - drainage for 30 minutes, where "- stirring for 30 minutes - aeration for 30 minutes -" has a total of 10 cycles, and the total operating cycle duration is 660 minutes. In the SBR biochemical reaction pool, an on-line dissolved oxygen (DO) monitor, an on-line oxidation-reduction potential (ORP) monitor, and an on-line conductivity (EC) monitor are installed. During the 600-minute reaction time excluding influent and effluent, the DO value, ORP value, and EC value of the water quality are collected every 1 minute. Every 10 minutes, the NH 3 -N value of the water quality is measured by chemical analysis methods, such as the NH 3 -N value at the 1st minute, 11th minute, and 21st minute, and the NH 3 -N values at other times are obtained by arithmetic interpolation method.
[0087] Run the SBR reactor for 100 complete cycles to obtain the DO value, ORP value, EC value, reaction temperature T, and pH data for 100 cycles. The data for each cycle consists of 600 groups of data, and the data characteristics within each cycle are time series. Each time series contains multiple input features as shown in Table 1, and 2 output features COD and NH 3 -N.
[0088] Table 1 Input Feature Description
[0089]
[0090] Product: EC*pH, EC*DO, EC*ORP, EC*T, EC*NTU, pH*DO, pH*ORP, pH*T, pH*NTU, DO*ORP, DO*T, DO*NTU, ORP*T, ORP*NTU, T*NTU;
[0091] Quotients: EC / pH, EC / DO, EC / ORP, EC / T, EC / NTU, Ph / DO, pH / ORP, pH / T, pH / NTU, DO / ORP, DO / T, DO / NTU, ORP / T, ORP / NTU, T / NTU.
[0092] Use the equal difference interpolation method to fill and process the monitoring data of wastewater quality indicators to obtain complete monitoring data of wastewater quality indicators;
[0093] Take the complete monitoring data of wastewater quality indicators and the deformed and extended data as independent variables, and take COD and NH 3 -N as the dependent variable to construct a variable data set. Each sample represents the water quality state at a specific time point, forming a time series data structure;
[0094] Input the variable data set into the wastewater quality indicator prediction model for training to obtain a trained wastewater quality indicator prediction model, a preliminary COD prediction value, and a preliminary NH 3 -N preliminary prediction value;
[0095] In 100 cycles, randomly select 80 cycles, sort them, use the data sets of the first 79 cycles as the training set, and the remaining 1-cycle data set as the test set. Use two different prediction methods (BPNN model, LSTM model, RNN model, RFR model, GBDT model) to train the 79 training sets respectively to obtain corresponding prediction models. Use the two prediction models to predict the 80th cycle respectively to obtain the prediction results of the 80th cycle, denoted as x 80 , y 80 . Repeat the above steps, use the data of each cycle in the 80 cycles as the test set, and the remaining 79 groups as the training set, and a total of 80 independent prediction results are obtained, denoted as set W, X, W={w 1 , w 2 , ……, w 80}, X={x 1 , x 2 , ……, x 80}. In this embodiment, the LSTM model and the RFR model are selected. Taking the prediction of COD as an example, the process of iteratively predicting the independent variables of the meta-model is as Figure 3 shown.
[0096] Input the preliminary COD prediction value, the preliminary NH3-N prediction value, the measured COD value, and the measured NH 3 -N value into the meta-model for training to obtain a three-dimensional prediction model;
[0097] Using the data of the remaining 20 cycles randomly selected from 100 cycles as the test set of the meta-model, and using the previously selected 80-cycle data as the training set. Through the BPNN model, LSTM model, RNN model, RFR model and GBDT model, the data of 20 cycles are predicted respectively to obtain the preliminary predicted values of COD and NH 3 -N preliminary predicted values, denoted as sets Y and Z, Y = {y 1 , y 2 , ……, y 20}, Z = {z 1 , z 2 , ……, z 20}. The sets Y and Z are used as input features on the meta-model test set, and the meta-model test and data processing flow are as Figure 4 shown.
[0098] S3. Input the wastewater quality monitoring data and the deformed and expanded data into the three-dimensional prediction model, and output the final predicted values of COD and NH 3 -N, complete the prediction of the sewage quality, and the comparison between the prediction results and the test values is as Figure 5 and Figure 6 shown.
[0099] Through the above steps, the present invention improves the prediction accuracy of time series data by comprehensively utilizing the advantages of multiple models and training and predicting in a stacked manner. The core of this method is to effectively integrate the prediction results of different models. Using the BPNN neural network as the meta-model, it fully excavates the prediction capabilities of each model, and finally obtains more accurate prediction results. This method has a wide application prospect in time series analysis and prediction tasks and is applicable to the data prediction needs of multiple fields.
[0100] The water quality intelligent prediction method of the present invention stacks deep learning and ensemble learning models, fully combines the advantages of the BPNN model, LSTM model, RNN model, RFR model, GBDT model and BPNN model, and maximally improves the prediction performance of the model. In this embodiment, the three-dimensional prediction model proposed by the present invention is compared with the prediction effects of a single BPNN model, LSTM model, RNN model and RFR model. The prediction performances of different models are shown in Table 2. It can be seen that in complex sewage water quality data, the present invention can better capture the time series characteristics and non-linear relationships in the data. In real-time prediction, it can improve the prediction accuracy, reduce overfitting, and provide a more reliable information basis for proposing feasible improvement schemes for wastewater treatment.
[0101] Table 2 Comparison of the prediction effects of the three-dimensional prediction model and other models on COD and ammonia nitrogen NH 3 -N on the test set
[0102]
[0103] The present invention can not only effectively promote the intellectualization and automation of water quality monitoring, but also provide a new idea and method for predictive analysis in other fields, having broad application potential and market value. Through in-depth research on model stacking, in the future, it can be replaced or extended according to different application requirements to achieve more excellent predictive performance.
[0104] Those of ordinary skill in the art will realize that the embodiments described herein are for helping readers understand the principles of the present invention, and it should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations that do not depart from the essence of the present invention based on these technical revelations disclosed by the present invention, and these deformations and combinations are still within the protection scope of the present invention.
Claims
1. A method for predicting sewage quality based on a pyramid-shaped three-dimensional model, characterized in that: The following steps are involved: S1. Obtain wastewater quality index monitoring data, and deform and expand the wastewater quality index monitoring data to obtain deformed and expanded data; S2. Construct a stereo prediction model based on multi-model fusion stacking; S3. Input the wastewater quality monitoring data and deformation expansion data into the three-dimensional prediction model, output the final COD prediction value and NH3-N prediction value, and complete the wastewater quality prediction; The three-dimensional prediction model includes a wastewater quality index prediction model and a meta-model; The wastewater quality index prediction model is used to receive the wastewater quality index monitoring data and output a preliminary prediction value of COD and a preliminary prediction value of NH3-N; The meta-model is used to receive the COD preliminary prediction value and the NH3-N preliminary prediction value, and output the final COD prediction value and the NH3-N prediction value; The wastewater quality index prediction model is a parallel RFR model, a CNN model, a GBDT model and a LSTM model; the input end of the RFR model, the input end of the CNN model, the input end of the GBDT model and the input end of the LSTM model are all input ends of the entire wastewater quality index prediction model; the output end of the RFR model, the output end of the CNN model, the output end of the GBDT model and the output end of the LSTM model are all connected to the input end of the meta-model; The meta-model is a BPNN neural network.
2. The method for predicting sewage quality based on a pyramid-shaped stereoscopic model according to claim 1, characterized in that: The wastewater quality index monitoring data in step S1 includes the wastewater's electrical conductivity EC, pH value, oxidation-reduction potential ORP, reaction temperature T, dissolved oxygen DO, and turbidity NTU.
3. The sewage quality prediction method based on the pyramid-type three-dimensional model according to claim 2, characterized in that: The deformation expansion data includes conductivity change rate △EC, pH change rate △pH, redox potential change rate △ORP, dissolved oxygen change rate △DO, reaction temperature change rate △T, turbidity change rate △NTU, conductivity accumulation value EC Cum , pH cumulative value pH Cum , Oxidation-reduction potential cumulative value ORP Cum , dissolved oxygen cumulative value DO Cum , reaction temperature cumulative value T Cum , Turbidity cumulative value NTU Cum , the product of the monitoring data of each wastewater quality index and the quotient of the monitoring data of each wastewater quality index.
4. The method for predicting sewage quality based on a pyramid-shaped stereoscopic model according to claim 1, characterized in that: The RFR model constructs multiple decision trees and averages the prediction results of the multiple decision trees to obtain each preliminary prediction value; the calculation formula of each preliminary prediction value is: in, represents the preliminary prediction values output by the RFR model, represents the total number of decision trees, Indicates the input wastewater quality index monitoring data, Indicates The prediction results of COD and NH3-N by the decision trees.
5. The method for predicting sewage quality based on a pyramid-shaped stereoscopic model according to claim 1, characterized in that: The GBDT model captures the nonlinear characteristics of wastewater quality index monitoring data by constructing a series of decision trees, specifically: Calculate the negative gradient of the current decision tree, and the calculation formula is: in, Indicates the mth iteration j The negative gradient of the wastewater quality index, represents the partial derivative, represents the loss function, Indicates j The true labels of the instances, Represents the decision tree of the m-1th iteration for the j The predicted values of wastewater quality indicators, Indicates j The input parameters of the wastewater quality indicators are: represents the derivative of the loss function with respect to the predicted value of the boost model; The negative gradient of the current decision tree is used as the target of the loss function to calculate the next decision tree. The calculation formula is: in, represents the predicted value of the decision tree at the mth iteration, represents the predicted value of the decision tree at the m-1th stage, Indicates The learning rate of a decision tree, Indicates The output of a decision tree, Indicates the monitoring data of wastewater quality indicators; Repeat the above steps until the preset number of iterations is reached to obtain the nonlinear characteristics of the wastewater quality index monitoring data.
6. The method for predicting sewage quality based on a pyramid-shaped stereoscopic model according to claim 1, characterized in that: The LSTM model captures the long-term dependency between wastewater quality monitoring data and the preliminary predicted values of COD and NH3-N by including an input gate, a forget gate, and an output gate; The input gate is used to determine how much current input information is retained; The forget gate is used to determine the information in the previous state that needs to be forgotten; The output gate is used to determine the impact of the current unit state on the output; The LSTM model specifically includes the following formula: in, represents the output parameter of the input gate, represents the output parameter of the forget gate, represents the output parameter of the output gate, Indicates that the LSTM model is The unit state at the moment, represents the sigmoid activation function, represents the weight of the input gate, express The hidden state of the moment, express The hidden state of the moment, express The wastewater quality index monitoring data is input into the LSTM model at all times. represents the bias of the input gate, represents the weight of the forget gate, represents the bias of the forget gate, Represents the adjustment factor dynamically calculated according to the current environmental conditions, represents the weight of the output gate, represents the bias of the output gate, Indicates that the LSTM model is The unit state at the moment, express The memory representation corresponding to the input information at each moment, represents the hyperbolic tangent activation function, represents the weight corresponding to the unit state update, Indicates the bias corresponding to the cell state update.
7. The method for predicting sewage quality based on a pyramid-shaped stereoscopic model according to claim 1, characterized in that: The output of the BPNN neural network is: in, It represents the final COD prediction value and NH3-N prediction value output by the BPNN neural network. represents the activation function, represents the weight of the output layer, represents the weight of the hidden layer, It represents the fusion features of the preliminary prediction values obtained by parallel fusion of the RFR model, CNN model, GBDT model and LSTM model. represents the bias term of the hidden layer, Represents the bias term of the output layer.
8. The method for predicting sewage quality based on a pyramid-shaped stereoscopic model according to claim 1, characterized in that: The step S2 specifically includes the following sub-steps: Perform expansion processing on the wastewater quality index monitoring data to obtain deformed expansion data; The wastewater quality index monitoring data is filled in using the arithmetic difference interpolation method to obtain complete wastewater quality index monitoring data; Construct variable data sets based on complete wastewater quality index monitoring data and deformation and expansion data; The variable data set is input into the wastewater quality index prediction model for training, and the trained wastewater quality index prediction model, COD preliminary prediction value and NH3-N preliminary prediction value are obtained; The preliminary predicted values of COD, preliminary predicted values of NH3-N, measured values of COD and measured values of NH3-N were input into the meta-model for training to obtain a three-dimensional prediction model.
Citation Information
Patent Citations
Offshore doubly fed wind power generator fault judging method based on GRA-LSTM-stacking model
CN111237134A
Water quality prediction method and device, electronic equipment and storage medium
CN113159456A