Method for predicting effluent total phosphorus in sewage treatment process based on hierarchical decomposition-integrated neural network

Through hierarchical decomposition and integrated neural network model, the total effluent phosphorus in sewage treatment is predicted, which solves the problem of low prediction accuracy of total effluent phosphorus in sewage treatment, and realizes accurate prediction of multi-cycle characteristics and precise control of sewage treatment process.

CN120108564APending Publication Date: 2025-06-06BEIJING UNIV OF TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510037495.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prediction accuracy of total phosphorus in effluent during sewage treatment is low, and due to the influence of multi-periodic and nonlinear characteristics, it is difficult to achieve accurate prediction.

Method used

The hierarchical decomposition-integrated neural network is used to decompose the time series into trend components, seasonal components and residual components through STL and CEEMDAN decomposition methods, and predict using the EA-GRU model to integrate the prediction results of all components to improve the prediction accuracy.

Benefits of technology

The prediction accuracy of total effluent phosphorus in the sewage treatment process is improved, and accurate prediction of total effluent phosphorus with multi-cycle characteristics can be achieved, assisting the sewage treatment plant to achieve accurate control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108564A_ABST
    Figure CN120108564A_ABST
Patent Text Reader

Abstract

The invention discloses a sewage treatment process effluent total phosphorus prediction method based on a hierarchical decomposition-integrated neural network, relates to the field of artificial intelligence, and solves the problems that the effluent total phosphorus concentration trend in the sewage treatment process is difficult to master, the prediction cost is high and the like. Aiming at the problems that the sewage treatment process is influenced by various environmental factors, the effluent total phosphorus embodies various characteristics such as multi-periodicity and nonlinearity, and the prediction precision is influenced, a hierarchical decomposition-integrated prediction model framework is provided, a time sequence is decomposed into a multi-periodicity component, a trend component and a residual error, and the prediction precision is influenced by the multi-periodicity component, the trend component and the residual error. Therefore, each component is accurately modeled, and the prediction performance of the prediction model is improved. Meanwhile, aiming at the problem that a common attention mechanism only pays attention to the internal importance of input variables, an improved attention mechanism is provided, the training efficiency is improved by removing redundant input components, the prediction precision of the model is improved by weighting important components, and the management operation of the sewage treatment plant is promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the field of artificial intelligence and is applied to the field of sewage treatment, in particular to a method for predicting total phosphorus in effluent from a sewage treatment process based on a hierarchical decomposition-integration neural network. Background Art

[0002] Total phosphorus is one of the main factors of eutrophication of water bodies and will have a negative impact on aquatic ecosystems. Therefore, accurate analysis of total phosphorus in the sewage treatment process is of great significance for controlling eutrophication of water bodies and protecting the ecological environment. The total phosphorus prediction results can provide a scientific basis for sewage treatment and management and guide the implementation of emission reduction measures. Since the sewage treatment process is affected by a variety of environmental factors, the total phosphorus in the effluent exhibits multiple characteristics such as multi-periodicity and nonlinearity, which affects the accuracy of the prediction of the total phosphorus in the effluent. Therefore, there is an urgent need for a method that can accurately predict the total phosphorus in the effluent based on the characteristics of the total phosphorus in the effluent during sewage treatment. Summary of the invention

[0003] 1. Technical problems that the present invention needs and can solve:

[0004] The present invention proposes a method for predicting total phosphorus in effluent from a sewage treatment process based on a hierarchical decomposition-integrated neural network. The hierarchical decomposition-integrated neural network model proposed in the present invention adopts a hierarchical decomposition method based on STL and CEEMDNAN and an ensemble attention-gate recurrent unit (EA-GRU) prediction model to specifically analyze the trend, multi-periodicity and other temporal characteristics of total phosphorus in effluent, thereby improving the prediction accuracy.

[0005] 2. Specific technical solutions of the present invention:

[0006] The main process of this method is as follows Figure 1 As shown, including:

[0007] S10, preprocessing the total phosphorus data and dividing them into a training set and a data set;

[0008] S20, based on the hierarchical decomposition method, decompose the temporal characteristics of the time series;

[0009] S30, constructing a prediction model corresponding to the trend component and residual component obtained by the first-level decomposition and the multi-period component obtained by the second-level decomposition;

[0010] S40, training the prediction model and making predictions, integrating the prediction results of all components, and verifying the accuracy of the model;

[0011] S50, inputting the real-time sewage data into the hierarchical decomposition-integrated neural network model to obtain the prediction result of the effluent total phosphorus concentration;

[0012] The sewage data include flow rate, pH value, conductivity, chemical oxygen demand (COD), water temperature, salinity, residual chlorine, total phosphorus, redox potential, dissolved oxygen, and total dissolved solids content;

[0013] The step S10 comprises the following steps:

[0014] S101, deleting abnormal data in the historical total phosphorus data using the 3σ criterion:

[0015] If the deviation v between the historical total phosphorus data and the average value of the historical total phosphorus data satisfies v>3σ, where σ is the standard deviation of the historical total phosphorus data, the historical total phosphorus data is considered to be abnormal data and the abnormal data is deleted.

[0016] S102, normalizing the cleaned historical total phosphorus data by using a Z-Score standardization method:

[0017] The formula is: where y Z is the normalized value of the historical total phosphorus data y, is the average value of historical total phosphorus data;

[0018] S103. Preferably, 70% of the data is used as a training set, and 30% of the data is used as a test set;

[0019] The method for predicting total phosphorus in effluent from sewage treatment process based on hierarchical decomposition-integrated neural network proposed in the present invention is as shown in the attached figure. Figure 2 As shown;

[0020] In order to effectively separate different time series features in a time series, a time series feature decomposition method based on hierarchical decomposition is constructed, wherein the first layer is decomposed in the time domain and the second layer is decomposed in the time-frequency domain; the step S20 includes the following steps:

[0021] S201, time domain decomposition based on STL decomposition, decomposing the time series into trend component, seasonal component and residual component in the time domain;

[0022] The STL decomposition algorithm takes the total phosphorus data y as input Z Decompose into trend component y_t, seasonal component y_s and residual component y_r, distinguish components with different characteristics in the input data, and analyze the three components separately to achieve detailed analysis of sewage time series data;

[0023] The specific steps of the STL data decomposition are:

[0024] STL executes the inner loop in each outer loop. The steps of the inner loop are:

[0025] S2011, detrending: in, It is nth d The detrended sequence of the inner cycle, It is nth d -1 inner cycle of trend sequence;

[0026] S2012, Smooth Cyclic Subsequence: Based on the nth o -1 external loop right Perform LOESS smoothing to obtain a temporary seasonal series in Ψ is the double square weighted formula;

[0027] S2013, low pass filter: Perform 3 sliding averages and then perform LOESS smoothing to obtain the residual trend

[0028] S2014, Decomposition of smooth cyclic subsequences: in It is nth d seasonal component of the sub-inner cycle;

[0029] S2015, Deseasonalization: in It is nth d Deseasonalized series of sub-inner cycles;

[0030] S2016, Smooth trend item: Based on right Perform LOESS smoothing to obtain the trend component

[0031] After all inner loops of an outer loop are completed, subtract the trend and seasonal components from the input data to get the nth o The residual term of the outer loop

[0032] After all the outer loops of STL are completed, input data y Z =y_t+y_s+y_r;

[0033] S202, time-frequency domain decomposition based on CEEMDAN, decomposing multi-periodic features in seasonal components;

[0034] The CEEMDAN algorithm is selected to decompose the seasonal component to obtain the multi-periodic features in the seasonal component y_s. The specific decomposition process of CEEMDAN is as follows:

[0035] S2021. Calculate the first IMF: Add M pairs of positive and negative Gaussian white noise to y_s to obtain y_s w , M is the number of white noise groups. According to experience, M is 250;

[0036] To y_s w Perform the first EMD decomposition, where imf j (1) is the jth modal component of the first decomposition; calculate the average value of the modal component to obtain the first modal component imf (1) :

[0037] Where J is the maximum number of decomposable modal components, which is determined by the decomposition process itself;

[0038] S2022. Calculate the first residual r (1) :r (1) =y_s-imf (1) ;

[0039] S2023. Calculate the kth IMF:

[0040] Let EMD (k) (·) is the k-th EMD decomposition; (k) Add M groups of paired positive and negative Gaussian white noise to get right Perform EMD decomposition to obtain the k+1th IMF:

[0041] Calculate the k+1th residual r (k+1) :r (k+1) =r (k) -imf (k+1) ;

[0042] If the stopping condition is met, that is, r (k+1) If it is a monotonic signal, the iteration stops and the CEEMDAN algorithm decomposition ends;

[0043] At this point, the original signal is the sum of k IMFs and the residual, completing the decomposition: Among them, imf (k) is the IMF obtained from the k-th decomposition, and r is the final residual obtained; since r represents noise that cannot be further decomposed, no further prediction is performed;

[0044] After the CEEMDAN decomposition, the seasonal component

[0045] For the trend component and residual component obtained by the first-layer decomposition and the multi-period component obtained by the second-layer decomposition, a multi-layer integrated deep learning model is constructed to make predictions respectively and integrate them to generate the final prediction results;

[0046] The step S30 comprises the following steps:

[0047] S301. Construct ARIMA model to predict trend component:

[0048] The ARIMA model is used to predict the stationary trend component obtained from the first level decomposition:

[0049]

[0050] Among them, h is the prediction step size, is the autoregressive coefficient, θ n is the moving average coefficient, p and q are the autoregressive order and the moving average order, ε t is white noise, c is a constant, is the predicted value of the trend component;

[0051] Preferably, according to the data characteristics, p and q are specified as 1, and the constant term is specified as 0; and θ n It is estimated by the model through the fitting process;

[0052] S302, construct EA-GRU model to predict periodic components:

[0053] The EA-GRU model is used to predict the IMF representing multi-period features obtained by the second-level decomposition. Figure 3 As shown;

[0054] First calculate the energy of each IMF:

[0055]

[0056] Among them, IMF t (k) It is IMF (k) The tth data point, N is the imf (k) Length;

[0057] In order to remove redundant IMFs, the energy ratio of each IMF is calculated:

[0058]

[0059] Then sort the IMFs in descending order, calculate the sum of the energy ratios starting from the first IMF, and get the cumulative energy ratio. When the cumulative energy ratio of the Lth IMF reaches more than 95%, the remaining IMFs will be regarded as redundant IMFs and eliminated;

[0060] For each filtered IMF, an improved attention-GRU module is established for prediction:

[0061] For the lth attention-GRU module, the GRU includes two gate structures, the update gate and the reset gate:

[0062] Update Gate t (l) Determine the discarding and updating of the state information at the previous moment:

[0063] Among them, W r (l) is the update gate weight matrix, is the input value of the lth GRU, sig is the sigmoid activation function;

[0064] Reset Gate Controls how much information about the previous state is written:

[0065] in, is the reset gate weight matrix;

[0066] Current hidden state That is, the output of the lth attention-GRU module is calculated as follows:

[0067]

[0068] The candidate hidden state in yes The weight matrix of

[0069] By further considering the importance of each IMF, the attention-GRU module introduces an improved attention mechanism to update the input of GRU. The improved attention weight calculation formula is:

[0070] The calculation formula is:

[0071] in, and is the learnable weight matrix of the Lth attention-GRU module, is the energy ratio of the Lth IMF, indicating the importance of IMF; therefore, the input of the Lth GRU Updated to:

[0072] After all l attention-GRU modules have completed their predictions, the second-layer decomposition and integration module adds up the prediction results of all multi-period components to obtain the predicted value of the seasonal component y_s

[0073] S303, construct the GRU model to predict the residual component:

[0074] Construct a GRU model to predict the residual component y_r because it has strong nonlinearity;

[0075] GRU consists of two gate structures, the update gate and the reset gate:

[0076] Update Gate t Determines the discarding and updating of the state information at the previous moment: r t =sig(W r [h t-1 ,y_r t ]);

[0077] Among them, W r is the update gate weight matrix, y_r t is the input value of GRU;

[0078] Reset gate z t Controls how much information of the previous state is written: z t =sig(W z [h t-1 ,y_r t ]);

[0079] Among them, W z is the reset gate weight matrix;

[0080] Current hidden state h t , which is the output of the GRU model The calculation is as follows:

[0081]

[0082] The candidate hidden state Where W h yes The weight matrix of

[0083] S304, the first-layer decomposition integration module adds the prediction results of all multi-period components, trend components and residual components to obtain the final prediction value:

[0084] 3. Compared with the prior art, the present invention has the following obvious advantages and beneficial effects:

[0085] 1. Aiming at the problem that sewage data exhibits multiple characteristics such as multi-periodicity and nonlinearity, which affect the prediction accuracy, the present invention proposes a hierarchical decomposition-integrated prediction model framework, which decomposes the time series into multi-periodic components, trend components and residuals, thereby accurately modeling each component and improving the prediction performance of the prediction model.

[0086] 2. Aiming at the problem that ordinary attention mechanism only focuses on the internal importance of input variables, the present invention proposes an improved attention mechanism, which improves the training efficiency by removing redundant input components and improves the prediction accuracy of the model by weighting important components.

[0087] 3. The method for predicting total phosphorus in effluent from sewage treatment process based on hierarchical decomposition-integrated neural network described in the present invention can accurately predict the total phosphorus in effluent with multi-period characteristics by using the hierarchical decomposition-integrated neural network model, realize real-time measurement of total phosphorus in effluent from sewage treatment process, and assist sewage treatment plants in achieving precise control of sewage treatment process. BRIEF DESCRIPTION OF THE DRAWINGS

[0088] The present invention is further described below in conjunction with the accompanying drawings and embodiments.

[0089] Figure 1 A schematic diagram of the process flow of the method for predicting total phosphorus in effluent from a sewage treatment process based on a hierarchical decomposition-integrated neural network provided by the present invention;

[0090] Figure 2 A schematic diagram of the structure of the hierarchical decomposition-integration neural network prediction model provided by the present invention;

[0091] Figure 3 A schematic diagram of the structure of the EA-GRU prediction model provided by the present invention;

[0092] Figure 4 This is a diagram showing the prediction effect of the total phosphorus concentration in effluent according to the present invention. DETAILED DESCRIPTION

[0093] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings.

[0094] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present invention should have the common meanings understood by persons having ordinary skills in the field to which the present invention belongs.

[0095] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.

[0096] Embodiment 1:

[0097] Embodiment 1 of the present invention provides a method for predicting total phosphorus in effluent from a sewage treatment process based on a hierarchical decomposition-integrated neural network, referring to Figure 1 ; Figure 1 This is a schematic flow chart of a method for predicting total phosphorus in effluent from a sewage treatment process based on a hierarchical decomposition-integrated neural network provided in Example 1.

[0098] In the first embodiment, the method for predicting total phosphorus in effluent from a sewage treatment process based on a hierarchical decomposition-integrated neural network comprises the following steps:

[0099] S10, preprocessing the sewage time series data and dividing it into a training set and a data set;

[0100] S20, based on the hierarchical decomposition method, decompose the temporal characteristics of the time series;

[0101] S30, constructing a prediction model corresponding to the trend component and residual component obtained by the first-level decomposition and the multi-period component obtained by the second-level decomposition;

[0102] S40, training the prediction model and making predictions, integrating the prediction results of all components, and verifying the accuracy of the model;

[0103] S50, inputting the real-time sewage data into the hierarchical decomposition-integrated neural network model to obtain the prediction result of the effluent total phosphorus concentration;

[0104] The sewage data include flow rate, pH value, conductivity, chemical oxygen demand (COD), water temperature, salinity, residual chlorine, total phosphorus, redox potential, dissolved oxygen, and total dissolved solids content;

[0105] The step S10 comprises the following steps:

[0106] S101, deleting abnormal data in the historical sewage data using the 3σ criterion:

[0107] If the deviation v between the historical total phosphorus data and the average value of the historical total phosphorus data satisfies v>3σ, where σ is the standard deviation of the historical total phosphorus data, the historical total phosphorus data is considered to be abnormal data and the abnormal data is deleted.

[0108] S102, normalizing the cleaned historical total phosphorus data by using a Z-Score standardization method:

[0109] The formula is: where y Z is the normalized value of the historical total phosphorus data y, is the average value of historical total phosphorus data;

[0110] S103. Preferably, 70% of the data is used as a training set, and 30% of the data is used as a test set;

[0111] The method for predicting total phosphorus in effluent from sewage treatment process based on hierarchical decomposition-integrated neural network proposed in the present invention is as shown in the attached figure. Figure 2 As shown;

[0112] In order to effectively separate different time series features in a time series, a time series feature decomposition method based on hierarchical decomposition is constructed, wherein the first layer is decomposed in the time domain and the second layer is decomposed in the time-frequency domain. Step S20 includes the following steps:

[0113] S201, time domain decomposition based on STL decomposition, decomposing the total phosphorus time series into trend component, seasonal component and residual component in the time domain;

[0114] The STL decomposition algorithm transforms the total phosphorus data y Z Decompose into trend component y_t, seasonal component y_s and residual component y_r, distinguish components with different characteristics in the input data, and analyze the three components separately to achieve detailed analysis of sewage time series data;

[0115] The specific steps of the STL data decomposition are:

[0116] STL executes the inner loop in each outer loop. The steps of the inner loop are:

[0117] S2011, detrending: in, It is nth d The detrended sequence of the inner cycle, It is nth d -1 inner cycle of trend sequence;

[0118] S2012, Smooth Cyclic Subsequence: Based on the nth o -1 external loop right Perform LOESS smoothing to obtain a temporary seasonal series in Ψ is the double square weighted formula;

[0119] S2013, low pass filter: Perform 3 sliding averages and then perform LOESS smoothing to obtain the residual trend

[0120] S2014, Decomposition of smooth cyclic subsequences: in It is nth d seasonal component of the sub-inner cycle;

[0121] S2015, Deseasonalization: in It is nth d Deseasonalized series of sub-inner cycles;

[0122] S2016, Smooth trend item: Based on right Perform LOESS smoothing to obtain the trend component

[0123] After all inner loops of an outer loop are completed, subtract the trend and seasonal components from the input data to get the nth o The residual term of the outer loop

[0124] After all the outer loops of STL are completed, input data y Z =y_t+y_s+y_r;

[0125] S202, time-frequency domain decomposition based on CEEMDAN, decomposing multi-periodic features in seasonal components;

[0126] The CEEMDAN algorithm is selected to decompose the seasonal component to obtain the multi-periodic features in the seasonal component y_s. The specific decomposition process of CEEMDAN is as follows:

[0127] S2021. Calculate the first IMF: Add M pairs of positive and negative Gaussian white noise to y_s to obtain y_s w , M is the number of white noise groups. According to experience, M is 250;

[0128] To y_s w Perform the first EMD decomposition, where imf j (1) is the jth modal component of the first decomposition; calculate the average value of the modal component to obtain the first modal component imf (1) :

[0129] Where J is the maximum number of decomposable modal components, which is determined by the decomposition process itself;

[0130] S2022. Calculate the first residual r (1) :r (1) =y_d-imf (1) ;

[0131] S2023. Calculate the kth IMF:

[0132] Let EMD (k) (·) is the k-th EMD decomposition; (k) Add M groups of paired positive and negative Gaussian white noise to get right Perform EMD decomposition to obtain the k+1th IMF:

[0133] Calculate the k+1th residual r (k+1) :r (k+1) =r (k) -imf (k+1) ;

[0134] If the stopping condition is met, that is, r (k+1) If it is a monotonic signal, the iteration stops and the CEEMDAN algorithm decomposition ends;

[0135] At this point, the original signal is the sum of k IMFs and the residual, completing the decomposition: Among them, imf (k) is the IMF obtained from the k-th decomposition, and r is the final residual obtained; since r represents noise that cannot be further decomposed, no further prediction is performed;

[0136] After the CEEMDAN decomposition, the seasonal component

[0137] For the trend component and residual component obtained by the first-layer decomposition and the multi-period component obtained by the second-layer decomposition, a multi-layer integrated deep learning model is constructed to make predictions respectively and integrate them to generate the final prediction results;

[0138] The step S30 comprises the following steps:

[0139] S301. Construct ARIMA model to predict trend component:

[0140] The ARIMA model is used to predict the stationary trend component obtained from the first level decomposition:

[0141]

[0142] Among them, h is the prediction step size, is the autoregressive coefficient, θ n is the moving average coefficient, p and q are the autoregressive order and the moving average order, ε t is white noise, c is a constant, is the predicted value of the trend component;

[0143] Preferably, according to the data characteristics, p and q are specified as 1, and the constant term is specified as 0; and θ nIt is estimated by the model through the fitting process;

[0144] S302, construct EA-GRU model to predict periodic components:

[0145] The EA-GRU model is used to predict the IMF representing multi-period features obtained by the second-level decomposition. Figure 3 As shown;

[0146] First calculate the energy of each IMF:

[0147]

[0148] Among them, IMF t (k) It is IMF (k) The tth data point, N is the imf (k) Length;

[0149] In order to remove redundant IMFs, the energy ratio of each IMF is calculated:

[0150]

[0151] Then sort the IMFs in descending order, calculate the sum of the energy ratios starting from the first IMF, and get the cumulative energy ratio. When the cumulative energy ratio of the Lth IMF reaches more than 95%, the remaining IMFs will be regarded as redundant IMFs and eliminated;

[0152] For each filtered IMF, an improved attention-GRU module is established for prediction:

[0153] For the lth attention-GRU module, the GRU includes two gate structures, the update gate and the reset gate:

[0154] Update Gate t (l) Determine the discarding and updating of the state information at the previous moment:

[0155] Among them, W r (l) is the update gate weight matrix, is the input value of the lth GRU, sig is the sigmoid activation function;

[0156] Reset Gate Controls how much information about the previous state is written:

[0157] in, is the reset gate weight matrix;

[0158] Current hidden state That is, the output of the lth attention-GRU module is calculated as follows:

[0159]

[0160] The candidate hidden state in yes The weight matrix of

[0161] By further considering the importance of each IMF, the attention-GRU module introduces an improved attention mechanism to update the input of GRU. The improved attention weight calculation formula is:

[0162] The calculation formula is:

[0163] in, and is the learnable weight matrix of the lth attention-GRU module, is the energy ratio of the lth IMF, indicating the importance of IMF; therefore, the input of the lth GRU Updated to:

[0164] Update Gate t (l) Determine the discarding and updating of the state information at the previous moment:

[0165] Among them, W r (l) is the update gate weight matrix, is the input value of the lth GRU, sig is the sigmoid activation function;

[0166] Reset Gate Controls how much information about the previous state is written:

[0167] in, is the reset gate weight matrix;

[0168] Current hidden state That is, the output of the lth attention-GRU module is calculated as follows:

[0169]

[0170] The candidate hidden state in yes The weight matrix of

[0171] By further considering the importance of each IMF, the attention-GRU module introduces an improved attention mechanism to update the input of GRU. The improved attention weight calculation formula is:

[0172] The calculation formula is:

[0173] in, and is the learnable weight matrix of the lth attention-GRU module, is the energy ratio of the lth IMF, indicating the importance of IMF; therefore, the input of the lth GRU Updated to:

[0174] After all l attention-GRU modules have completed their predictions, the second-layer decomposition and integration module adds up the prediction results of all multi-period components to obtain the predicted value of the seasonal component y_s

[0175] S303, construct the GRU model to predict the residual component:

[0176] Construct a GRU model to predict the residual component y_r because it has strong nonlinearity;

[0177] GRU consists of two gate structures, the update gate and the reset gate:

[0178] Update Gate t Determines the discarding and updating of the state information at the previous moment: r t =sig(W r [h r-1 ,y_r t ]);

[0179] Among them, W r is the update gate weight matrix, y_r t is the input value of GRU;

[0180] Reset gate z t Controls how much information of the previous state is written: z t =sig(w z [H T-1 ,Y_r t ]);

[0181] Among them, W z is the reset gate weight matrix;

[0182] Current hidden state h t , which is the output of the GRU model The calculation is as follows:

[0183]

[0184] The candidate hidden state Where W h yes The weight matrix of

[0185] S304, the first-layer decomposition integration module adds the prediction results of all multi-period components, trend components and residual components to obtain the final prediction value:

[0186] S40, train the integrated deep learning model and verify the accuracy of the model;

[0187] The hierarchical decomposition-integration neural network model training process provided in the embodiment includes the following steps:

[0188] S401. Use the STL decomposition algorithm for input variables:

[0189] The STL decomposition algorithm is used to decompose the effluent ammonia nitrogen into trend component, seasonal component and residual component;

[0190] S402, determine the auxiliary variables for predicting effluent ammonia nitrogen:

[0191] The Pearson correlation coefficient represents the covariance of two variables divided by the product of their standard deviations, which can achieve translation invariance. It measures the linear correlation between two variables and gives a value between -1 and 1. The specific formula for calculating the Pearson correlation coefficient ρ(I,y) between each auxiliary variable I and the effluent ammonia nitrogen parameter y is as follows:

[0192] Among them I l is the lth data of the auxiliary variable, y l is the lth data of the effluent ammonia nitrogen parameter, where l = (1, 2, ..., N``), N`` is the number of measured parameters, is the mean value of the auxiliary variable, is the average value of effluent ammonia nitrogen parameters;

[0193] The input auxiliary variables of effluent ammonia nitrogen were selected by comparing the Pearson correlation coefficient of each auxiliary variable, and five input auxiliary variables were calculated: pH value, conductivity, water temperature, COD, and residual chlorine;

[0194] S402, importing the training set into the model for training;

[0195] Step 1, train the ARIMA model using the trend component;

[0196] Step 2: Use CEEMDAN to decompose the seasonal components, select the IMFs based on the energy of each decomposed IMF, and train the EA-GRU model;

[0197] Step 3, use the residual component to train the GRU model;

[0198] Repeat steps 1-3 until all batches of the training dataset have participated in model training and the specified number of iterations has been reached;

[0199] Preferably, the specified number of iterations is determined to be 250 based on multiple experiments;

[0200] S403, importing the test set data into the trained model to verify the model training effect;

[0201] At the end of each round of cyclic training, the test set is imported into the integrated deep learning model to evaluate the model training effect, and the model prediction result is output;

[0202] Input y_t into ARIMA to get the forecast result

[0203] Decompose y_s by CEEMDAN and then input it into EA-GRU to get the prediction result

[0204] Input y_r and five input variables into the GRU neural network to get the prediction results

[0205] The outputs of all models are integrated and denormalized to obtain the final prediction result:

[0206] Compare the prediction results output by the verified integrated deep learning model with the test set;

[0207] When the prediction error does not decrease for 5 consecutive times, the training is stopped and the optimal hierarchical decomposition-integrated neural network model is obtained; otherwise, iterative training is performed again to update the network weights and bias of the hierarchical decomposition-integrated neural network model until the prediction error does not decrease for 5 consecutive times and the optimal hierarchical decomposition-integrated neural network model is obtained;

[0208] S50, inputting the real-time sewage data into the hierarchical decomposition-integration neural network to obtain the prediction results of the sewage water quality index.

[0209] Real-time sewage data is obtained, and data is deleted and normalized. The processed data samples are input into the trained effluent total phosphorus prediction model, and the water quality index of the effluent total phosphorus in the sewage treatment process is predicted to obtain the final sewage water quality soft measurement results.

[0210] In this embodiment, Figure 4 The predicted result of effluent COD; X-axis: test sample, unit is piece; Y-axis: predicted value of effluent total phosphorus concentration, unit is mg / L, the solid line is the measured value of effluent total phosphorus concentration, and the dotted line is the predicted value of effluent total phosphorus concentration;

[0211] It can be seen that the predicted value of the effluent total phosphorus concentration obtained by the present invention is basically fitted with its actual value curve, with high accuracy, and can achieve accurate prediction of the effluent total phosphorus.

[0212] The method for predicting total phosphorus in effluent from a sewage treatment process based on a hierarchical decomposition-integrated neural network provided by the present invention collects real-time water quality parameters in the sewage treatment process and environmental factors that affect water quality indicators, and predicts the total phosphorus concentration in the effluent from the sewage treatment process based on an STL decomposition, a CEEMDNAN decomposition method and an EA-GRU prediction model, and compares the theoretical sewage parameters with the actual sewage parameters. If the error is too large, it is considered that the sewage treatment process needs to be adjusted. Experimental data proves that the method for predicting total phosphorus in effluent from a sewage treatment process based on a hierarchical decomposition-integrated neural network proposed by the present invention has good prediction effect, and can help improve the measurement and control accuracy of sewage treatment plants over sewage treatment processes.

[0213] The specific implementation methods described above provide a detailed description of the technical solutions and beneficial effects of the present invention. It should be understood that the above is only the most preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, supplements and equivalent substitutions made within the scope of the principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for predicting total phosphorus in effluent from sewage treatment process based on hierarchical decomposition-integrated neural network, characterized in that: include: S10, preprocessing the sewage time series data and dividing it into a training set and a data set; S20, based on the hierarchical decomposition method, decompose the temporal characteristics of the time series; S30, constructing a prediction model corresponding to the trend component and residual component obtained by the first-level decomposition and the multi-period component obtained by the second-level decomposition; S40, training the prediction model and making predictions, integrating the prediction results of all components, and verifying the accuracy of the model; S50, inputting the real-time sewage data into the hierarchical decomposition-integrated neural network model to obtain the prediction result of the effluent total phosphorus concentration; The sewage data include flow rate, pH value, conductivity, chemical oxygen demand (COD), water temperature, salinity, residual chlorine, total phosphorus, redox potential, dissolved oxygen, and total dissolved solids content; The step S10 comprises the following steps: S101, deleting abnormal data in the historical sewage data using the 3σ criterion: If the deviation v between the historical total phosphorus data and the average value of the historical total phosphorus data satisfies v>3σ, where σ is the standard deviation of the historical total phosphorus data, the historical total phosphorus data is considered to be abnormal data and the abnormal data is deleted; S102, normalizing the cleaned historical total phosphorus data by using a Z-Score standardization method: The formula is: where y Z is the normalized value of the historical total phosphorus data y, is the average value of historical total phosphorus data; A time series feature decomposition method based on hierarchical decomposition is constructed, wherein the first layer decomposition is in the time domain and the second layer decomposition is in the time-frequency domain, and the step S20 includes the following steps: S201, time domain decomposition based on STL decomposition, decomposing the total phosphorus time series into trend component, seasonal component and residual component in the time domain; The STL decomposition algorithm transforms the total phosphorus data y Z Decompose into trend component y_t, seasonal component y_s and residual component y_r, distinguish components with different characteristics in the input data, and analyze the three components separately to achieve detailed analysis of sewage time series data; The specific steps of the STL data decomposition are: STL executes the inner loop in each outer loop. The steps of the inner loop are: S2011, detrending: in, It is nth d The detrended sequence of the inner cycle, It is nth d -1 inner cycle of trend sequence; S2012, Smooth Cyclic Subsequence: Based on the nth o -1 external loop right Perform LOESS smoothing to obtain a temporary seasonal series in Ψ is the double square weighted formula; S2013, low pass filter: Perform 3 sliding averages and then perform LOESS smoothing to obtain the residual trend S2014, Decomposition of smooth cyclic subsequences: in It is nth d seasonal component of the sub-inner cycle; S2015, Deseasonalization: in It is nth d Deseasonalized series of sub-inner cycles; S2016, Smooth trend item: Based on right Perform LOESS smoothing to obtain the trend component After all inner loops of an outer loop are completed, subtract the trend and seasonal components from the input data to get the nth o The residual term of the outer loop After all the outer loops of STL are completed, input data y Z =y_t+y_s+y_r; S202, time-frequency domain decomposition based on CEEMDAN, decomposing multi-periodic features in seasonal components; The CEEMDAN algorithm is selected to decompose the seasonal component to obtain the multi-periodic features in the seasonal component y_s. The specific decomposition process of CEEMDAN is as follows: S2021. Calculate the first IMF: Add M pairs of positive and negative Gaussian white noise to y_s to obtain y_s w , M is the number of white noise groups, M is 250; To y_s w Perform the first EMD decomposition, where imf j (1) is the jth modal component of the first decomposition; calculate the average value of the modal component to obtain the first modal component imf (1) : Where J is the maximum number of decomposable modal components, which is determined by the decomposition process itself; S2022. Calculate the first residual r (1) :r (1) =y_s-imf (1) ; S2023. Calculate the kth IMF: Let EMD (k) (·) is the k-th EMD decomposition; (k) Add M groups of paired positive and negative Gaussian white noise to get right Perform EMD decomposition to obtain the k+1th IMF: Calculate the k+1th residual r (k+1) :r (k+1) =r (k) -imf (k+1) ; If the stopping condition is met, that is, r (k+1) If it is a monotonic signal, the iteration stops and the CEEMDAN algorithm decomposition ends; At this point, the original signal is the sum of k IMFs and the residual, completing the decomposition: Among them, imf (k) is the IMF obtained from the k-th decomposition, and r is the final residual obtained; since r represents noise that cannot be further decomposed, no further prediction is performed; After the CEEMDAN decomposition, the seasonal component For the trend component and residual component obtained by the first-layer decomposition and the multi-period component obtained by the second-layer decomposition, a multi-layer integrated deep learning model is constructed to make predictions respectively and integrate them to generate the final prediction results; The step S30 comprises the following steps: S301. Construct ARIMA model to predict trend component: The ARIMA model is used to predict the stationary trend component obtained from the first level decomposition: Among them, h is the prediction step size, is the autoregressive coefficient, θ n is the moving average coefficient, p and q are the autoregressive order and the moving average order, ε t is white noise, c is a constant, is the predicted value of the trend component; and θ n It is estimated by the model through the fitting process; S302, construct EA-GRU model to predict periodic components: The EA-GRU model is used to predict the IMF representing multi-period features obtained by the second-level decomposition; First calculate the energy of each IMF: Among them, IMF t (k) It is IMF (k) The tth data point, N is the imf (k) Length; In order to remove redundant IMFs, the energy ratio of each IMF is calculated: Then sort the IMFs in descending order, calculate the sum of the energy ratios starting from the first IMF, and get the cumulative energy ratio. When the cumulative energy ratio of the Lth IMF reaches more than 95%, the remaining IMFs will be regarded as redundant IMFs and eliminated; For each filtered IMF, an improved attention-GRU module is established for prediction: For the lth attention-GRU module, the GRU includes two gate structures, the update gate and the reset gate: Update Gate t (l) Determine the discarding and updating of the state information at the previous moment: in, is the update gate weight matrix, is the input value of the lth GRU, sig is the sigmoid activation function; Reset Gate Controls how much information from the previous state is written: in, is the reset gate weight matrix; Current hidden state That is, the output of the lth attention-GRU module is calculated as follows: The candidate hidden state in yes The weight matrix of By further considering the importance of each IMF, the attention-GRU module introduces an improved attention mechanism to update the input of GRU. The improved attention weight calculation formula is: The calculation formula is: in, and is the learnable weight matrix of the lth attention-GRU module, is the energy ratio of the lth IMF, indicating the importance of IMF; therefore, the input of the lth GRU Updated to: Update Gate t (l) Determine the discarding and updating of the state information at the previous moment: in, is the update gate weight matrix, is the input value of the lth GRU, sig is the sigmoid activation function; Reset Gate Controls how much information from the previous state is written: in, is the reset gate weight matrix; Current hidden state That is, the output of the lth attention-GRU module is calculated as follows: The candidate hidden state in yes The weight matrix of The attention-GRU module introduces an improved attention mechanism to update the input of GRU. The improved attention weight calculation formula is: The calculation formula is: in, and is the learnable weight matrix of the lth attention-GRU module, is the energy ratio of the lth IMF, indicating the importance of IMF; therefore, the input of the lth GRU Updated to: After all l attention-GRU modules have completed their predictions, the second-layer decomposition and integration module adds up the prediction results of all multi-period components to obtain the predicted value of the seasonal component y_s S303, construct the GRU model to predict the residual component: Construct a GRU model to predict the residual component y_r because it has strong nonlinearity; GRU consists of two gate structures, the update gate and the reset gate: Update Gate t Determines the discarding and updating of the state information of the previous moment: r t =sig(W r [h t-1 ,y_r t ]); Among them, W r is the update gate weight matrix, y_r t is the input value of GRU; Reset gate z t Controls how much information of the previous state is written: z t =sig(W z [h t-1 ,y_r t ]); Among them, W z is the reset gate weight matrix; Current hidden state h t , which is the output of the GRU model The calculation is as follows: The candidate hidden state Where W h yes The weight matrix of S304, the first-layer decomposition integration module adds the prediction results of all multi-period components, trend components and residual components to obtain the final prediction value: S40, train the integrated deep learning model and verify the accuracy of the model; The hierarchical decomposition-integration neural network model training process provided in the embodiment includes the following steps: S401. Use the STL decomposition algorithm on the input variables: The STL decomposition algorithm is used to decompose the effluent ammonia nitrogen into trend component, seasonal component and residual component; S402, determine the auxiliary variables for predicting effluent ammonia nitrogen: The Pearson correlation coefficient represents the covariance of two variables divided by the product of their standard deviations, which can achieve translation invariance. It measures the linear correlation between two variables and gives a value between -1 and 1. The specific formula for calculating the Pearson correlation coefficient ρ(I,y) between each auxiliary variable I and the effluent ammonia nitrogen parameter y is as follows: Among them I l is the lth data of the auxiliary variable, y l is the lth data of the effluent ammonia nitrogen parameter, where l = (1, 2, ..., N``), N`` is the number of measured parameters, is the mean value of the auxiliary variable, is the average value of effluent ammonia nitrogen parameters; The input auxiliary variables of effluent ammonia nitrogen were selected by comparing the Pearson correlation coefficient of each auxiliary variable, and five input auxiliary variables were calculated: pH value, conductivity, water temperature, COD, and residual chlorine; S402, importing the training set into the model for training; Step 1, train the ARIMA model using the trend component; Step 2: Use CEEMDAN to decompose the seasonal components, select the IMFs based on the energy of each decomposed IMF, and train the EA-GRU model; Step 3, use the residual component to train the GRU model; Repeat steps 1-3 until all batches of the training data set have participated in model training and the specified number of iterations has been reached; the specified number of iterations is more than 250; S403, importing the test set data into the trained model to verify the model training effect; At the end of each round of cyclic training, the test set is imported into the integrated deep learning model to evaluate the model training effect, and the model prediction result is output; Input y_t into ARIMA to get the forecast result Decompose y_s by CEEMDAN and then input it into EA-GRU to get the prediction result Input y_r and five input variables into the GRU neural network to get the prediction results The outputs of all models are integrated and denormalized to obtain the final prediction result: Compare the prediction results output by the verified integrated deep learning model with the test set; When the prediction error does not decrease for 5 consecutive times, the training is stopped and the optimal hierarchical decomposition-integrated neural network model is obtained; otherwise, iterative training is performed again to update the network weights and bias of the hierarchical decomposition-integrated neural network model until the prediction error does not decrease for 5 consecutive times and the optimal hierarchical decomposition-integrated neural network model is obtained; S50, inputting the real-time sewage data into the hierarchical decomposition-integration neural network to obtain the prediction results of the sewage water quality index.

Citation Information

Cited By

  • Ecological environment big data processing method based on mining area pollutant migration behavior modeling

    CN120408565A

  • Ecological environment big data processing method based on modeling of pollutant migration behavior in mining area

    CN120408565B

  • Method for predicting effluent total phosphorus index in sewage treatment process

    CN122455144A