A method for predicting total phosphorus index of effluent in sewage treatment process

CN122455144BActive Publication Date: 2026-09-29JIANGXI NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610904373.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-23
Publication Date
2026-09-29
Estimated Expiration
2046-06-23

AI Technical Summary

Technical Problem

[0005]针对现有技术中存在的预测滞后、难以进行前瞻性调控问题,本发明提供了一种污水处理过程中出水总磷指标的预测方法,包括如下步骤:

Benefits of technology

本发明构建深度学习主干预测+残差二次校正的两阶段闭环架构,第一阶段通过深度学习网络捕捉出水总磷变化的宏观时序规律,第二阶段通过梯度提升树模型挖掘深度学习无法捕捉的局部非线性耦合关系与异常波动特征,二者协同实现了趋势预测与细节校正的兼顾。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122455144B_ABST
    Figure CN122455144B_ABST
Patent Text Reader

Abstract

The application discloses a method for predicting total phosphorus index of effluent in a sewage treatment process. The method comprises the following steps: constructing a data set containing multi-dimensional time sequence characteristics and preprocessing; performing batch normalization and working condition random mask on input data; adopting exponential moving average to decompose into trend item data and residual error item data; inputting the trend item data and the residual error item data into two LSTM-Attention branches respectively to extract long-period trend and multi-feature coupling dependence; dynamically fusing trend item features and residual error item features through an adaptive gating mechanism; finally, performing initial prediction through a deep learning backbone network, and performing secondary correction on the prediction residual by a lightweight gradient boosting machine model to output a final total phosphorus prediction value. Through time sequence decoupling, double-branch customized extraction, adaptive fusion and two-stage hybrid correction, the application significantly improves the prediction accuracy and robustness, and is suitable for early warning of total phosphorus of effluent in a sewage plant.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent control and water quality time-series prediction technology for wastewater treatment, specifically a method for predicting the total phosphorus index in effluent during wastewater treatment. Background Technology

[0002] Wastewater treatment, as a core component of water environment management and water resource recycling, uses total phosphorus in its effluent as a key constraint for controlling the risk of eutrophication. Accurate and proactive prediction of total phosphorus is a crucial prerequisite for pre-emptive control of process parameters, stabilizing effluent quality, and reducing energy and chemical consumption. Traditional laboratory testing and online instrument monitoring suffer from significant time delays, supporting only post-event traceability and failing to provide pre-emptive decision-making support. With the development of industrial IoT and artificial intelligence technologies, time-series data-driven predictive models have become a research hotspot in the industry, but they still face multiple technical bottlenecks in practical engineering applications.

[0003] Existing models are insufficient in capturing long-distance temporal dependencies and biochemical lags, making it difficult to effectively model the lag effects caused by the hydraulic retention time of several hours in oxidation ditches. Furthermore, the dual-flow architecture oversimplifies the handling of trend components, resulting in insufficient long-term correlation mining. Meanwhile, single-flow models struggle to balance trend fitting during stable water quality periods with fluctuating responses during influent shock periods, easily leading to overfitting during stable periods and a sharp drop in accuracy during shock periods. The contradiction between the error accumulation of iterative multi-step prediction and the high parameter count of direct multi-step prediction further restricts the feasibility of model implementation at edge environments. Existing feature fusion strategies also cannot dynamically adapt to the switching between "trend-driven" and "fluctuation-driven" operating conditions, amplifying prediction bias.

[0004] Furthermore, existing solutions are insufficient in exploring the strong nonlinear coupling relationships among multiple variables such as influent orthophosphate concentration, polyaluminum chloride dosage, and dissolved oxygen. This makes it difficult to capture hidden single-feature time-series dependencies and multi-feature correlations from residual fluctuations, resulting in weak predictive capabilities for abnormal fluctuations in effluent total phosphorus and early warning of exceedance risks. At the same time, Transformer-type models have poor adaptability to small sample data and lack customized robust designs for industrial scenarios such as sensor failure, data loss, and equipment maintenance. They are prone to prediction failure when data quality is poor. In addition, the lack of error correction mechanisms for strong noise and strong coupling characteristics makes it difficult to meet the engineering requirements for stable 24 / 7 operation of wastewater treatment plants. Summary of the Invention

[0005] To address the problems of prediction lag and difficulty in proactive control in existing technologies, this invention provides a method for predicting total phosphorus levels in wastewater during wastewater treatment, comprising the following steps: Step S1: Construct a dataset for predicting total phosphorus in wastewater effluent and perform data preprocessing to obtain time-series data; divide the dataset into training and validation sets according to the proportions. Step S2: Input the time series data into the pre-built dual-process time series prediction model, perform batch normalization and condition-based random mask preprocessing, and obtain the preprocessed data; Step S3: Perform exponential moving average decomposition on the preprocessed data to obtain trend term data and residual term data; Step S4: Input the trend data into the first LSTM-Attention processing flow to capture the long-term trend and time lag characteristics of the change in total phosphorus in the effluent, and output the trend sequence features; Step S5: Input the residual term data into the second LSTM-Attention processing flow to obtain the time dependency information within a single data feature and the time dependency information between multiple data features, and output the residual sequence features; Step S6: Input the trend sequence features and residual sequence features into the prediction module of the dual-process time series prediction model. The prediction module includes an adaptive gated feature fusion module and a two-stage hybrid prediction layer. The adaptive gated feature fusion module dynamically fuses the trend sequence features and residual sequence features through a gating mechanism to generate fused time series features. The two-stage hybrid prediction layer includes a deep learning backbone network and a residual correction model. The deep learning backbone network outputs the initial prediction value based on the fused time series features, and the residual correction model learns and corrects the residual of the initial prediction value to obtain the final total phosphorus prediction result in the effluent.

[0006] Further, in step S2, the condition-based random mask preprocessing operation is as follows: First, the operating conditions are divided into three categories: stable operating conditions, shock operating conditions, and abnormal operating conditions according to the influent flow rate and influent orthophosphate concentration; the mask probability for stable operating conditions is 1% and the duration of continuous masking follows a normal distribution N(2,0.25), the mask probability for shock operating conditions is 3% and the duration of continuous masking follows a normal distribution N(4,1), and the mask probability for abnormal operating conditions is 5% and the duration of continuous masking follows a normal distribution N(6,2.25).

[0007] Further, in step S3, the trend term and residual term of the exponential moving average decomposition are expressed as follows: ; ; Where t represents the current time; c is the feature index, representing different water quality monitoring parameters; This represents the water quality fluctuation coefficient at feature index c at the current time t. This represents the original monitoring value of feature index c at the current time t; The trend term of the feature index c at the current time t is the smoothed value after the exponential moving average. The trend term representing the feature index c of the previous time step t-1 is used for recursive calculation; The residual term for the feature index c at the current time t represents the degree to which the original value deviates from the trend term.

[0008] Further, in step S4, the first LSTM-Attention processing procedure includes: using an LSTM network to perform temporal encoding on the trend item data, and outputting the temporal encoded features through the gating mechanism of the LSTM network. The temporal coding features output by the LSTM network are weighted through an additive attention layer; the attention calculation process is as follows: ; ; ; in, denoted as attention score; W represents the first linear layer; tanh is the activation function; V represents the second linear layer; attn_weights are the attention weights, softmax is the activation function; dim is the dimension and channel; attn_out is the weighted aggregated feature; For time-series coding features; seq_len is the sequence length.

[0009] Further, in step S5, the second LSTM-Attention processing flow includes: performing temporal encoding on the residual term data through an LSTM network; performing cross-feature weighting on the multi-feature temporal encoding output by the LSTM network through an additive attention layer; adding a dropout layer to the attention output to suppress overfitting; and mapping the features back to the original feature dimension through a linear projection layer to output residual sequence features.

[0010] Furthermore, in step S6, the adaptive gated feature fusion module specifically performs the following operations: concatenating the trend sequence features with the residual sequence features to obtain the concatenated features. The gate weights are generated using a linear layer and a sigmoid activation function, and the calculation formula is as follows: ; in, For gating weights; This is for splicing features; Linear() is a linear layer processing operation; Dynamic feature fusion is performed based on gating weights, and the fused temporal features are output. The calculation formula is as follows: ; in, Features of trend sequences; Features of the residual sequence; To integrate temporal features.

[0011] Further, in step S6, the two-stage hybrid prediction layer includes: The first stage constructs a multilayer perceptron linear prediction head, with the input being the feature of the last time step fused with temporal features, and the output being the initial prediction value of the deep learning backbone; The second stage involves constructing a residual correction model. This model takes as input three types of features: the flattened features of the original time-series data, the last-step features of the fused time-series features, and the time-series mean features of the fused time-series features. The model is trained with the residual between the deep learning predictions and the actual values ​​on the training set as the target. During inference, the residual correction model outputs the residual correction value. The initial deep learning prediction value is added to the residual correction value to obtain the final prediction result of total phosphorus in the effluent.

[0012] Furthermore, the method for predicting the total phosphorus index in effluent during wastewater treatment also includes a prediction model for the total phosphorus index in effluent during wastewater treatment, comprising: The input layer is used for batch normalization and condition-based random mask preprocessing of the input data; The exponential moving average decomposition module is used to perform exponential moving average decomposition on the preprocessed data to obtain trend term data and residual term data. The first LSTM-Attention processing module is used to input the trend item data into the first LSTM-Attention processing flow and output the trend sequence features. The second LSTM-Attention processing module is used to input the residual term data into the second LSTM-Attention processing flow and output residual sequence features. An adaptive gating feature fusion module is connected to the outputs of the first LSTM-Attention processing module and the second LSTM-Attention processing module, respectively. It dynamically fuses the trend sequence features and residual sequence features through an adaptive gating mechanism to generate fused temporal features. The two-stage hybrid prediction layer includes a deep learning backbone network and a residual correction model. The deep learning backbone network performs initial predictions on the fused temporal features, and the residual correction model corrects the prediction residuals to output the final total phosphorus prediction result in the effluent.

[0013] Compared with the prior art, the present invention has the following beneficial effects: This invention constructs a two-stage closed-loop architecture of deep learning backbone prediction + residual secondary correction. The first stage captures the macro-temporal pattern of total phosphorus change in effluent through a deep learning network. The second stage uses a gradient boosting tree model to mine local nonlinear coupling relationships and abnormal fluctuation characteristics that deep learning cannot capture. The two work together to achieve both trend prediction and detailed correction.

[0014] This invention designs isomorphic but functionally differentiated dual LSTM-Attention parallel branches. The trend term branch specifically captures long-term temporal dependencies caused by biochemical reaction lags, improving the insufficient trend capture capability of existing linear flow methods. The residual term branch deeply mines the internal temporal patterns of single features and the coupling correlations between multiple features, effectively extracting water quality fluctuation information. The dual-branch collaboration achieves accurate fitting of trends during stable periods and early capture of fluctuations during shock periods, providing a reliable basis for pre-process control.

[0015] All modules of this invention are customized for the actual engineering scenarios of wastewater treatment plants: input features are screened based on biochemical mechanisms to fit the actual operating rules of phosphorus removal processes in wastewater treatment plants; the condition-based random mask module simulates continuous data loss caused by equipment maintenance, enabling the model to make stable predictions even when there are sensor failures or poor data quality. Attached Figure Description

[0016] Figure 1 This is a schematic diagram of the prediction model process of the present invention. Detailed Implementation

[0017] To make the technical solution of the present invention easier to understand and implement, the specific implementation steps are further described below in conjunction with the embodiments and accompanying drawings, but the embodiments are not intended to limit the present invention.

[0018] This invention provides a method for predicting the total phosphorus index in effluent during wastewater treatment, comprising the following steps: Step S1: Construct a dataset for predicting total phosphorus in wastewater effluent and perform data preprocessing to obtain time-series data; divide the dataset into training and validation sets according to the proportions. Step S2: Input the time series data into the pre-built dual-process time series prediction model, perform batch normalization and condition-based random mask preprocessing, and obtain the preprocessed data; Step S3: Perform exponential moving average decomposition on the preprocessed data to obtain trend term data and residual term data; Step S4: Input the trend data into the first LSTM-Attention processing flow to capture the long-term trend and time lag characteristics of the change in total phosphorus in the effluent, and output the trend sequence features; Step S5: Input the residual term data into the second LSTM-Attention processing flow to obtain the time dependency information within a single data feature and the time dependency information between multiple data features, and output the residual sequence features; Step S6: Input the trend sequence features and residual sequence features into the prediction module of the dual-process time series prediction model. The prediction module includes an adaptive gated feature fusion module and a two-stage hybrid prediction layer. The adaptive gated feature fusion module dynamically fuses the trend sequence features and residual sequence features through a gating mechanism to generate fused time series features. The two-stage hybrid prediction layer includes a deep learning backbone network and a residual correction model. The deep learning backbone network outputs the initial prediction value based on the fused time series features, and the residual correction model learns and corrects the residual of the initial prediction value to obtain the final total phosphorus prediction result in the effluent.

[0019] Further, in step S1, the dataset is constructed, which is specifically divided into the following sub-steps: Step S101, Data Acquisition: Based on the existing PLC (Programmable Logic Controller) sensing system and operating database of the wastewater treatment plant, historical operating data stored in the InfluxDB database is automatically acquired using a Python script and saved as a CSV file to obtain the dataset. Based on the biochemical mechanism of biological phosphorus removal + chemical phosphorus removal in wastewater treatment, Pearson correlation analysis is used to select the 13 input features with the strongest correlation to the total phosphorus index in the effluent. These are: secondary dosage of polyaluminum chloride in the east zone, secondary dosage of polyaluminum chloride in the west zone, dosage of polyacrylamide in the east zone, dosage of polyacrylamide in the west zone, ammonia nitrogen in the influent, pH value in the influent, sludge concentration in the aerobic tank, orthophosphate concentration in the influent, temperature in the middle section of the aerobic tank, oxidation-reduction potential in the anaerobic tank, dissolved oxygen concentration in the middle section of the aerobic tank, chemical oxygen demand in the influent, and total nitrogen in the influent. All indicators are expressed as actual operating monitoring values.

[0020] Step S102: Delete or replace unreasonable data in the dataset using a script. This is divided into two categories. One category involves treating large gaps exceeding 4 hours in the dataset as data breakpoints. The original time series is then cut into multiple independent valid data segments along the gap boundaries. Each valid data segment is treated as an independent time series sample, and attempts are no longer made to connect the data before and after the gap. After cutting, if the length of a valid data segment is less than the minimum time window length required by the downstream model (in this embodiment, the sliding window length is 7), the segment is discarded; otherwise, it is retained for subsequent processing. Another approach is to replace data with randomly generated data when there are obvious outliers or a few missing values ​​in the dataset. The random range of replacement values ​​is determined by the mean and standard deviation. The specific values ​​are determined based on the overall data distribution of the dataset, ensuring that 70% of the data falls within this range. Therefore, it is believed that data randomly generated within this range can maintain good consistency with the original data.

[0021] The method for identifying obvious outliers is similar to that described above. The outlier detection range only includes 3% of the total data, and this range is used to determine the real coefficients. Data falling within this range is considered an obvious outlier.

[0022] Step S103, Time Series Sample Construction: The time series training samples are constructed using the sliding window method. The length of the input sequence is set to 7 time steps (corresponding to 14 hours of historical data, covering the complete hydraulic retention cycle of the wastewater treatment plant), and the prediction step size is 4 time steps (corresponding to the total phosphorus concentration in the effluent in the next 8 hours). The sliding window step size is 1, generating three-dimensional time series samples in the format [batch, seq_len, in_channels], where batch is the sample batch, seq_len is the length of the input sequence, and in_channels is the original input feature dimension. In this embodiment, the input feature dimension is 13.

[0023] Step S104: Divide the time-series samples generated in step S103 into a training set and a validation set according to the time order, with a ratio of 8:2. Specifically, take the first 80% of the samples as the training set and the last 20% of the samples as the validation set.

[0024] In step S2, the dataset prepared in step S1 is input into the dual-process time series prediction model, and the data is processed as follows: Step S201: Batch normalization: Each batch of data is input into the dual-process time series prediction model for max-min normalization, scaling the data to a fixed interval [0,1] to better analyze indicator data at different orders of magnitude; during normalization, the minimum and maximum values ​​of each feature on the training set are used as scaling parameters. Step S202: Random Masking Module: This module uses a probabilistic algorithm to mask the input data with a low probability, simulating data loss scenarios caused by anomalies in the real world. First, based on the influent flow rate and influent orthophosphate concentration, the operating conditions are divided into three categories: stable operating conditions, shock operating conditions, and abnormal operating conditions. The specific classification rules are as follows: First, four statistics are pre-calculated based on the training set data: the median and 90th quantile of the influent flow rate, and the median and 90th quantile of the influent orthophosphate concentration. For any time window of 7 sampling points, the average influent orthophosphate concentration, the average influent flow rate, and the standard deviation of the influent orthophosphate concentration within the window are calculated. If the average influent orthophosphate concentration within the window reaches or exceeds the 90th percentile of the influent orthophosphate concentration in the training set, or the average influent flow rate reaches or exceeds the 90th percentile of the flow rate in the training set, or the standard deviation of the influent orthophosphate concentration within the window reaches or exceeds 0.15, it is judged as an abnormal operating condition. If the abnormal conditions are not met, but the ratio of the average influent orthophosphate concentration within the window to the median influent orthophosphate concentration in the training set reaches or exceeds 1.15, or the ratio of the average influent flow rate to the median flow rate in the training set reaches or exceeds 1.15, or the standard deviation of the influent orthophosphate concentration within the window reaches or exceeds 0.08, it is judged as a shock condition. All other cases are judged as stable operating conditions. The probability of masking under stable operating conditions is 1% and the duration of continuous masking follows a normal distribution N(2,0.25). The probability of masking under impact conditions is 3% and the duration of continuous masking follows a normal distribution N(4,1). The probability of masking under abnormal operating conditions is 5% and the duration of continuous masking follows a normal distribution N(6,2.25).

[0025] In step S3, the time series decomposition module adopts EMA (Exponential Moving Average) decomposition, which introduces the exponential exhaustion weight of historical data on the basis of exponential moving average decomposition. The more recent the data point, the greater the weight, and the more distant the data point, the smaller the weight, but it will never be zero.

[0026] The core idea of ​​EMA decomposition is to break down a complex time series into several components with different characteristics that are easier to understand and model. Its purpose is to reduce modeling difficulty, improve prediction accuracy, and enhance model interpretability by separating and isolating different components in the time series.

[0027] Furthermore, the trend term and residual term of the exponential moving average decomposition in step S3 are expressed as follows: ; ; Where t represents the current time; c is the feature index, representing different water quality monitoring parameters; This represents the original monitoring value of feature index c at the current time t; The trend term of the feature index c at the current time t is the smoothed value after the exponential moving average. The trend term representing the feature index c of the previous time step t-1 is used for recursive calculation; The residual term for the feature index c at the current time t represents the degree to which the original value deviates from the trend term; This represents the water quality fluctuation coefficient at feature index c at the current time t. The calculation formula is: ; in, This represents the minimum and maximum values ​​of the standard deviation of this feature calculated based on all samples in the current training set. The water quality fluctuation coefficient represented by feature index c at the current time t reflects the standard deviation of this feature over the most recent 7 time steps. The calculation formula is as follows: ; in, For feature index; This represents the original monitoring value / real-time data of feature index c at the current time t; This represents the average of the feature indices c from the 7 time steps prior to the current time t. As shown in Table 1, the complete feature index table contains 13 input features and 1 prediction target. The total phosphorus concentration (OutTP) in the effluent is only used during model training for subsequent error correction and is not included in the model input. Table 1 Feature Index Table

[0028] For key characteristics such as the secondary dosage of polyaluminum chloride in the eastern zone of PAC1, the secondary dosage of polyaluminum chloride in the western zone of PAC2, and the orthophosphate concentration in the influent of InTP2, an additional shock coefficient correction is applied: ; in, ; The biological phosphorus removal coefficient is used in this embodiment. =0.12; Let be the shock coefficient of the influent orthophosphate concentration at time t, calculated using the following formula: ; in, This refers to the orthophosphate concentration in the influent. is the shock coefficient of influent orthophosphate concentration at time t, used to quantify the degree of abnormal fluctuation in influent orthophosphate concentration; The truncation function restricts the calculation results to the interval [1.0, 2.0] to avoid interference from extreme values; The sequence of influent orthophosphate concentrations for the seven time steps prior to time t; This represents the local average concentration of orthophosphate in the influent over the past seven steps. This represents the global average concentration of orthophosphate in the influent.

[0029] Furthermore, in step S4, the trend term processing procedure adopts an LSTM-Attention fusion structure, specifically targeting the stationary temporal characteristics of the trend term, to capture the long-term trend and biochemical reaction lag characteristics of the total phosphorus change in the effluent. The specific steps are as follows: Step S401, LSTM network temporal encoding: The trend term is temporally encoded using an LSTM network. The original input feature dimension is 13. Through the gating mechanism of the LSTM network, the long-distance temporal dependence and the lag correlation caused by the hydraulic residence time in the trend term are captured. The output temporal encoding feature lstm_out has the shape [batch, seq_len, hidden_dim], where batch is the sample batch, seq_len is the length of the input sequence, and hidden_dim is the hidden layer dimension. In this embodiment, the hidden layer dimension is 64.

[0030] Step S402, Attention Weighting: The temporal coding features output by the LSTM network are weighted through an additive attention layer, automatically focusing on the key time steps that have the greatest impact on the trend of total phosphorus in the effluent, and weakening the interference of irrelevant time segments; the attention layer contains two linear layers, namely the first linear layer W and the second linear layer V, and the attention calculation process is as follows: ; ; ; in, denoted as attention score; W represents the first linear layer; tanh is the activation function; V represents the second linear layer; attn_weights are the attention weights, softmax is the activation function; dim is the dimension and channel; attn_out is the weighted aggregated feature; For time-series coding features; seq_len is the sequence length.

[0031] Step S403, Dropout Regularization and Linear Projection: Add a dropout layer (dropout rate 0.1) to the attention output attn_out to suppress overfitting; at the same time, map it to the original input feature dimension through a linear projection layer; map the temporal encoded features output by the LSTM network back to the original feature dimension, and output the trend sequence feature trend_seq with shape [batch, seq_len, in_channels] for subsequent feature fusion.

[0032] In step S5, the residual term processing adopts an LSTM-Attention structure isomorphic to the trend term branch, specifically targeting the nonlinear fluctuation characteristics of the residual term, and deeply mining the two types of core temporal dependency information hidden in the residual data. The specific steps are as follows: Step S501, Extraction of time dependence within a single feature: The residual term is temporally encoded using an LSTM network to capture the temporal fluctuation pattern of each feature, including the lag effect of PAC dosage, the periodic characteristics of dissolved oxygen changes, and other time dependence information within a single feature, and outputs the temporally encoded feature.

[0033] Step S502, Extraction of Coupling Dependencies Among Multiple Features: Through an additive attention layer (same as step S402 above), cross-feature weighting is performed on the multi-feature temporal encoding output by the LSTM network to mine the nonlinear coupling correlations among multiple features such as influent orthophosphate concentration, dosage, dissolved oxygen, and influent flow rate, including the response relationship between influent orthophosphate concentration shock and PAC dosage, the correlation between dissolved oxygen changes and biological phosphorus removal efficiency, etc. The features and time steps that have the greatest impact on the fluctuation of total phosphorus in the effluent are automatically weighted.

[0034] Step S503, Dropout Regularization and Linear Projection: Add a dropout layer (dropout rate 0.1) to the attention output to suppress overfitting. Map the features back to the original feature dimension through a linear projection layer and output the residual sequence feature res_seq with shape [batch, seq_len, in_channels], which is aligned with the trend sequence feature dimension for subsequent feature fusion.

[0035] Furthermore, step S6 is divided into two core components: adaptive gating feature fusion and two-stage hybrid prediction. The specific steps are as follows: Step S601, Adaptive Gated Feature Fusion: A gating mechanism is used to dynamically fuse the trend sequence feature trend_seq and the residual sequence feature res_seq, replacing the simple additive fusion of existing technologies. This achieves adaptive adjustment of feature contribution under different water quality conditions. The specific process is as follows: Step 1: Concatenate the trend sequence features and residual sequence features in the last dimension to obtain the concatenated features. The shape is [batch, seq_len, in_channels*2]; Step 2: Generate adaptive gate weights using a linear layer and a sigmoid activation function. The calculation formula is as follows: ; Where gate is the gate weight; This is for concatenating features; Linear() is a linear layer processing operation, with the input dimension of the Linear layer being in_channels*2 and the output dimension being in_channels; Step 3: Perform dynamic feature fusion based on gating weights, and output the fused temporal feature fused_seq. The calculation formula is as follows: ; in, To integrate temporal features; Features of trend sequences; The residual sequence features are used; the gate weight can be dynamically adjusted according to the water quality conditions: during the stable water quality period, the weight of the trend term features is increased to weaken the noise interference of the residual; during the water inflow shock period, the weight of the residual term features is increased to strengthen the capture of fluctuation information and achieve adaptive adaptation for all operating conditions.

[0036] Step S602, First Stage: Deep Learning Backbone Prediction. A two-layer MLP linear prediction head is constructed. The input is the last time step feature fused_seq[:, -1, :] of the fused temporal features. The final temporal state of the sequence is captured, and the initial prediction result of the deep learning backbone is output, corresponding to the total phosphorus concentration of the effluent in the next 8 hours. The structure of the linear prediction head includes: a linear layer used to map the dimension of the input features from the dimension of the fused temporal features to the dimension of the hidden layer; after passing through the ReLU activation function and the dropout layer; finally, through another linear layer, the hidden layer features are mapped to the prediction step size dimension. In this embodiment, the prediction step size is 4.

[0037] Step S603, Second Stage: In this embodiment, the LightGBM (Light Gradient Boosting Machine) model is used as the residual correction model for secondary residual correction. The specific process is as follows: Three types of features are extracted as input to the LightGBM model: flattening features of the original time series data, the final feature of the fused time series features, and the time series mean feature of the fused time series features. The three types of features are concatenated to form the input features of the LightGBM model. Model training is divided into two stages. The first stage completes the end-to-end training of the deep learning backbone network and calculates the residual between the deep learning predictions and the true values ​​on the training set. The second stage uses the above-mentioned concatenated features as input and the residual between the deep learning predictions and the true values ​​on the training set as the target value to train the LightGBM model. During model inference, the residual correction value is predicted by the trained LightGBM model. The initial prediction value of deep learning is added to the residual correction value to obtain the prediction result under the normalized scale. Then, the inverse normalization calculation is performed to obtain the predicted value of total phosphorus in the effluent with physical dimensions.

[0038] Furthermore, the method for predicting the total phosphorus index in effluent during wastewater treatment also includes a prediction model for the total phosphorus index in effluent during wastewater treatment, comprising: The input layer is used for batch normalization and condition-based random mask preprocessing of the input data; The exponential moving average decomposition module is used to perform exponential moving average decomposition on the preprocessed data to obtain trend term data and residual term data. The first LSTM-Attention processing module is used to input the trend item data into the first LSTM-Attention processing flow and output the trend sequence features. The second LSTM-Attention processing module is used to input the residual term data into the second LSTM-Attention processing flow and output residual sequence features. An adaptive gating feature fusion module is connected to the outputs of the first LSTM-Attention processing module and the second LSTM-Attention processing module, respectively. It dynamically fuses the trend sequence features and residual sequence features through an adaptive gating mechanism to generate fused temporal features. The two-stage hybrid prediction layer includes a deep learning backbone network and a residual correction model. The deep learning backbone network performs initial predictions on the fused temporal features, and the residual correction model corrects the prediction residuals to output the final total phosphorus prediction result in the effluent.

[0039] Reference Figure 1 In this embodiment, the prediction model processing flow is as follows: First, input multivariate time series data; then, using a mechanism-guided exponential moving average decomposition method, the input data is separated into trend components and residual components. The data then enters a dual-branch processing stage: Trend branch: using an LSTM network combined with an additive attention mechanism, focusing on capturing the evolutionary patterns over time; Residual branch: using an LSTM network combined with an additive attention mechanism, focusing on mining the correlation information in the feature dimension. Next, the trend features and residual features are dynamically fused through an adaptive gating mechanism, and then the initial prediction results of the deep learning stage are output through a multilayer perceptron network. Then, a two-stage hybrid correction stage is entered: the actual observed values ​​are compared with the initial prediction results, and the residuals between the two are calculated; key indicators are screened for the fusion features, and a LightGBM model is constructed, which is specifically used to fit and predict residual changes. Finally, the initial deep learning predictions are added to the residual corrections, and inverse normalization is performed to output the final predicted total phosphorus value in the effluent.

[0040] The above embodiments merely illustrate several implementation methods of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of protection of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the appended claims.

Claims

1. A method for predicting the total phosphorus index in effluent during wastewater treatment, characterized in that, Includes the following steps: Step S1: Construct a dataset for predicting total phosphorus in wastewater effluent and perform data preprocessing to obtain time-series data; divide the dataset into training and validation sets according to the proportions. Step S2: Input the time series data into the pre-built dual-process time series prediction model, perform batch normalization and condition-based random mask preprocessing, and obtain the preprocessed data; Step S3: Perform exponential moving average decomposition on the preprocessed data to obtain trend term data and residual term data; Step S4: Input the trend data into the first LSTM-Attention processing flow to capture the long-term trend and time lag characteristics of the change in total phosphorus in the effluent, and output the trend sequence features; The first LSTM-Attention processing flow includes: using an LSTM network to perform temporal encoding on the trend term data, and outputting the temporal encoded features through the gating mechanism of the LSTM network. The temporal coding features output by the LSTM network are weighted through an additive attention layer; the attention calculation process is as follows: ; ; ; in, denoted as attention score; W represents the first linear layer; tanh is the activation function; V represents the second linear layer; attn_weights are the attention weights, softmax is the activation function; dim is the dimension and channel; attn_out is the weighted aggregated feature; These are time-series encoded features; seq_len is the sequence length; Step S5: Input the residual term data into the second LSTM-Attention processing flow to obtain the time dependency information within a single data feature and the time dependency information between multiple data features, and output the residual sequence features; the second LSTM-Attention processing flow includes: performing temporal encoding on the residual term data through an LSTM network, performing cross-feature weighting on the multi-feature temporal encoding output by the LSTM network through an additive attention layer; adding a dropout layer to the attention output to suppress overfitting, and mapping the features back to the original feature dimension through a linear projection layer to output the residual sequence features; Step S6: Input the trend sequence features and residual sequence features into the prediction module of the dual-process time series prediction model. The prediction module includes an adaptive gated feature fusion module and a two-stage hybrid prediction layer. The adaptive gated feature fusion module dynamically fuses the trend sequence features and residual sequence features through a gating mechanism to generate fused time series features. The two-stage hybrid prediction layer includes a deep learning backbone network and a residual correction model. The deep learning backbone network outputs the initial prediction value based on the fused time series features, and the residual correction model learns and corrects the residual of the initial prediction value to obtain the final total phosphorus prediction result in the effluent.

2. The method for predicting total phosphorus levels in wastewater during wastewater treatment according to claim 1, characterized in that, In step S2, the condition-based random mask preprocessing operation is as follows: First, the operating conditions are divided into three categories: stable operating conditions, shock operating conditions, and abnormal operating conditions according to the influent flow rate and influent orthophosphate concentration. The mask probability for stable operating conditions is 1% and the duration of continuous masking follows a normal distribution N(2,0.25), the mask probability for shock operating conditions is 3% and the duration of continuous masking follows a normal distribution N(4,1), and the mask probability for abnormal operating conditions is 5% and the duration of continuous masking follows a normal distribution N(6,2.25).

3. The method for predicting total phosphorus levels in wastewater during wastewater treatment according to claim 1, characterized in that, In step S3, the trend term and residual term of the exponential moving average decomposition are expressed as follows: ; ; Where t represents the current time; c is the feature index, representing different water quality monitoring parameters; This represents the water quality fluctuation coefficient at feature index c at the current time t. This represents the original monitoring value of feature index c at the current time t; The trend term of the feature index c at the current time t is the smoothed value after the exponential moving average. The trend term representing the feature index c of the previous time step t-1 is used for recursive calculation; The residual term for the feature index c at the current time t represents the degree to which the original value deviates from the trend term.

4. The method for predicting total phosphorus levels in wastewater during wastewater treatment according to claim 1, characterized in that, In step S6, the adaptive gated feature fusion module specifically performs the following operations: concatenating the trend sequence features and the residual sequence features to obtain the concatenated features. ; The gate weights are generated using a linear layer and a sigmoid activation function, and the calculation formula is as follows: ; in, For gating weights; This is for splicing features; Linear() is a linear layer processing operation; Dynamic feature fusion is performed based on gating weights, and the fused temporal features are output. The calculation formula is as follows: ; in, Features of trend sequences; Features of the residual sequence; To integrate temporal features.

5. The method for predicting total phosphorus levels in wastewater during wastewater treatment according to claim 1, characterized in that, In step S6, the two-stage hybrid prediction layer includes: The first stage constructs a multilayer perceptron linear prediction head, with the input being the feature of the last time step fused with temporal features, and the output being the initial prediction value of the deep learning backbone; The second stage involves constructing a residual correction model. This model takes as input three types of features: the flattened features of the original time-series data, the last-step features of the fused time-series features, and the time-series mean features of the fused time-series features. The model is trained with the residual between the deep learning predictions and the actual values ​​on the training set as the target. During inference, the residual correction model outputs the residual correction value. The initial deep learning prediction value is added to the residual correction value to obtain the final prediction result of total phosphorus in the effluent.

Citation Information

Patent Citations

  • COD (Chemical Oxygen Demand) prediction method based on two-channel convolution time sequence mechanism

    CN119557631A

  • Method for predicting water quality of industrial wastewater

    CN121963944A