Water quality prediction method for water treatment plant based on neural network algorithm

By employing a water quality prediction method based on neural network algorithms and utilizing a long short-term memory network model with dynamic feature extraction and dual-flow adaptive attention mechanism, the nonlinear time-varying and abrupt change conditions in water quality prediction for water treatment plants were solved. This enabled accurate prediction and intelligent control of effluent COD, improving the refinement and energy efficiency of operation management.

CN121256506BActive Publication Date: 2026-05-12GUANGDONG FORCON ENG TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGDONG FORCON ENG TECH
Filing Date
2025-11-15
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing water quality prediction methods for water treatment plants suffer from low accuracy in predicting effluent COD and inaccurate operation control due to nonlinear and time-varying water quality, unpredictable critical conditions, delayed response to sudden changes in operating conditions, and crude regulation.

Method used

A water quality prediction method based on neural network algorithm is adopted. By collecting operational data and water quality data in real time, dynamic feature extraction, sedimentation tank sludge thickness hidden state encoding, static feature screening and interactive feature construction are used. Combined with a long short-term memory network model with dual-flow adaptive attention mechanism, the predicted value of chemical oxygen demand is output and the process parameters are dynamically adjusted.

Benefits of technology

It significantly improves the accuracy and robustness of COD time-series prediction, enhances the model's time sensitivity and physical interpretability, ensures stable compliance of effluent water quality, optimizes energy and chemical consumption, and achieves intelligent and refined operation management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121256506B_ABST
    Figure CN121256506B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of water quality prediction, in particular to a water quality prediction method for water treatment plants based on a neural network algorithm. The method comprises the following steps: collecting operation data and water quality data of a water treatment plant under different working conditions in real time, and preprocessing the operation data and the water quality data; based on the preprocessed operation data and water quality data, extracting multi-dimensional features reflecting the variation law of chemical oxygen demand through dynamic feature extraction, hidden state coding of sludge thickness in a sedimentation tank, static feature screening and interaction feature construction; based on the multi-dimensional features, predicting the water quality data by using a long short-term memory network model based on a double-flow adaptive attention mechanism, and outputting the predicted value of the chemical oxygen demand in a future period. In the feature extraction stage, the dynamic variation features of operation parameters, the static correlation screening features and the interactive coupling features between parameters are comprehensively considered.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of water quality prediction technology, and more specifically, to a water quality prediction method for water treatment plants based on neural network algorithms. Background Technology

[0002] Urban wastewater treatment involves complex biochemical reactions, sludge-water separation, and mass transfer mechanisms, resulting in multiphase coupling. Effluent quality (Chemical Oxygen Demand, COD) is significantly influenced by various factors, including influent flow rate, water quality fluctuations, temperature changes, sludge state, and process parameters (such as aeration intensity, chemical dosage, and reflux ratio), exhibiting significant nonlinearity, time-varying characteristics, and uncertainty. Traditional water quality prediction methods primarily rely on mechanistic models (such as the ASM series of activated sludge models). However, these models are complex, have numerous parameters, and depend on precise initial settings and continuous calibration, making them difficult to apply stably in practical engineering. Conventional statistical models (such as multiple linear regression and support vector machines) have limited expressive power and struggle to capture the complex dynamic relationships within the system, leading to insufficient prediction accuracy. Furthermore, key process states such as sludge thickness and settling performance in sedimentation tanks directly affect effluent quality, but the lack of reliable online monitoring methods makes real-time acquisition difficult, resulting in a lack of awareness of the system's internal state and limiting the accurate prediction of effluent COD trends. Meanwhile, wastewater treatment plants frequently face operational disturbances such as rainy season confluence, sudden changes in influent load, and abnormalities in the return flow system. Existing single-flow neural network models lack the ability to effectively distinguish between steady-state operating characteristics and transient disturbance characteristics, making it difficult to achieve sensitive responses to sudden operating conditions and easily leading to prediction lags or inaccuracies. More importantly, most water plants currently still adopt a passive control mode of "adjusting after exceeding standards," which has significant lag in regulation. This not only makes it difficult to ensure stable effluent compliance but also easily leads to problems such as excessive aeration and waste of chemicals, resulting in unnecessary consumption of energy and resources. Therefore, this paper proposes a water quality prediction method for water treatment plants based on neural network algorithms. Summary of the Invention

[0003] The purpose of this invention is to provide a water quality prediction method for water treatment plants based on neural network algorithms, so as to solve the problems of low accuracy in effluent COD prediction and inaccurate operation control in existing sewage treatment systems mentioned in the background art, which are caused by nonlinear time-varying water quality, unmeasurable critical states, delayed response to sudden operating conditions, and extensive regulation.

[0004] To achieve the above objectives, the present invention aims to provide a water quality prediction method for water treatment plants based on neural network algorithms, comprising the following steps:

[0005] S1. Real-time acquisition of operating data and water quality data of water treatment plants under different operating conditions, and preprocessing of operating data and water quality data;

[0006] S2. Based on the pre-processed operational data and water quality data, multi-dimensional features reflecting the change law of chemical oxygen demand are extracted through dynamic feature extraction, hidden state encoding of sludge thickness in sedimentation tank, static feature screening and interactive feature construction.

[0007] S3. Based on multidimensional features, a long short-term memory network model based on a dual-stream adaptive attention mechanism is used to predict water quality data and output the predicted value of chemical oxygen demand for future periods.

[0008] S4. Dynamically adjust the process parameters of the water treatment plant based on the predicted value of chemical oxygen demand to reduce pollution emissions.

[0009] As a further improvement to this technical solution, the feature is that: in S1, the operating data includes at least the influent flow rate, the dosage of chemicals, the aeration intensity, and the sludge discharge volume, and the water quality data includes at least the influent temperature, the historical effluent chemical oxygen demand concentration, the suspended solids concentration in the influent, the suspended solids concentration in the returned sludge, the suspended solids concentration in the effluent, and the suspended solids concentration in the discharged sludge.

[0010] As a further improvement to this technical solution, the feature is that: in step S2, extracting multidimensional features reflecting the changing pattern of chemical oxygen demand includes the following steps:

[0011] S2.1 Divide the running data into time series, set time windows, and calculate the dynamic characteristics of each running parameter in each window;

[0012] S2.2. For the sludge thickness in the sedimentation tank, the dynamic modal features of the thickness are extracted using a sludge settling dynamic coding method based on a sludge thickness simulator.

[0013] S2.3. Based on dynamic features, using the Pearson correlation coefficient, input variables related to chemical oxygen demand are selected as static features, and interactive features are constructed.

[0014] S2.4. The thickness dynamic modal features, dynamic features, static features, and interactive features are normalized using the Z-score method.

[0015] S2.5. The normalized thickness dynamic modal features, dynamic features, static features, and interaction features are concatenated into a multi-dimensional feature vector in chronological order. .

[0016] As a further improvement to this technical solution, the feature is that: in step S2.2, for the sludge thickness in the sedimentation tank, the dynamic modal features of the thickness are extracted using a sludge settling dynamic coding method based on a sludge thickness simulator, including the following steps:

[0017] S2.21. Based on the suspended solids concentration in the influent, the influent flow rate, the suspended solids concentration and return flow rate in the returned sludge, the sludge discharge rate, and the suspended solids concentration and effluent flow rate in the effluent, calculate the virtual net sludge flux at each time point. ;

[0018] S2.22. Construct a lightweight recurrent neural network as a sludge thickness simulator, and run it in parallel with a long short-term memory network model based on a dual-stream adaptive attention mechanism to generate hidden states. From hidden state Estimated suspended solids concentration in effluent is output via the fully connected layer. Estimated chemical oxygen demand of effluent And through the estimated value of suspended solids concentration in the effluent Estimated chemical oxygen demand of effluent Joint training is performed using auxiliary loss functions;

[0019] S2.23. After training, extract the hidden state vectors generated by the sludge thickness simulator at each time step as the thickness dynamic modal features of the sludge thickness.

[0020] As a further improvement to this technical solution, the feature is that: in S2.22, the construction of a lightweight recurrent neural network as a sludge thickness simulator involves the following specific steps: [The text abruptly shifts to a different topic] and its rate of change As input to the sludge thickness simulator, it reflects the dynamic accumulation and discharge trend of sludge in the sedimentation tank. Through time-step recursive calculation, the input virtual net sludge flux is... The sequence is encoded as a hidden state sequence. Hidden state The hidden state is explicitly defined as a representation of the normalized virtual sludge thickness, and a sludge thickness simulator is trained using an auxiliary loss function. The time evolution characteristics can characterize the dynamic response of suspended solids concentration and chemical oxygen demand in water.

[0021] As a further improvement to this technical solution, the feature is that: in step S3, a long short-term memory network model based on a dual-flow adaptive attention mechanism is used to predict water quality data and output a predicted value of chemical oxygen demand for future periods, including the following steps:

[0022] S3.1, Multidimensional feature vectors Arrange the input sequence in chronological order. ; and for the input sequence Perform normalization processing;

[0023] S3.2 Construct a long short-term memory network model based on a dual-stream adaptive attention mechanism;

[0024] S3.3. The mean squared error is used to train the long short-term memory network model based on the dual-stream adaptive attention mechanism;

[0025] S3.4, Input sequence The input is fed into a trained long short-term memory network model based on a two-stream adaptive attention mechanism;

[0026] S3.5, Outputting the Future Predicted values ​​of chemical oxygen demand at each time step.

[0027] As a further improvement to this technical solution, the feature is that: in step S3.2, constructing a long short-term memory network model based on a dual-stream adaptive attention mechanism involves the following specific steps: converting multidimensional feature vectors... Separate into process steady-state feature vectors and reflux dynamic feature vector At each time step of the Long Short-Term Memory Network model, the sensitivity weight of the reflux dynamic feature is calculated, the input gate, forget gate, and output gate of the Long Short-Term Memory Network model are extended into a two-stream structure, and a reflux mutation early warning and compensation mechanism is added to monitor the change in the reflux ratio in real time.

[0028] Among them, the dynamic characteristic flow of reflux includes at least the instantaneous change rate of reflux ratio, the reflux cumulative effect index, and the coupling characteristics of reflux-sludge thickness.

[0029] As a further improvement to this technical solution, the following steps are involved in calculating the sensitivity weights of the backflow dynamic features at each time step of the Long Short-Term Memory network model:

[0030] At each time step The hidden state from the previous step Current backflow dynamic feature vector Perform normalization processing and change the hidden state from the previous step. Current backflow dynamic feature vector Concatenate into attention input vector , input attention vector By performing a linear mapping and nonlinear activation, an intermediate representation is obtained. ; to represent the middle With bias term Calculate the attention score and obtain the sensitivity weights using the sigmoid activation function. ;Sensitivity weight The weighted reflux feature is obtained by multiplying the original reflux feature element by element. And input it into the extended dual-stream gated long short-term memory network model.

[0031] As a further improvement to this technical solution, the feature is that: the expansion of the input gate, forget gate, and output gate of the Long Short-Term Memory network model into a two-stream structure involves the following specific steps:

[0032] The steady-state feature vector of the process and reflux dynamic feature vector Each gate is constructed independently, and the activation values ​​of the input gate, forget gate, and output gate of the long short-term memory network model are calculated separately.

[0033] The results from the gated channels are weighted and combined using a gated fusion unit, and the weights are based on the sensitivity of the backflow dynamic characteristics. Adaptively adjust the fusion weights during the weighted combination process;

[0034] Calculate the steady-state eigenvectors of the process respectively and reflux dynamic feature vector Candidate memory cells in the gated channel;

[0035] The gated outputs of candidate memory units are weighted and fused, and the candidate memory units are updated based on the weighted and fused outputs, thereby realizing the collaborative memory and adaptive update of steady-state and dynamic information.

[0036] Output the hidden state at the current moment.

[0037] As a further improvement to this technical solution, the feature is that: in step S4, the process parameters of the water treatment plant are dynamically adjusted based on the predicted value of chemical oxygen demand, including the following steps:

[0038] S4.1 Compare the predicted chemical oxygen demand (COD) value obtained by the long short-term memory model with the set target COD range, and calculate the predicted COD deviation.

[0039] S4.2. Based on the predicted direction and magnitude of the deviation in water chemical oxygen demand, determine the control strategy and generate control instructions;

[0040] S4.3 Transmit the control command to the automatic control system to drive the aeration and reflux device to perform parameter adjustments in order to reduce pollutant emissions.

[0041] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0042] 1. The water quality prediction method for water treatment plants based on neural network algorithms disclosed in this invention comprehensively considers the dynamic change characteristics of operating parameters (such as moving average and rate of change), static correlation screening characteristics, and interactive coupling characteristics between parameters during the feature extraction stage. It also introduces a sludge thickness simulator based on GRU, utilizing its hidden states to implicitly and dynamically encode the sludge settling process in the sedimentation tank. This effectively captures the deep-seated evolutionary patterns of internal biological reactions and solid-liquid separation processes that are difficult to directly reflect in traditional monitoring data. This multi-dimensional feature fusion strategy fully explores the key factors affecting effluent COD and their nonlinear correlation mechanisms, providing a highly discriminative input space for subsequent prediction models and significantly improving the accuracy and robustness of COD time-series prediction.

[0043] 2. The water quality prediction method for water treatment plants based on neural network algorithms involved in this invention divides input features into steady-state process flow and dynamic return flow through a constructed dual-flow adaptive attention LSTM network. Differential processing and adaptive fusion of these two types of information are achieved through sensitivity weights and a gating fusion mechanism. Specifically, the model introduces a return flow mutation early warning compensation mechanism to monitor instantaneous changes in key parameters such as the return ratio in real time, dynamically adjusting attention weights and memory update paths. This enables the model to respond quickly to influent load shocks or operational anomalies, avoiding prediction inaccuracies due to delayed learning. This mechanism not only enhances the model's temporal sensitivity and physical interpretability but also provides a reliable basis for the forward-looking control of subsequent process parameters, effectively ensuring stable effluent quality compliance while optimizing energy and chemical consumption, achieving intelligent and refined operation management. Attached Figure Description

[0044] Figure 1 This is a flowchart of the overall method of the present invention. Detailed Implementation

[0045] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0046] Example: Please refer to Figure 1 As shown, this embodiment provides a water quality prediction method for water treatment plants based on neural network algorithms, including the following steps:

[0047] S1. Real-time acquisition of operational and water quality data from water treatment plants (municipal wastewater treatment plants) under different operating conditions, and preprocessing of the operational and water quality data (preprocessing includes: imputing missing values ​​and handling outliers to remove obvious measurement noise and abnormal fluctuations; then time alignment and resampling of each data to ensure that multi-source data correspond at the same time step; then normalization of continuous features (such as Z-score standardization) to make data of different dimensions comparable).

[0048] In this embodiment, the operational data includes at least the influent flow rate, chemical dosage, aeration intensity, and sludge discharge volume, and the water quality data includes at least the influent temperature, historical effluent chemical oxygen demand (COD) concentration, suspended solids concentration in the influent, suspended solids concentration in the returned sludge, suspended solids concentration in the effluent, and suspended solids concentration in the discharged sludge.

[0049] S2. Based on the pre-processed operational data and water quality data, multi-dimensional features reflecting the variation law of chemical oxygen demand (COD) are extracted through dynamic feature extraction, hidden state encoding of sludge thickness in sedimentation tank, static feature screening and interactive feature construction.

[0050] In this embodiment, the extraction of multidimensional features reflecting the variation pattern of chemical oxygen demand (COD) includes the following steps:

[0051] S2.1. The operating data is divided into time series segments, and time windows are set. Within each window, the rate of change, first-order difference, moving average, and moving standard deviation of each operating parameter are calculated to characterize the trend of COD changes over time. Specifically, within each time window, the rate of change (the ratio of the difference between adjacent time points to the value of the previous time point) of each operating parameter sequence is calculated to reflect the instantaneous fluctuation characteristics of the operating condition; the first-order difference is calculated to characterize the time series change trend of the parameters; and the moving average and moving standard deviation are calculated using the moving window method to extract the statistical characteristics of stationarity and volatility; finally, these dynamic features are mapped to the sample labels (historical effluent COD concentration) corresponding to the time window to form a feature set reflecting the changing pattern of the operating state, providing basic data support for the input of the subsequent neural network model.

[0052] S2.2 Regarding the sludge thickness in the sedimentation tank, a dynamic coding method for sludge settling based on a sludge thickness simulator is used to extract dynamic modal features of the thickness. The purpose is to extract modal features that can reflect the dynamic evolution of the sludge settling process in the sedimentation tank by encoding the implicit temporal features of sludge thickness changes, so as to provide a deep state characterization for subsequent water quality prediction models.

[0053] Among them, the method of extracting dynamic modal features of sludge thickness using a sludge thickness simulator-based dynamic encoding method for sludge settling specifically solves the problem that traditional methods are unable to accurately and directly quantify the dynamic impact of the complex physical and biochemical process of sludge settling in sedimentation tanks on the effluent chemical oxygen demand (COD). In the wastewater treatment process, sludge thickness and its rate of change are key intrinsic variables reflecting the operating status of sedimentation tanks and solid-liquid separation efficiency. It is affected by a combination of factors such as influent load, reflux ratio, and sludge discharge operation. Its dynamic process is highly nonlinear and has significant time lag. Direct measurement by traditional sensors is not only costly and prone to failure, but the single thickness value obtained cannot fully reveal its internal dynamic evolution law and the deep correlation with effluent COD. The advantage and role of this method is that it abandons the limitations of relying on direct measurement by physical sensors or simply using historical thickness data as input. By constructing a lightweight GRU network as a "sludge thickness simulator", driven by easily obtainable process parameters such as virtual net sludge flux, it implicitly learns and encodes the complete dynamic process of sludge accumulation-compression-discharge. Its core function is to transform the continuous physical state (sludge settling dynamics) that is difficult to observe directly into a low-dimensional hidden state vector sequence rich in semantic information (i.e., thickness dynamic modal features). This feature not only compensates for the shortcomings of direct measurement, but also, through the constraint of the auxiliary loss function, contains the dynamic response law that is directly related to the changes in effluent suspended solids and COD concentration. This provides a more essential and discriminative process state representation for the subsequent COD prediction model, significantly improving the prediction accuracy and generalization ability of the model under complex working conditions.

[0054] For the sludge thickness in the sedimentation tank, a dynamic modal feature of the thickness is extracted using a sludge settling dynamic coding method based on a sludge thickness simulator, including the following steps:

[0055] S2.21. Based on the suspended solids concentration in the influent, the influent flow rate, the suspended solids concentration and return flow rate in the returned sludge, the sludge discharge rate, and the suspended solids concentration and effluent flow rate in the effluent, calculate the values ​​for each time step. Virtual net sludge flux :

[0056] ;

[0057] S2.22. Construct a lightweight recurrent neural network (GRU) as a sludge thickness simulator, and run it in parallel with a long short-term memory network model based on a dual-stream adaptive attention mechanism to generate hidden states. From hidden state Estimated suspended solids concentration in effluent is output via the fully connected layer. Estimated chemical oxygen demand of effluent And through the estimated value of suspended solids concentration in the effluent Estimated chemical oxygen demand of effluent The auxiliary loss function is used for joint training to optimize the representation ability of the hidden state;

[0058] The lightweight recurrent neural network (GRU) architecture is based on virtual net sludge flux. and its rate of change As the input sequence, the current input and the hidden state from the previous time step are processed through an input gate and a reset gate. Gating operations are performed, and the intermediate layer uses an update gate to fuse new input and historical state information to generate the current hidden state. This hidden state encodes the dynamic process of sludge accumulation, compression, and discharge in the sedimentation tank, and is also passed to subsequent modules as a time-series representation feature. At the same time, the hidden state is mapped to the predicted output through a fully connected layer. and The network is trained on the output through an auxiliary loss function, so that the hidden state can effectively represent the dynamic evolution of sludge thickness, achieving lightweight and low computational complexity time series modeling.

[0059] Furthermore, constructing a lightweight recurrent neural network (GRU) as a sludge thickness simulator involves the following specific steps: [The text then abruptly shifts to a seemingly unrelated topic about virtual net sludge flux.] and its rate of change As input to the sludge thickness simulator, it reflects the dynamic accumulation and discharge trend of sludge in the sedimentation tank. Through time-step recursive calculation, the input virtual net sludge flux is... The sequence is encoded as a hidden state sequence. Hidden state The hidden state is explicitly defined as a representation of the normalized virtual sludge thickness. This state captures the dynamic process of sludge accumulation, compression, and discharge in the sedimentation tank. The sludge thickness simulator is trained using an auxiliary loss function to make the hidden state... The time evolution characteristics can characterize the dynamic response of suspended solids concentration and chemical oxygen demand (COD) in water, thereby realizing implicit modeling of the sludge settling process in sedimentation tanks.

[0060] Among them, the auxiliary loss function for:

[0061] ;

[0062] In the formula, the auxiliary loss function The optimization objective used when training the sludge thickness simulator (GRU) measures the error between the model's predicted effluent suspended solids concentration and the actual chemical oxygen demand (COD). For at any time Actual measured concentration of suspended solids in the effluent (true value). For at any time Actual measured chemical oxygen demand (COD) of effluent (true value). To control the weighting coefficient of the suspended solids concentration prediction error in the loss, To control the weighting coefficient of chemical oxygen demand prediction error in the loss, This represents the total number of time steps.

[0063] This auxiliary loss function minimizes the squared errors of the model-predicted effluent suspended solids concentration and chemical oxygen demand relative to the true values, thereby reducing the hidden state generated by the GRU. To better reflect the intrinsic relationship between the dynamic process of sludge in sedimentation tanks and changes in effluent quality;

[0064] After training, the hidden state sequence It is directly used as the thickness dynamic modal feature input to the subsequent feature extraction module for chemical oxygen demand (COD) prediction;

[0065] S2.23. After training, extract the hidden state vectors generated by the sludge thickness simulator at each time step as the thickness dynamic modal features of the sludge thickness.

[0066] S2.3. Based on dynamic characteristics, using the Pearson correlation coefficient, input variables related to chemical oxygen demand (COD) are selected as static features, and interactive features are constructed to reveal the coupling effect between different operating parameters and their potential impact on COD changes. Specifically, using historical effluent COD concentration as the target variable, the Pearson correlation coefficient between the dynamic characteristics of each operating parameter and COD is calculated. (In the formula, This refers to the dynamic characteristic sequence of operating parameters (such as the rate of change of influent flow rate, the moving average of chemical dosage, etc.). The target variable is the historical chemical oxygen demand (COD) concentration series in the effluent. For variables The standard deviation reflects its fluctuation range. For variables The standard deviation of the correlation coefficient was used to select features whose absolute values ​​exceeded a preset threshold (e.g., 0.6) as static input variables significantly related to COD. Subsequently, these static features were combined in pairs to construct interactive features. The product terms were used to characterize the coupling relationship and synergistic effect between operating parameters (e.g., the product term of chemical dosage and aeration intensity can reflect the synergistic effect of chemical dosing and biochemical reaction efficiency, and the product term of influent flow rate and temperature can reflect the linkage effect of load shock and reaction rate). This formed a composite feature set that could more fully characterize the process influence mechanism, providing a highly correlated and discriminative input space for subsequent prediction models.

[0067] S2.4. The thickness dynamic modal features, dynamic features, static features, and interactive features are normalized using the Z-score method to make features of different dimensions comparable and to make each type of feature have zero mean and unit variance on the numerical scale.

[0068] S2.5. The normalized thickness dynamic modal features, dynamic features, static features, and interaction features are concatenated into a multi-dimensional feature vector in chronological order. In the formula, Indicates at time step The Each feature component Indicates the feature component index, from 1 to The data originate from different feature types, including: thickness dynamic modal features (reflecting the dynamic law of sludge settling), dynamic features (rate of change of operating parameters, moving average, etc.), static features (steady-state input variables that are highly correlated with COD), and interactive features (product coupling terms between operating parameters). This represents the total number of dimensions of the input features at this time step, that is, the total number of features contained in the concatenated multidimensional feature vector.

[0069] S3. Based on multidimensional features, a long short-term memory (LSTM) network model based on a dual-stream adaptive attention mechanism is used to predict water quality data and output the predicted value of chemical oxygen demand (COD) for future periods (the chemical oxygen demand (COD) here refers to the chemical oxygen demand (COD) of the effluent).

[0070] In this embodiment, a Long Short-Term Memory (LSTM) network model based on a dual-flow adaptive attention mechanism is used to predict water quality data. This is mainly to address the problems of low prediction accuracy and poor adaptability of traditional models in water quality prediction for water treatment plants, caused by complex operating conditions, coupling of multiple factors (especially dynamic parameters such as return sludge), and varying timeliness of their effects. Specifically, it addresses the complex nonlinear, time-varying, and strongly coupled relationship between process parameters such as influent load, return ratio, and aeration intensity and effluent COD. Traditional single LSTM models struggle to distinguish between steady-state patterns and dynamic disturbances, especially exhibiting slow response to critical events such as abrupt changes in the return process, leading to significant lag or deviation in prediction results when operating conditions fluctuate. The core advantage of this dual-flow adaptive attention mechanism LSTM lies in its ability to proactively and adaptively identify and distinguish between two different types of information flows: "process steady-state characteristics" and "return dynamic characteristics." Its function is to enable the model to focus on learning long-term stable operating rules through independent gating channels and sensitivity weight calculations, while responding quickly to short-term dynamic disturbances such as sudden changes in the reflux ratio. The attention mechanism and early warning compensation mechanism are like installing an "intelligent scheduler" in the model, which can dynamically adjust the attention to different features and times, thereby significantly improving the model's prediction accuracy, robustness and real-time performance under complex and time-varying conditions, and providing a more reliable decision-making basis for subsequent precise regulation.

[0071] A long short-term memory (LSTM) network model based on a dual-stream adaptive attention mechanism is used to predict water quality data and output the predicted value of chemical oxygen demand (COD) for future periods. The steps include:

[0072] S3.1, Multidimensional feature vectors Arrange the input sequence in chronological order. ;

[0073] in, The length of the time window;

[0074] And for the input sequence Normalization (Z-score method) is performed to ensure that all features are within the same scale, which facilitates network training convergence.

[0075] S3.2 Construct a Long Short-Term Memory (LSTM) network model based on a dual-stream adaptive attention mechanism;

[0076] The architecture of the Long Short-Term Memory (LSTM) network model consists of three parts: an input layer, a hidden layer (intermediate flow), and an output layer. The input layer receives multi-dimensional feature vectors arranged in a time series, with each time step containing both steady-state and dynamic features. The hidden layer is composed of several LSTM units, each containing an input gate, a forget gate, and an output gate. The gating mechanism controls the memory and forgetting of information in the time dimension. A two-stream structure or attention mechanism can be introduced to enhance the responsiveness to different feature dimensions and time steps. The output layer maps the hidden state to the target predicted value, namely the chemical oxygen demand (COD) concentration in the future time period, through a fully connected layer, thus achieving the goal of time-series prediction and dynamic regulation.

[0077] Furthermore, constructing a Long Short-Term Memory (LSTM) network model based on a dual-stream adaptive attention mechanism involves the following specific steps: converting multi-dimensional feature vectors... Separate into process steady-state feature vectors and reflux dynamic feature vector At each time step of the Long Short-Term Memory (LSTM) network model, the sensitivity weight of the reflux dynamic feature is calculated, the input gate, forget gate, and output gate of the LSTM network model are extended into a two-stream structure, and a reflux mutation early warning compensation mechanism is added to monitor the change in the reflux ratio in real time.

[0078] Among them, the dynamic characteristic flow of reflux includes at least the instantaneous change rate of reflux ratio, the reflux cumulative effect index, and the coupling characteristics of reflux-sludge thickness;

[0079] Furthermore, at each time step of the Long Short-Term Memory (LSTM) network model, the sensitivity weights of the backflow dynamic features are calculated, involving the following steps:

[0080] At each time step The hidden state from the previous step Current backflow dynamic feature vector Perform normalization processing and change the hidden state from the previous step. Current backflow dynamic feature vector Concatenate into attention input vector , input attention vector By performing a linear mapping and nonlinear activation, an intermediate representation is obtained. (In the formula, This is the attention weight matrix, used to weight the input vector. Linear mapping to an intermediate representation space The hyperbolic tangent activation function is used to introduce nonlinearity and make the intermediate representation... Capable of capturing complex feature relationships, and representing intermediate... With bias term Calculate the attention score and apply it through the sigmoid activation function (i.e., ... Obtain sensitivity weights (This weight) (used to measure the influence of each dimension of backflow feature on the current state update), sensitivity weights. The weighted reflux feature is obtained by multiplying the original reflux feature element by element. And input it into the extended dual-stream gated long short-term memory (LSTM) network model;

[0081] During training, The parameters of the Long Short-Term Memory (LSTM) network model are learned through backpropagation, while an L1 sparse regularization term is introduced. To suppress excessive focus on irrelevant features, a time smoothing constraint term was added. To ensure the continuity and physical rationality of sensitivity weights over time, This is the attention weight matrix, used to weight the concatenated attention input vector. Linear mapping to The intermediate representation space of the dimension; For the corresponding bias term; This is the projection matrix, used to represent the intermediate representation. Projection to and reflow dynamic characteristics Use spaces of the same dimension to calculate attention scores; These are the bias terms for the projection operation; these parameters, along with the LSTM network model parameters, are learned and optimized through backpropagation; dimensionality. , , These are the dimensions of the LSTM hidden state, the reflow dynamic features, and the attention intermediate representation, respectively, and are set according to the specific application scenario;

[0082] Furthermore, the input gate, forget gate, and output gate of the Long Short-Term Memory (LSTM) network model are extended into a two-stream structure, involving the following specific steps:

[0083] The steady-state feature vector of the process and reflux dynamic feature vector Each gate is constructed independently, and the activation values ​​of the input gate, forget gate, and output gate of the Long Short-Term Memory (LSTM) network model are calculated separately. (In the LSTM network, the input features are divided into process steady-state features.) and reflux dynamic characteristics Two types of features are constructed, with independent gating channels for each type. This means that the activation values ​​of their respective input gate, forget gate, and output gate are calculated separately. This allows the steady-state feature channel to focus on capturing the long-term stable behavior of the system, while the dynamic feature channel focuses on reflecting short-term fluctuations caused by reflux and sludge changes. Then, the gating outputs of the two types of features are weighted and combined through a gating fusion unit, enabling the model to simultaneously retain steady-state and dynamic information and achieve adaptive prediction of chemical oxygen demand (COD) changes.

[0084] ;

[0085] ;

[0086] ;

[0087] ;

[0088] ;

[0089] ;

[0090] In the formula, The steady-state feature vector of the process The activation value of the input gate, The steady-state feature vector of the process The activation value of the forget gate, The steady-state feature vector of the process The activation value of the output gate, For the dynamic feature vector of reflux The activation value of the input gate, For the dynamic feature vector of reflux The activation value of the forget gate, For the dynamic feature vector of reflux The activation value of the output gate, The input gate pair process steady-state feature vector The weight matrix is ​​used to control the proportion of candidate memory writes. Forget gate pairs of process steady-state feature vectors The weight matrix is ​​used to control the proportion of memory retained from the previous time step. The output gate is the steady-state feature vector of the process. The weight matrix, The input gate corresponds to the previous hidden state. The weight matrix, To forget the previous hidden state The weight matrix, For the output gate, the previous hidden state The weight matrix, For the input gate bias term, For the forget gate bias term, This is the output gate bias term;

[0091] The results from the gated channels are weighted and combined using a gated fusion unit, and the weights are based on the sensitivity of the backflow dynamic characteristics. Adaptively adjust the fusion weights during the weighted combination process, enabling the model to dynamically respond to the impact of sudden backflow changes, for example:

[0092] ;

[0093] ;

[0094] ;

[0095] In the formula, For the fused input gate, For the Gate of Oblivion after fusion, For the output gate after fusion, For input gate fusion weights, For the forgetting gate, weights are integrated. For output gate fusion weights;

[0096] Calculate the steady-state eigenvectors of the process respectively and reflux dynamic feature vector Candidate memory cells in the gated channel:

[0097] , ;

[0098] In the formula, As a candidate memory cell for steady-state channels, Candidate memory units for dynamic channels. This is the input weight matrix for the candidate memory cells in the steady-state channel. This is the input weight matrix for dynamic channel candidate memory cells. This is the hidden state weight matrix of the candidate memory cells in the steady-state channel. This is the hidden state weight matrix for candidate memory cells in the dynamic channel. The bias term for candidate memory cells in the steady-state channel. This is the bias term for candidate memory cells in the dynamic channel. Weighted backflow dynamic feature vector , ;

[0099] The gated outputs of candidate memory units are then weighted and fused, and the candidate memory units are updated based on the weighted and fused outputs, thereby achieving collaborative memory and adaptive updating of steady-state and dynamic information:

[0100] ;

[0101] In the formula, This is a memory unit from the previous time step, storing long-term information. The memory unit for the current time step is obtained by fusing steady-state and dynamic channel candidate memory units and gating information for updating. The fusion weights for candidate memory units;

[0102] Output the hidden state at the current moment:

[0103] ;

[0104] Furthermore, the process of adding a reflux mutation early warning and compensation mechanism to monitor the change in the reflux ratio in real time is as follows: at each time step First, calculate the dynamic feature vector of the reflux. Instantaneous rate of change of reflux ratio (i.e., relative to reflux ratio) Perform first-order difference or relative change calculations to obtain the change magnitude at each time step and the cumulative change index (i.e., accumulate or weighted sum the changes in the reflux ratio within the historical time window, reflecting the cumulative effect of the reflux ratio change in the short term), and compare it with the hidden state of the previous time step. Input the attention module together to generate sensitivity weights Subsequently, the rate of change of the reflux ratio is analyzed to see if it exceeds a preset threshold. If it does, a sudden change warning signal is triggered, and this signal is used as a compensation factor in the weighted reflux characteristics. (When the reflux ratio becomes abnormal or changes abruptly, the warning signal will adjust the attention weight.) The corresponding dimension's value is used to enhance or suppress the backflow feature of that dimension on candidate memory units. The influence of this makes the LSTM network sensitive to mutations and responds promptly during memory updates, ensuring that the dynamic channel can capture the potential impact of backflow changes on chemical oxygen demand (COD). Through the fusion process of input gate, forget gate and candidate memory units, the hidden state update is dynamically adjusted to ensure that the LSTM network responds more sensitively to backflow mutations, thereby improving the stability and accuracy of COD prediction.

[0105] S3.3. The mean squared error (MSE) is used to train the long short-term memory (LSTM) network model based on the two-stream adaptive attention mechanism. Specifically, the input sequences in the training set are used to train the model. The data is fed sequentially into a trained two-stream LSTM network model (i.e., a Long Short-Term Memory (LSTM) network model based on a two-stream adaptive attention mechanism), and the model outputs a predicted sequence. Subsequently, the predicted values ​​were compared with the actual effluent chemical oxygen demand (COD) series. Stepwise alignment, calculate the predicted value at each time step. Compared with the true value Square error between The mean squared error (MSE) is obtained by averaging over all time steps and used as the loss function. Then, the gradient of the MSE loss with respect to all trainable parameters in the network (including LSTM gate parameters, two-stream fusion weights, attention weight matrix, etc.) is calculated using the backpropagation algorithm, and the parameters are iteratively updated using an optimizer (such as Adam or SGD) to gradually reduce the prediction error of the model, thereby completing the network training.

[0106] S3.4, Input sequence The input is fed into a trained long short-term memory network model based on a two-stream adaptive attention mechanism;

[0107] S3.5, Outputting the Future Predicted values ​​of chemical oxygen demand (COD) at each time step.

[0108] S4. Dynamically adjust the process parameters of the water treatment plant based on the predicted value of chemical oxygen demand (COD) to reduce pollution discharge and ensure that the effluent meets the discharge standards.

[0109] In this embodiment, the process parameters of the water treatment plant are dynamically adjusted based on the predicted value of chemical oxygen demand (COD), including the following steps:

[0110] S4.1 Compare the predicted chemical oxygen demand (COD) value obtained by the long short-term memory (LSTM) model with the set target chemical oxygen demand (COD) range, calculate the deviation of predicted water chemical oxygen demand, and reflect the degree of deviation of water quality.

[0111] S4.2. Based on the predicted direction and magnitude of the deviation of water chemical oxygen demand, determine the control strategy and generate control instructions; when the deviation of water chemical oxygen demand is positive and the deviation increases, increase the aeration rate or extend the aeration time to enhance the oxidation reaction of organic matter; when the deviation of water chemical oxygen demand is negative and the deviation decreases, appropriately reduce the aeration rate or the reflux ratio to avoid excessive treatment and energy waste.

[0112] S4.3 Transmit the control command to the automatic control system, drive the aeration and reflux device to perform parameter adjustment, and monitor the response trend of the adjustment results to ensure that the chemical oxygen demand (COD) of the effluent is consistently lower than the discharge standard, thereby reducing pollutant emissions.

[0113] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely preferred examples and are not intended to limit the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention.

Claims

1. A water quality prediction method for water treatment plants based on neural network algorithms, characterized in that, Includes the following steps: S1. Real-time acquisition of operating data and water quality data of water treatment plants under different operating conditions, and preprocessing of operating data and water quality data; S2. Based on the pre-processed operational data and water quality data, multi-dimensional features reflecting the change law of chemical oxygen demand are extracted through dynamic feature extraction, hidden state encoding of sludge thickness in sedimentation tank, static feature screening and interactive feature construction. S3. Based on multidimensional features, a long short-term memory network model based on a dual-stream adaptive attention mechanism is used to predict water quality data and output the predicted value of chemical oxygen demand for future periods. S4. Dynamically adjust the process parameters of the water treatment plant based on the predicted value of chemical oxygen demand to reduce pollution emissions; In step S2, extracting multidimensional features reflecting the changing patterns of chemical oxygen demand includes the following steps: S2.1 Divide the running data into time series, set time windows, and calculate the dynamic characteristics of each running parameter in each window; S2.

2. For the sludge thickness in the sedimentation tank, the dynamic modal features of the thickness are extracted using a sludge settling dynamic coding method based on a sludge thickness simulator. S2.

3. Based on dynamic features, using the Pearson correlation coefficient, input variables related to chemical oxygen demand are selected as static features, and interactive features are constructed. S2.

4. The thickness dynamic modal features, dynamic features, static features, and interactive features are normalized using the Z-score method. S2.

5. The normalized thickness dynamic modal features, dynamic features, static features, and interaction features are concatenated into a multi-dimensional feature vector in chronological order. ; In step S2.2, regarding the sludge thickness in the sedimentation tank, the dynamic modal features of the thickness are extracted using a sludge settling dynamic coding method based on a sludge thickness simulator, including the following steps: S2.

21. Based on the suspended solids concentration in the influent, the influent flow rate, the suspended solids concentration and return flow rate in the returned sludge, the sludge discharge rate, and the suspended solids concentration and effluent flow rate in the effluent, calculate the virtual net sludge flux at each time point. ; S2.

22. Construct a lightweight recurrent neural network as a sludge thickness simulator, and run it in parallel with a long short-term memory network model based on a dual-stream adaptive attention mechanism to generate hidden states. From hidden state Estimated suspended solids concentration in effluent is output via the fully connected layer. Estimated chemical oxygen demand of effluent And through the estimated value of suspended solids concentration in the effluent Estimated chemical oxygen demand of effluent Joint training is performed using auxiliary loss functions; S2.

23. After training, extract the hidden state vectors generated by the sludge thickness simulator at each time step as the thickness dynamic modal features of the sludge thickness. In step S3, a long short-term memory network model based on a dual-flow adaptive attention mechanism is used to predict water quality data and output the predicted value of chemical oxygen demand for future periods, including the following steps: S3.1, Multidimensional feature vectors Arrange the input sequence in chronological order. ; and for the input sequence Perform normalization processing; S3.2 Construct a long short-term memory network model based on a dual-stream adaptive attention mechanism; S3.

3. The mean squared error is used to train the long short-term memory network model based on the dual-stream adaptive attention mechanism; S3.4, Input sequence The input is fed into a trained long short-term memory network model based on a two-stream adaptive attention mechanism; S3.5, Outputting the Future Predicted values ​​of chemical oxygen demand at each time step.

2. The water quality prediction method for water treatment plants based on neural network algorithms according to claim 1, characterized in that: In S1, the operating data includes at least the influent flow rate, chemical dosage, aeration intensity, and sludge discharge volume, and the water quality data includes at least the influent temperature, historical effluent chemical oxygen demand concentration, suspended solids concentration in the influent, suspended solids concentration in the returned sludge, suspended solids concentration in the effluent, and suspended solids concentration in the discharged sludge.

3. The water quality prediction method for water treatment plants based on neural network algorithms according to claim 1, characterized in that: In step S2.22, the construction of a lightweight recurrent neural network as a sludge thickness simulator involves the following specific steps: [The text then abruptly shifts to a different topic:] ...virtual net sludge flux... and its rate of change As input to the sludge thickness simulator, it reflects the dynamic accumulation and discharge trend of sludge in the sedimentation tank. Through time-step recursive calculation, the input virtual net sludge flux is... The sequence is encoded as a hidden state sequence. Hidden state The hidden state is explicitly defined as a representation of the normalized virtual sludge thickness, and a sludge thickness simulator is trained using an auxiliary loss function. The time evolution characteristics can characterize the dynamic response of suspended solids concentration and chemical oxygen demand in water.

4. The water quality prediction method for water treatment plants based on neural network algorithms according to claim 1, characterized in that: In step S3.2, constructing a long short-term memory network model based on a dual-stream adaptive attention mechanism involves the following specific steps: converting multi-dimensional feature vectors... Separate into process steady-state feature vectors and reflux dynamic feature vector At each time step of the Long Short-Term Memory Network model, the sensitivity weight of the reflux dynamic feature is calculated, the input gate, forget gate, and output gate of the Long Short-Term Memory Network model are extended into a two-stream structure, and a reflux mutation early warning and compensation mechanism is added to monitor the change in the reflux ratio in real time. Among them, the dynamic characteristic flow of reflux includes at least the instantaneous change rate of reflux ratio, the reflux cumulative effect index, and the coupling characteristics of reflux-sludge thickness.

5. The water quality prediction method for water treatment plants based on neural network algorithms according to claim 4, characterized in that: The calculation of the sensitivity weights of the backflow dynamic features at each time step of the Long Short-Term Memory network model involves the following steps: At each time step The hidden state from the previous step Current backflow dynamic feature vector Perform normalization processing and change the hidden state from the previous step. Current backflow dynamic feature vector Concatenate into attention input vector , input attention vector By performing a linear mapping and nonlinear activation, an intermediate representation is obtained. ; to represent the middle With bias term Calculate the attention score and obtain the sensitivity weights using the sigmoid activation function. ;Sensitivity weight The weighted reflux feature is obtained by multiplying the original reflux feature element by element. And input it into the extended dual-stream gated long short-term memory network model.

6. The water quality prediction method for water treatment plants based on neural network algorithms according to claim 4, characterized in that: The process of extending the input gate, forget gate, and output gate of the Long Short-Term Memory (LSTM) network model into a two-stream structure involves the following specific steps: The steady-state feature vector of the process and reflux dynamic feature vector Each gate is constructed independently, and the activation values ​​of the input gate, forget gate, and output gate of the long short-term memory network model are calculated separately. The results from the gated channels are weighted and combined using a gated fusion unit, and the weights are based on the sensitivity of the backflow dynamic characteristics. Adaptively adjust the fusion weights during the weighted combination process; Calculate the steady-state eigenvectors of the process respectively and reflux dynamic feature vector Candidate memory cells in the gated channel; The gated outputs of candidate memory units are weighted and fused, and the candidate memory units are updated based on the weighted and fused outputs, thereby realizing the collaborative memory and adaptive update of steady-state and dynamic information. Output the hidden state at the current moment.

7. The water quality prediction method for water treatment plants based on neural network algorithms according to claim 1, characterized in that: In step S4, the process parameters of the water treatment plant are dynamically adjusted based on the predicted value of chemical oxygen demand, including the following steps: S4.1 Compare the predicted chemical oxygen demand (COD) value obtained by the long short-term memory model with the set target COD range, and calculate the predicted COD deviation. S4.

2. Based on the predicted direction and magnitude of the deviation in water chemical oxygen demand, determine the control strategy and generate control instructions; S4.3 Transmit the control command to the automatic control system to drive the aeration and reflux device to perform parameter adjustments in order to reduce pollutant emissions.