Wastewater recycling process optimization method and system based on recycling effect prediction
By using a wastewater reuse process optimization method based on long short-term memory networks and a recycling effect predictor, the problem of failing to identify abnormal risks in advance in the wastewater reuse process was solved, achieving precise and low-cost optimization and control, and reducing operating costs and equipment wear.
Patent Information
- Application Number
- CN202610857314.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-15
- Publication Date
- 2026-08-25
AI Technical Summary
Existing wastewater reuse processes lack the ability to identify future abnormal risk periods in advance, leading to delayed responses and blind interventions, which increases control costs and equipment wear and tear, and makes it impossible to achieve precise and low-cost optimization.
Based on historical wastewater reuse operation monitoring data, a multidimensional influent characteristic data sequence is predicted through a long short-term memory network. Combined with a recycling effect predictor and sliding window anomaly detection, abnormal operating periods are identified and optimized.
It has achieved forward-looking optimization of wastewater reuse processes, avoiding delayed responses and blind interventions, reducing operating costs and equipment wear and tear, and improving the accuracy and efficiency of regulation.
Smart Images

Figure CN122634449A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of process optimization technology, and more specifically to a method and system for optimizing wastewater reuse processes based on the prediction of recycling effects. Background Technology
[0002] Current optimization and control of wastewater reuse processes mostly rely on passive adjustments based on real-time operational monitoring data. These adjustments are primarily based on empirical or real-time response control of current influent characteristics, effluent indicators, and operating parameters, typically using indicators such as current water quality, quantity, or recovery rate as a basis. Process adjustments are made only after anomalies are detected. Some methods also attempt to use offline simulation or steady-state models for decision support, but overall, passive response and post-event adjustments remain the main approaches.
[0003] Existing wastewater reuse process optimization methods lack the ability to identify future abnormal risk periods in advance, resulting in significant defects in process optimization decisions: on the one hand, there is a delayed response, with adjustments only made after problems such as excessive effluent quality and decreased recovery rate occur, which not only results in high adjustment costs but also causes slow recovery of the wastewater reuse system; on the other hand, there is blind premature intervention, with excessive adjustments to process parameters during the stable operation period, leading to waste of reagents and energy, while accelerating the wear and tear of treatment equipment, making it impossible to achieve accurate and low-cost optimization of wastewater reuse processes. Summary of the Invention
[0004] This invention provides a method and system for optimizing wastewater reuse processes based on recycling effect prediction, aiming to solve the technical problem that existing technologies cannot achieve accurate and low-cost optimization of wastewater reuse processes.
[0005] In view of the above problems, the present invention provides a method and system for optimizing wastewater reuse processes based on the prediction of recycling effect.
[0006] In a first aspect, the present invention provides a method for optimizing wastewater reuse processes based on the prediction of recycling effects, including: Based on historical wastewater reuse operation monitoring data, predictive multidimensional influent characteristic data sequences are obtained within a preset operation window; Based on the predicted multidimensional influent characteristic data sequence, the fluctuation analysis of the predicted recycling effect is performed. Based on the predicted fluctuation coefficient, the pre-trained recycling effect predictor is called. Combined with the expected operating parameters of wastewater reuse, several predicted recycling effect index sequences within the preset operating window are predicted. Sliding window anomaly detection is performed on each of the predicted recovery effect index sequences. Based on the deviation between the predicted value and the theoretical expected value of each node, abnormal running nodes within the preset running window are identified, and abnormal running segment sequences are determined. The abnormal runtime sequence is taken as the time period to be optimized within the preset runtime window. During the wastewater reuse process, the wastewater reuse process is optimized and controlled in the time period to be optimized by combining the preset monitoring-feedback mechanism.
[0007] Secondly, the present invention provides a wastewater reuse process optimization system based on recycling effect prediction, comprising: The influent characteristic prediction module is used to predict and obtain a sequence of multidimensional influent characteristic data within a preset operating window based on historical wastewater reuse operation monitoring data. The recycling effect prediction module is used to perform a volatility analysis of the recycling effect prediction based on the predicted multidimensional influent feature data sequence, call the pre-trained recycling effect predictor based on the prediction volatility coefficient, and combine the expected operating parameters of wastewater reuse to predict a series of predicted recycling effect indicators within the preset operating window. An abnormal time period identification module is used to perform sliding window anomaly detection on each of the predicted recovery effect index sequences. Based on the deviation between the predicted value and the theoretical expected value of each node, it identifies the abnormal running nodes within the preset running window and determines the abnormal running time sequence. The process optimization and control module is used to take the abnormal running segment sequence as the time period to be optimized within the preset running window, and to perform optimization and control of the wastewater reuse process within the time period to be optimized in the wastewater reuse treatment process in combination with the preset monitoring-feedback mechanism.
[0008] One or more technical solutions provided in this invention have at least the following technical effects or advantages: This invention provides a method and system for optimizing wastewater reuse processes based on recovery effect prediction. It predicts a multi-dimensional influent characteristic data sequence within a preset operating window based on historical data, and performs fluctuation analysis on the predicted recovery effect. A pre-trained recovery effect predictor is then used to obtain multiple predicted recovery effect index sequences. Subsequently, a sliding window anomaly detection is performed on each index sequence, accurately identifying abnormal operating nodes and abnormal time periods by utilizing the deviation between the predicted and theoretical expected values. Finally, this time period is designated as the optimization period, and precise process control is implemented using a monitoring-feedback mechanism. This invention can pre-determine the time window when performance degradation may occur, avoiding delayed responses and blind interventions. It achieves pre-prediction, dynamic positioning, and targeted optimization of wastewater reuse processes, reducing operating costs and equipment wear. Attached Figure Description
[0009] Figure 1 A schematic flowchart of a wastewater reuse process optimization method based on recycling effect prediction provided in an embodiment of the present invention; Figure 2This is a schematic diagram of the wastewater reuse process optimization system based on recycling effect prediction provided in an embodiment of the present invention; The components represented by each number in the attached diagram are explained below: 11 Influent characteristic prediction module, 12 Recycling effect prediction module, 13 Abnormal period identification module, and 14 Process optimization and control module. Detailed Implementation
[0010] This invention provides a method and system for optimizing wastewater reuse processes based on recycling effect prediction, which addresses the technical problem that existing technologies cannot achieve accurate and low-cost optimization of wastewater reuse processes.
[0011] Example 1, as Figure 1 As shown, this invention provides a method for optimizing wastewater reuse processes based on recycling effect prediction, the method comprising: S100: Based on historical wastewater reuse operation monitoring data, predict and obtain the predicted multidimensional influent characteristic data sequence within the preset operation window.
[0012] In this embodiment of the invention, the influent water quality, quantity, and temperature in the wastewater reuse process exhibit strong temporal dynamics and nonlinear fluctuations. Traditional process optimization methods rely solely on current or historical measured data for feedback adjustment, failing to anticipate future influent trends and resulting in a lack of prior information to support subsequent recovery effect prediction and anomaly risk identification. Without accurate prediction of future multidimensional influent characteristics, true advance warning and precise control are difficult to achieve. Therefore, it is necessary to construct a predictive model based on historical operational monitoring data that can capture the temporal dependencies of influent characteristics, obtaining a predicted multidimensional influent characteristic data sequence within a preset operating window. This provides fundamental data input for subsequent fluctuation analysis of recovery effect prediction and anomaly period identification.
[0013] Step S100 in the method provided in this embodiment of the invention includes: Based on historical wastewater reuse operation monitoring data, historical multidimensional influent characteristic data sequences are collected within the historical operation window. These historical multidimensional influent characteristic data sequences include historical influent water quality characteristic sequences, historical influent water volume sequences, and historical influent temperature sequences. The historical operation window and the preset operation window have the same time span. A long short-term memory network was trained until convergence by collecting a sample training set, and a multi-dimensional water inflow feature time series prediction model was generated. The historical influent water quality characteristic sequence, historical influent water volume sequence, and historical influent temperature sequence are input into the multidimensional influent characteristic time series prediction model to obtain the predicted multidimensional influent characteristic data sequence within the preset operating window.
[0014] First, based on historical wastewater reuse operation monitoring data, a historical multidimensional influent characteristic data sequence is collected within the historical operation window. This historical multidimensional influent characteristic data sequence includes historical influent water quality characteristic sequences, historical influent water quantity sequences, and historical influent temperature sequences. The historical operation window and the preset operation window have the same time span. Historical wastewater reuse operation monitoring data refers to the real-time collection and storage of comprehensive monitoring data on influent, process operation, and effluent during the wastewater reuse plant's past stable operation. The historical operation window is the historical data time range used for training the multidimensional influent characteristic time series prediction model. Its time span is completely consistent with the preset operation window, serving as the foundational period for the multidimensional influent characteristic time series prediction model to learn the influent characteristic time series patterns. The preset operation window refers to the future target operation period that needs to be predicted in advance and for which process optimization should be carried out. The historical multidimensional influent characteristic data sequence refers to the set of influent water quality, water quantity, and temperature characteristic data arranged in a fixed time node order within the historical operation window.
[0015] The influent water quality characteristic sequence can further include time series of multiple water quality indicators such as chemical oxygen demand (COD), ammonia nitrogen (NH3-N), total phosphorus (TP), and suspended solids (SS). The influent flow rate sequence is the instantaneous or cumulative flow rate time series entering the treatment unit, and the influent temperature sequence is the influent water temperature time series.
[0016] For example, the preset operating window is set to the next 24 hours, with one monitoring node per hour, for a total of 24 nodes. The past 24 consecutive hours are selected as the historical operating window, and the collected historical influent water quality characteristic sequence contains 24 time points. The COD value at each point is [45, 48, 52, ..., 60] mg / L; the historical influent volume sequence is [120, 115, 118, ..., 125] m³. 3 / h; the historical influent temperature sequence is [18.2, 18.1, 18.0, ..., 17.8]℃.
[0017] Secondly, a sample training set is collected to train a Long Short-Term Memory (LSTM) network until convergence, generating a multi-dimensional influent feature time-series prediction model. The LSTM network is a recurrent neural network adept at processing time-series data, effectively capturing long-term temporal dependencies of influent features, avoiding the gradient vanishing problem of traditional neural networks, and adapting to the time-series prediction of fluctuating wastewater influent. The multi-dimensional influent feature time-series prediction model is based on the LSTM architecture and is a dedicated time-series model specifically designed for simultaneously predicting influent water quality, quantity, and temperature in future periods.
[0018] The training set consists of several pairs of input-output sequences. Each sample's input is a sequence of multidimensional water inflow characteristic data within a historical operating window, and its output is the actual multidimensional water inflow characteristic data sequence within the next preset operating window of the same time span. Multiple time window pairs are slidably extracted from the historical monitoring data to form the training samples.
[0019] First, a multidimensional water inflow feature time-series prediction model is constructed. Its main structure is a one-layer Long Short-Term Memory (LSTM) network with 64 neurons in the hidden layer and the activation function is the hyperbolic tangent function. After the LSTM layer, a fully connected output layer is connected. The output dimension is consistent with the feature dimension of the predicted sequence and the activation function is a linear function, which is used to map the hidden state of the LSTM to the predicted value.
[0020] Secondly, the training dataset is derived from at least one year of historical operational monitoring data of the wastewater reuse plant, including time series data of influent water quality (COD), water quantity, and temperature. The training method employs supervised learning, inputting the multidimensional features within the historical operational window of each sample into the model and outputting the corresponding multidimensional features within a preset operational window. The mean squared error (MSE) between the model output and the true label sequence is used as the loss function.
[0021] Next, the optimizer Adam was chosen, with an initial learning rate of 0.001, and a mini-batch gradient descent method was employed with a batch size of 32. During training, the sample set was divided into training, validation, and test sets at a ratio of 80%, 10%, and 10%, respectively. After each training epoch, the loss value on the validation set was calculated. When the validation set loss value no longer decreased within 20 consecutive epochs, the model was considered to have converged, training was stopped, and the current model parameters were saved, thus obtaining the trained multi-dimensional water inflow feature time-series prediction model.
[0022] For example, from historical data of the past year, a sample is extracted every 6 hours: the input is the COD, water volume, and temperature sequence within a certain 24-hour window, and the output is the actual COD, water volume, and temperature sequence within the next 24-hour window. Approximately 1200 samples are obtained. 80% of these are used for training, 10% for validation, and 10% for testing. The LSTM network uses the aforementioned single-layer 64-neuron structure. After 200 training iterations, the validation set loss function value drops below 0.002 and does not decrease further for multiple consecutive iterations, indicating model convergence. This yields the trained multi-dimensional water inflow feature time-series prediction model.
[0023] Finally, the historical influent water quality characteristic sequences, historical influent water volume sequences, and historical influent temperature sequences are input into the multidimensional influent characteristic time series prediction model to obtain the predicted multidimensional influent characteristic data sequence within the preset operating window. The collected multidimensional influent characteristic data sequence within the latest historical operating window is then input into the trained multidimensional influent characteristic time series prediction model according to the model's input format. The multidimensional influent characteristic time series prediction model performs forward calculations and outputs the multidimensional predicted values for each time point within the preset operating window. The output results are then inversely normalized to obtain the predicted multidimensional influent characteristic data sequence in actual physical units.
[0024] For example, the input data is normalized and then fed into a trained multidimensional influent feature time series prediction model. The model outputs prediction results for the next 24 hours. After inverse normalization, the predicted COD sequence is obtained as [62, 65, 70, ..., 55] mg / L, and the predicted water volume sequence is [130, 135, 128, ..., 118] m³ / L. 3 / h, predicting the temperature sequence [17.6, 17.5, 17.4, ..., 17.0]℃. This predicted multidimensional influent feature data sequence serves as the basic input for subsequent steps.
[0025] In this embodiment of the invention, the multi-dimensional influent feature prediction of the future operating window is completed based on the LSTM time series model. It can accurately capture the temporal fluctuation patterns of influent water quality, water quantity and temperature, and provide forward-looking and high-precision input data for subsequent recovery effect prediction. It solves the prediction lag problem caused by the reliance on real-time data in traditional methods from the source, and ensures the accuracy of subsequent recovery effect prediction and abnormal period identification.
[0026] S200: Based on the predicted multidimensional influent characteristic data sequence, the fluctuation analysis of the predicted recycling effect is performed. Based on the predicted fluctuation coefficient, the pre-trained recycling effect predictor is called. Combined with the expected operating parameters of wastewater reuse, several predicted recycling effect index sequences within the preset operating window are predicted.
[0027] In this embodiment of the invention, a volatility analysis of the predicted recycling effect is performed based on the predicted multidimensional influent characteristic data sequence. A pre-trained recycling effect predictor is invoked based on the predicted volatility coefficient, and combined with the expected operating parameters of wastewater reuse, several predicted recycling effect index sequences within a preset operating window are predicted. In the wastewater reuse process, the degree of volatility in influent characteristics directly affects the stability and predictability of the subsequent recycling effect. If the influent fluctuates drastically, a single fixed prediction model is difficult to accurately capture the changing trend of the recycling effect, requiring the use of diverse prediction branches to adapt to uncertainty; conversely, if the influent is stable, the number of prediction branches can be reduced to lower computational overhead. Therefore, it is necessary to perform volatility analysis on the predicted multidimensional influent characteristic data sequence, calculate the predicted volatility coefficient, and dynamically invoke different numbers of recycling effect prediction branches accordingly, thereby achieving a balance between prediction accuracy and computational efficiency. Simultaneously, the recycling effect predictor needs to be pre-trained using historical data to establish a complex mapping relationship between multidimensional influent characteristics, operating parameters, and recycling effect indicators. The purpose of this step is to achieve the adaptive invocation of the above volatility analysis and predictor, providing multiple possible recycling effect index sequences for subsequent anomaly detection.
[0028] Step S200 in the method provided in this embodiment of the invention includes: Among them, the volatility analysis of recovery effect prediction based on the predicted multidimensional influent characteristic data sequence includes: Parametric fluctuation analysis was performed on the predicted influent water quality characteristic sequence, predicted influent water quantity sequence, and predicted influent temperature sequence in the predicted multidimensional influent characteristic data sequence, and the coefficient of variation of water quality characteristic, influent water quantity, and influent temperature were calculated. The coefficients of variation of water quality characteristics, influent flow rate, and influent temperature are weighted and fused to obtain the comprehensive variability, which is used as the predictive fluctuation coefficient for predicting the recovery effect.
[0029] First, parameter fluctuation analysis was performed on the predicted influent water quality characteristic sequence, predicted influent water quantity sequence, and predicted influent temperature sequence from the predicted multidimensional influent characteristic data sequence, respectively, to calculate the coefficient of variation (CV) for water quality characteristics, influent water quantity, and influent temperature. The coefficient of variation (CV) is a statistic that measures the dispersion of data, defined as the ratio of the standard deviation to the mean of the sequence. The CV is dimensionless, facilitating comparisons of the degree of fluctuation between parameters with different dimensions. The CV for water quality characteristics is calculated for the predicted influent water quality characteristic sequence. The CV for influent water quantity is calculated for the predicted influent water quantity sequence. The CV for influent temperature is calculated for the predicted influent temperature sequence.
[0030] Specifically, for the predicted influent water quality characteristic sequence within the preset operating window obtained by S100, its mean and standard deviation are calculated. The standard deviation divided by the mean yields the coefficient of variation of the water quality characteristics. Similarly, the coefficients of variation are calculated for the predicted influent flow rate sequence and the predicted influent temperature sequence, respectively.
[0031] For example, with a preset operating window of the next 24 hours, the predicted COD sequence is [62, 65, 70, ..., 55] mg / L, the calculated mean is 60.5 mg / L, and the standard deviation is 5.2 mg / L. Then, the coefficient of variation of the water quality characteristics... The predicted water volume series has a mean of 125 m³. 3 / h, standard deviation is 8m 3 / h, coefficient of variation of water volume The predicted temperature series has a mean of 17.3℃, a standard deviation of 0.4℃, and a temperature variation coefficient of... .
[0032] Secondly, the coefficients of variation for water quality characteristics, influent flow rate, and influent temperature are weighted and fused to obtain the comprehensive variability, which serves as the prediction fluctuation coefficient for the recovery effect. Weighted fusion involves assigning different weights to each coefficient of variation based on their impact on the recovery effect, and then calculating a weighted sum. The comprehensive variability is a single indicator after weighted fusion, reflecting the degree of fluctuation in the overall influent characteristics within the prediction window. The prediction fluctuation coefficient, i.e., the comprehensive variability, is used to subsequently determine how many prediction branches to invoke.
[0033] Specifically, the weights w1 for the coefficient of variation of water quality characteristics, w2 for the coefficient of variation of water quantity, and w3 for the coefficient of variation of temperature are determined in advance through experience or correlation analysis. Generally, water quality has the greatest impact, followed by water quantity, and then temperature, and the sum of the three is 1. The predicted fluctuation coefficient = w1 × coefficient of variation of water quality characteristics + w2 × coefficient of variation of influent water quantity + w3 × coefficient of variation of influent temperature.
[0034] For example, the weights are set as follows: water quality characteristics have the greatest impact, so w1=0.6; water quantity has the second greatest impact, so w2=0.3; and temperature has a relatively small impact, so w3=0.1. The predicted fluctuation coefficient is then calculated. .
[0035] The training steps for the recycling effect predictor include: Collect multidimensional influent characteristic data sequences of several samples, as well as operating parameters of several samples of wastewater reuse, and obtain several historical actual recovery effect index sequences as recovery effect index sequences of several samples. Among them, the operating parameters include the dosage of chemicals and the operating control parameters, and the operating control parameters include at least membrane flux, backwashing frequency and aeration intensity. Several multidimensional influent feature data sequences of several samples, several sample operating parameters, and several historical actual recovery effect index sequences are used as training data, divided into J equal parts, and cross-selected with replacement to obtain J sample training sets. Using J sample training sets, the Long Short-Term Memory Network is trained under supervision until convergence, generating J recycling effect prediction branches, which are then combined to obtain the recycling effect predictor.
[0036] The sample recovery effect index sequence includes actual recovery effect indicators at several time points within a historical time window, including effluent quality, recovery rate, and energy consumption cost.
[0037] First, collect multidimensional influent characteristic data sequences of several samples, as well as operating parameters of several samples of wastewater reuse, and obtain several historical actual recovery effect index sequences as recovery effect index sequences of several samples. Among them, the operating parameters include the dosage of chemicals and the operating control parameters, which at least include membrane flux, backwashing frequency and aeration intensity.
[0038] The sample recovery effect index sequence includes actual recovery effect indicators at several time points within a historical time window, including effluent quality, recovery rate, and energy consumption cost.
[0039] The sample multidimensional influent characteristic data sequence refers to a multidimensional influent characteristic sequence with the same structure as that in S100, extracted from historical monitoring data, including water quality, water quantity, and temperature. Sample operating parameters include reagent dosage and operating control parameters, which at least include membrane flux, backwashing frequency, and aeration intensity; these are adjustable control variables during process operation. The sample recovery effect index sequence refers to the actual recovery effect index at each time point within the historical time window, including effluent water quality, recovery rate, and energy cost. Effluent water quality includes compliance status such as COD and ammonia nitrogen; recovery rate includes permeate production rate; and energy cost includes energy consumption per unit of permeate production.
[0040] Specifically, data from multiple time windows are extracted from the historical database of the wastewater reclamation plant. Each sample contains: a multidimensional influent characteristic data sequence within a historical time window, the actual operating parameters used within that window, and the corresponding sequence of actual recycling performance indicators within that window. All samples are then aggregated to form a training dataset.
[0041] For example, a sample is taken every 12 hours from data from the past two years. Each sample includes: COD, water volume, and temperature sequences for the past 24 hours; the average chemical dosage (e.g., 8 mg / L) and average membrane flux (e.g., 18 L / (m²)) for the past 24 hours. 2•h), average backwashing frequency (e.g., once every 2 hours), average aeration intensity (e.g., 1.2m). 3 / (m 2 •h). And the actual effluent COD, recovery rate, and energy cost sequence at each time point within that 24-hour period. A total of 2000 samples were collected.
[0042] Secondly, several multidimensional influent feature data sequences, several sample operating parameters, and several historical actual recovery effect index sequences are used as training data, divided into J equal parts, and cross-selected with replacement to obtain J sample training sets. Dividing into J equal parts means randomly shuffling all training data and uniformly dividing it into J non-overlapping subsets, each containing approximately the same number of samples. Cross-selection with replacement uses the Bagging approach, randomly selecting a number of subsets with replacement from the J parts each time to construct a sample training set. That is, for each training set, random sampling with replacement is performed from all data, with the number of samples equal to the original data size, resulting in J different training sets. Each training set has the same size as the original dataset, but with duplicate samples.
[0043] Specifically, let the total number of samples be N, and divide it into J equal parts, each with approximately N / J samples. Then, perform J sampling with replacement: each time, randomly select one sample from the J parts or draw N samples with replacement from the entire dataset, ultimately obtaining J training sets. Each training set is used to train an independent prediction branch.
[0044] For example, let J=10. Randomly shuffle the 2000 samples and divide them into 10 equal parts, each with 200 samples. Then perform 10 rounds of sampling with replacement: each time, randomly select 10 parts with replacement from the 10 parts, that is, draw 2000 samples. It is allowed for the same part to be drawn multiple times, resulting in 10 training sets, each with 2000 samples.
[0045] Next, using J training samples, the Long Short-Term Memory (LSTM) network is trained under supervised supervision until convergence, generating J recovery effect prediction branches. These branches are then combined to form the recovery effect predictor. Each recovery effect prediction branch is an independent LSTM model, whose input is a multi-dimensional influent feature data sequence and operating parameters, and whose output is a sequence of predicted recovery effect indicators. The recovery effect predictor is a set of J recovery effect prediction branches, from which some branches are subsequently selected for prediction based on the predicted fluctuation coefficient.
[0046] Specifically, an LSTM network structure similar to S100 but with a different output is constructed. The input layer receives multi-dimensional water inflow features and operating parameters. The operating parameters can be concatenated as additional features to each time step or as a global vector input. The output layer outputs a sequence of recovery performance indicators corresponding to a preset operating window. For the i-th sample training set (i=1…J), an LSTM network is independently trained using this training set. The loss function is the mean squared error (MSE), the optimizer is Adam, and the convergence condition is that the loss on the validation set no longer decreases. After training, the model parameters of this branch are saved. Finally, J trained branches are obtained, which together form the recovery performance predictor.
[0047] For example, a training set of J=10 is used, with each training set training an independent LSTM model. Each model structure consists of two LSTM layers: the first layer has 64 neurons, and the second layer has 32 neurons, both using the tanh activation function. This is followed by a fully connected layer that outputs three recovery performance metrics, predicting values for each time point within the next 24 hours (24×3=72 dimensions). Each training set is trained for 150 epochs, converging when the validation set loss drops below 0.003. After training the 10 branches, 10 model files are saved, collectively forming the recovery performance predictor.
[0048] Specifically, based on the predicted fluctuation coefficient, a pre-trained recycling effect predictor is invoked, and combined with the expected operating parameters of wastewater reuse, several predicted recycling effect index sequences within a preset operating window are predicted, including: The optimal number of branches H is obtained by multiplying the ratio of the predicted volatility coefficient to the maximum historical volatility coefficient within the historical time period by J and rounding up. H is less than or equal to J. If H is less than 3, then H is equal to 3. H recovery effect prediction branches are randomly selected from J recovery effect prediction branches. Based on the predicted multidimensional influent characteristic data sequence and expected operating parameters, H predicted recovery effect index sequences within the preset operating window are obtained as several predicted recovery effect index sequences.
[0049] First, the ratio of the predicted volatility coefficient to the maximum historical volatility coefficient within the historical time period is multiplied by J and rounded up to obtain the optimal number of branches H. H is less than or equal to J; if H is less than 3, it is set to 3. The maximum historical volatility coefficient refers to the maximum value among all comprehensive variability calculated from historical data over a relatively long period, used to normalize the current volatility coefficient. The optimal number of branches H is the number of prediction branches to be used, dynamically determined based on the current volatility intensity. The greater the volatility, the larger H is to enhance prediction diversity; the smaller the volatility, the smaller H is to improve computational efficiency. The lower limit of H is set to 3 to ensure at least some diversity.
[0050] Specifically, the maximum historical volatility coefficient CV_max is calculated from the historical database. The predicted volatility coefficient calculated in this step is denoted as CV_cur. (Ratio) Then the optimal number of branches. If H > J, then H = J; if H < 3, then H = 3.
[0051] For example, the historical maximum volatility coefficient Current predicted volatility coefficient ,but . Rounding up, we get H=3.
[0052] Secondly, H recovery effect prediction branches are randomly selected from the J recovery effect prediction branches. Based on the predicted multidimensional influent characteristic data sequence and expected operating parameters, H predicted recovery effect index sequences within the preset operating window are obtained, serving as several predicted recovery effect index sequences. Expected operating parameters refer to the planned operating parameters of the wastewater reuse process within the preset operating window, including reagent dosage, membrane flux, backwashing frequency, aeration intensity, etc., which can be preset by the process engineer or provided by the results of the previous optimization round. The predicted recovery effect index sequence refers to the recovery effect index sequence output by each selected prediction branch at each time point within the preset operating window.
[0053] Specifically, from the J pre-trained recovery effect prediction branches, H branches are randomly selected without replacement. For each selected branch, the predicted multidimensional inflow feature data sequence obtained from S100 and the given expected operating parameters are input into that branch, forward inference is performed, and the corresponding predicted recovery effect index sequence is output. The outputs of the H branches are saved as H different prediction sequences. Due to the different training data of the branches, the different prediction sequences will exhibit certain differences, reflecting the prediction uncertainty caused by inflow fluctuations.
[0054] For example, given J=10 and H=3, three recovery effect prediction branches 2, 5, and 8 are randomly selected. The expected operating parameters are set as follows: reagent dosage 9 mg / L, membrane flux 20 L / (m²·h), backwashing frequency once every 1.5 hours, and aeration intensity 1.0 m³ / (m²·h). The predicted 24-hour COD, water volume, and temperature sequences from S100 are concatenated with the above expected operating parameters and input into recovery effect prediction branches 2, 5, and 8 respectively. Branch 2 outputs the predicted recovery effect sequence: effluent COD [8,9,10,…] mg / L, recovery rate [78%,77%,75%,…], energy cost [0.45,0.46,0.48,…] yuan / m³. 3 Branches 5 and 8 output slightly different sequences. Finally, three sequences of predicted recovery performance indicators are obtained, which serve as inputs for step S300.
[0055] In this embodiment of the invention, firstly, a volatility analysis is performed on the predicted multidimensional influent characteristics to quantify the degree of variation in influent water quality, quantity, and temperature within the future window, and a comprehensive predicted volatility coefficient is obtained through weighted fusion. Based on this, multiple recovery effect prediction branches are trained using the Bagging method to construct an integrated recovery effect predictor. The number of branches invoked is dynamically determined according to the predicted volatility coefficient, achieving an adaptive balance between prediction accuracy and computational cost. Finally, combined with the expected operating parameters, multiple sequences of predicted recovery effect indicators with differences are obtained, providing rich probabilistic prediction results for subsequent sliding window anomaly detection. This avoids the systematic bias that a single prediction model may produce when influent fluctuations are severe, improving the robustness and accuracy of anomaly node identification.
[0056] S300: Perform sliding window anomaly detection on each predicted recovery effect index sequence. Based on the deviation between the predicted value and the theoretical expected value of each node, identify the abnormal running nodes within the preset running window and determine the abnormal running segment sequence.
[0057] In this embodiment of the invention, a sliding window anomaly detection is performed on each predicted recovery performance index sequence. Based on the deviation between the predicted value and the theoretical expected value of each node, abnormal operating nodes within a preset operating window are identified, and the abnormal operating segment sequence is determined. After obtaining multiple predicted recovery performance index sequences, it is necessary to identify the specific time periods within the preset operating window where performance degradation may occur. Traditional anomaly detection methods typically set a fixed threshold for the predicted value, but ignore the autocorrelation and dynamic evolution of the recovery performance index over time. This invention employs a sliding window mechanism, using the predicted values of L consecutive nodes preceding the current node, and calculates the theoretical expected value of the node using a specially trained index theoretical predictor. Then, the original predicted value is compared with the theoretical expected value, and anomalies are determined based on the magnitude of the deviation. This method can effectively capture local abnormal fluctuations in the predicted sequence, avoiding false alarms or missed alarms caused by improper global threshold settings. Furthermore, by statistically analyzing the time periods and duration proportions of consecutive abnormal nodes, truly engineering-significant abnormal operating segments are screened out, thereby providing accurate time positioning for subsequent optimization and control.
[0058] Step S300 in the method provided in this embodiment of the invention includes: The first predicted recovery performance index sequence is randomly selected from several predicted recovery performance index sequences; A sliding window anomaly detection is performed on the first predicted recovery effect index sequence. Abnormal running nodes are identified based on the deviation between the predicted value and the theoretical expected value of each node, and the first initial abnormal running segment sequence is determined. Traverse the remaining sequences that are not selected from several predicted recovery effect index sequences, perform the same sliding window anomaly detection on each sequence, determine the corresponding initial anomaly runtime sequence, and obtain several initial anomaly runtime sequence sequences. Merge several initial abnormal runtime segment sequences to generate the final abnormal runtime segment sequence.
[0059] First, a first predicted recovery performance index sequence is randomly selected from several predicted recovery performance index sequences. The first predicted recovery performance index sequence refers to one of the H predicted recovery performance index sequences randomly selected as the current sequence to be detected. Subsequent steps will perform sliding window anomaly detection on this sequence, and so on for the remaining sequences. Let the H predicted recovery performance index sequences be a set {Seq1, Seq2, ..., Seq_H}. A sequence is randomly selected from this set, denoted as Seq_first. For example, H=3, and the three sequences are denoted as Seq_A, Seq_B, and Seq_C respectively. Seq_B is randomly selected as the first predicted recovery performance index sequence.
[0060] Secondly, a sliding window anomaly detection is performed on the first predicted recovery effect index sequence. Based on the deviation between the predicted value and the theoretical expected value of each node, abnormal running nodes are identified, and the first initial abnormal running segment sequence is determined.
[0061] Specifically, a sliding window anomaly detection is performed on the first predicted recovery effect index sequence. Based on the deviation between the predicted value and the theoretical expected value of each node, abnormal running nodes are identified, and the first initial abnormal running segment sequence is determined, including: The first L consecutive data points of the current node are used as the reference sliding window, where L is not less than 5; Based on the predicted recovery performance index segment within the baseline sliding window, the theoretical expected value of the recovery performance index of the current node is predicted, and the relative deviation between the predicted value and the theoretical expected value of the current node is calculated. If the relative deviation exceeds the preset deviation threshold, the current node is marked as an abnormal running node. Traverse all nodes within the preset running window according to the baseline sliding window, and output the first abnormal running node set of the first predicted recovery effect index sequence within the preset running window. Obtain the time period consisting of all consecutive abnormal running nodes in the first abnormal running node set, and calculate the proportion of the time span of each consecutive abnormal running node time period to the total duration of the preset running window. When the proportion exceeds the preset abnormal proportion threshold, the consecutive abnormal running node time period is determined as an initial abnormal running segment. Summarize all consecutive time periods that meet the preset abnormal proportion threshold, and output the first initial abnormal running segment sequence corresponding to the first predicted recovery effect index sequence.
[0062] First, the L consecutive data points preceding the current node are used as the baseline sliding window, where L is not less than 5. The current node refers to a specific time point within the preset operating window, such as the t-th time point. The baseline sliding window is the input data window used to predict the theoretical expected value of the current node, and it consists of the predicted recovery performance index data of the L consecutive nodes preceding the current node. L is the window length, which is at least 5 to ensure statistical significance.
[0063] Specifically, for a time node t within the preset operating window, t≥L+1, starting from the (L+1)th node, the predicted recovery effect index values of the preceding L nodes (tL, t-L+1, ..., t-1) are taken to form a sequence segment of length L, which serves as the baseline sliding window. For example, let L=5. The preset operating window is 24 hours, with node numbers 1-24. When the current node is t=6, the baseline sliding window includes the predicted recovery effect indexes of nodes 1-5, with each node containing values for three dimensions: effluent quality, recovery rate, and energy consumption cost. When t=7, the window slides to nodes 2-6, and so on.
[0064] Secondly, based on the predicted recovery effect index segment within the baseline sliding window, the theoretical expected value of the recovery effect index of the current node is predicted, and the relative deviation between the predicted value and the theoretical expected value of the current node is calculated. If the relative deviation exceeds the preset deviation threshold, the current node is marked as an abnormal running node.
[0065] Specifically, based on the predicted recovery performance index segment within the baseline sliding window, the theoretical expected value of the recovery performance index for the current node is predicted, and the relative deviation between the predicted value and the theoretical expected value for the current node is calculated. If the relative deviation exceeds a preset deviation threshold, the current node is marked as an abnormal operating node, including: Using the node span of the baseline sliding window as a constraint, several sample recovery effect index segments are collected, and the historical actual recovery effect index of the next adjacent node of different sample recovery effect index segments is collected as the theoretical expected value of the sample, thus obtaining several sample theoretical expected values. Using several sample recovery performance index fragments and several sample theoretical expected values, a long short-term memory network is trained until convergence to generate an index theoretical predictor; Using an index theory predictor, the theoretical expected value of the recovery performance index for the current node is predicted based on the predicted recovery performance index segment within the baseline sliding window. Calculate the relative deviation between the predicted value and the theoretical expected value of the current node to obtain the deviation of effluent water quality characteristics, recovery rate deviation and energy consumption cost deviation; If at least two of the following deviations—the deviation in effluent water quality characteristics, the deviation in recovery rate, and the deviation in energy consumption cost—exceed their respective preset deviation thresholds, the current node will be marked as an abnormal operating node.
[0066] First, constrained by the node span of the baseline sliding window, several sample recovery performance index segments are collected. The historical actual recovery performance index of the next adjacent node for each sample recovery performance index segment is then collected as the theoretical expected value of the sample, resulting in several theoretical expected values. The node span refers to the length L of the baseline sliding window, meaning each segment contains data from L consecutive nodes. A sample recovery performance index segment is a sequence of historical actual recovery performance indices for L consecutive nodes extracted from historical operational data, following the same node span L. The next adjacent node refers to the actual recovery performance index of the next time node immediately following the sample recovery performance index segment. The theoretical expected value of the sample serves as the label for supervised learning, i.e., the historical actual recovery performance index of the next adjacent node corresponding to that segment.
[0067] Specifically, from the historical operational database, long-term series actual recovery performance index data, which are from the same source as the training data of the recovery performance predictor, are extracted. All continuous segments of length L are truncated with a step size of 1, and each segment corresponds to the actual value of the next node, forming a training sample pair: (segment, actual value of the next node). The actual value of the next node is used as the theoretical expected value of the sample.
[0068] For example, the historical actual recovery performance data includes hourly data from the past year, totaling 8760 nodes. Taking L=5, the 6th node is predicted from nodes 1-5, the 7th node is predicted from nodes 2-6, and so on, resulting in approximately 8754 sample segments. Each segment is a matrix of 5 time nodes × 3 indicators, and the corresponding theoretical expected value of the sample is a three-dimensional vector of the actual effluent COD, recovery rate, and energy consumption cost at the 6th node.
[0069] Secondly, using several sample recovery performance index segments and several sample theoretical expected values, a Long Short-Term Memory (LSTM) network is trained until convergence to generate a theoretical predictor of the performance index. The theoretical predictor is built based on an LSTM model, taking a recovery performance index segment of length L as input and outputting the theoretical expected value of the recovery performance index for the next node of that segment. The theoretical predictor is used to simulate the temporal evolution of the recovery performance index under normal operating conditions.
[0070] Specifically, an LSTM network is constructed with an input dimension of L×3, L time steps, and three metric metrics per step. The output dimension is 3, corresponding to the three metrics of the next node. One to two hidden layers are used, and the number of neurons is adjusted based on the amount of data, for example, 32 or 64. Mean squared error (MSE) is used as the loss function, and the Adam optimizer is employed. Supervised training is performed using sample segments as input and the corresponding theoretical expected values of the samples as labels. Training stops when the loss on the validation set no longer decreases, the model parameters are saved, and the theoretical predictor of the metrics is obtained.
[0071] For example, using 8754 collected samples, the dataset was divided into training, validation, and test sets at 80%, 10%, and 10% respectively. A single-layer LSTM was constructed with 64 hidden neurons, an activation function of tanh, and a fully connected linear output layer. After 150 training epochs, the validation set MSE dropped below 0.001, indicating model convergence, and the model was saved as a theoretical predictor of metrics.
[0072] Secondly, using the index theory predictor, based on the predicted recovery performance index segment within the baseline sliding window, the theoretical expected value of the recovery performance index for the current node is predicted. The theoretical expected value refers to the recovery performance index value that should appear at the next node under the historical evolution pattern of the current baseline sliding window, reflecting the expected performance under normal operating conditions.
[0073] Specifically, for the current node t, the predicted recovery performance index values of the previous L nodes are taken to form an input tensor of shape [1,L,3]. This tensor is then input into the trained index theory predictor, and the model forward calculates and outputs a three-dimensional vector, which represents the theoretical expected value of the effluent water quality, the theoretical expected value of the recovery rate, and the theoretical expected value of the energy consumption cost for the current node t.
[0074] For example, the current node t=6, and the baseline sliding window is the predicted recovery performance indicators for nodes 1-5. Inputting this 5×3 matrix into the indicator theory predictor yields the following outputs: theoretical expected effluent COD = 8.2 mg / L, theoretical expected recovery rate = 80.5%, and theoretical expected energy cost = 0.42 yuan / m³. 3 .
[0075] Furthermore, the relative deviation between the predicted value and the theoretical expected value of the current node is calculated to obtain the deviation of effluent water quality characteristics, recovery rate deviation, and energy consumption cost deviation. The predicted value refers to the recovery effect index value of the current node directly output by the recovery effect prediction branch in S200. The relative deviation - |predicted value - theoretical expected value| / theoretical expected value is used to eliminate the influence of different index dimensions. For indices close to zero, a small constant can be added to prevent division by zero.
[0076] Specifically, for the current node t, obtain its predicted value in the first predicted recovery effect index sequence (denoted as P_water quality, P_recovery, P_energy consumption), and its theoretical expected value (denoted as E_water quality, E_recovery, E_energy consumption). Calculate the following respectively: effluent water quality characteristic deviation = |P_water quality - E_water quality| / E_water quality; recovery rate deviation = |P_recovery - E_recovery| / E_recovery; energy cost deviation = |P_energy consumption - E_energy consumption| / E_energy consumption.
[0077] For example, the predicted values for node 6 are: effluent COD = 12 mg / L, recovery rate = 75%, and energy cost = 0.55 yuan / m³. 3Theoretical expected values: effluent COD = 8.2 mg / L, recovery rate = 80.5%, energy cost = 0.42 yuan / m³ 3 Therefore: the effluent water quality deviates. Recovery rate deviation Energy consumption deviation .
[0078] Subsequently, if at least two of the deviations in effluent water quality characteristics, recovery rate, and energy cost exceed their respective preset deviation thresholds, the current node is marked as an abnormal operating node. The preset deviation thresholds refer to the upper limits of allowable relative deviations set for effluent water quality, recovery rate, and energy cost, respectively. Exceeding the threshold indicates that the indicator is abnormal at that node. The thresholds can be obtained through statistical analysis of the deviation distribution under historical normal operating conditions, such as by taking the 95th percentile. An abnormal operating node is a time point determined to be where the recovery effect significantly deviates from normal expectations, indicating that there may be process problems or influent shocks at that moment.
[0079] Specifically, three thresholds are set: Th_water quality, Th_recovery, and Th_energy consumption, for example, 0.2, 0.1, and 0.25 respectively. The three deviations of the current node are compared with their corresponding thresholds, and the number of deviations exceeding the thresholds is counted. If the number of deviations exceeds the threshold by ≥2, the current node is marked as an abnormal operating node; otherwise, it is a normal node.
[0080] For example, let the preset deviation thresholds be: effluent water quality 0.2, recovery rate 0.1, and energy cost 0.25. The deviation of node 6 is: water quality >0.2, recovery rate <0.1, energy consumption >0.25. Two indicators exceeded the threshold: water quality and energy consumption. Since at least two of them were met, node 6 was marked as an abnormal operating node.
[0081] Next, traverse all nodes within the preset running window according to the baseline sliding window, and output the first abnormal running node set within the preset running window for the first predicted recovery effect index sequence. Traversal refers to all nodes within the preset running window that meet the sliding window start condition. The first abnormal running node set refers to the set that records the positions of all nodes marked as abnormal running nodes. Initialize an empty set Set_abnormal. For nodes t=L+1 to T, where T is the total number of nodes in the preset running window, perform the above steps. If it is determined to be abnormal, add t to the set. After the traversal is completed, output Set_abnormal as the first abnormal running node set.
[0082] For example, for node 6, L=5, T=24, traverse nodes 6 to 24, a total of 19 nodes. After detection, nodes 6, 9, 10, 11, 12, and 18 are marked as abnormal, so the first abnormal running node set is {6,9,10,11,12,18}.
[0083] Furthermore, the time period consisting of all consecutive abnormal running nodes in the first abnormal running node set is obtained, and the proportion of the time span of each consecutive abnormal running node time period to the total duration of the preset running window is calculated. When the proportion exceeds the preset abnormal proportion threshold, the consecutive abnormal running node time period is determined as an initial abnormal running segment. All consecutive time periods that meet the preset abnormal proportion threshold are summarized, and the first initial abnormal running segment sequence corresponding to the first predicted recovery effect index sequence is output.
[0084] In this context, consecutive abnormal running nodes refer to a sequence of nodes that are adjacent on the time axis and all belong to abnormal running nodes. The time span of a time period refers to the start and end time difference of consecutive abnormal nodes; for example, from node a to node b, the span is (b-a+1) × sampling interval. The preset abnormality proportion threshold refers to the minimum proportion of a consecutive abnormal time period to the total duration of the entire preset running window. Only time periods exceeding this proportion are considered meaningful abnormal running segments, avoiding interference from scattered single-point abnormalities. Initial abnormal running segments are consecutive time periods that meet the above proportion conditions, and each time period is defined by a start node and an end node. The first initial abnormal running segment sequence refers to the set of all initial abnormal running segments detected from the first predicted recovery effect index sequence, which may contain multiple discontinuous time periods.
[0085] Specifically, the nodes in the first set of abnormal running nodes are sorted by time to identify all continuous segments. For each continuous segment, the time span is calculated as (number of nodes within the segment × sampling interval). The preset total running window duration is calculated as T × sampling interval. The proportion is calculated as time span / total duration. If the proportion is greater than the preset abnormal percentage threshold, the continuous segment is identified as an initial abnormal running segment, and its start and end nodes are recorded. All time periods that meet the conditions are summarized, and the sequence of the first initial abnormal running segments is output.
[0086] For example, the sampling interval is 1 hour, and the preset total running window duration is 24 hours. In the first abnormal running node set {6,9,10,11,12,18}, the continuous segments are: [6]: single point, spanning 1 hour, accounting for 1 / 24≈4.2%; [9,10,11,12]: 4 hours, accounting for 4 / 24≈16.7%;
[18] : 1 hour, accounting for 4.2%. Assuming the preset abnormal percentage threshold is 10%, only the continuous segment [9,10,11,12] meets the condition and is determined to be an initial abnormal running segment (starting node 9, ending node 12). Therefore, the first initial abnormal running segment sequence = {(9,12)}.
[0087] Next, iterate through the remaining sequences that were not selected from the predicted recovery performance index sequences, and perform the same sliding window anomaly detection on each to determine their corresponding initial anomaly runtime segment sequences, resulting in several initial anomaly runtime segment sequences. The remaining sequences refer to the H-1 sequences other than the randomly selected first predicted recovery performance index sequence. For each remaining predicted recovery performance index sequence, repeat the above steps to obtain the initial anomaly runtime segment sequence corresponding to each sequence. Finally, H initial anomaly runtime segment sequences are obtained.
[0088] For example, H=3, the first predicted recovery performance index sequence is Seq_B, and its initial abnormal runtime sequence is {(9,12)}. Repeating S302~S309 for Seq_A yields its initial abnormal runtime sequence {(8,11),(20,22)}. Repeating this process for Seq_C yields its initial abnormal runtime sequence {(10,13)}. Therefore, a total of 3 initial abnormal runtime sequences are obtained.
[0089] Finally, several initial abnormal runtime segment sequences are merged to generate the final abnormal runtime segment sequence. The time segments from all initial abnormal runtime segment sequences are combined, and overlapping or adjacent time segments are merged to form the final abnormal runtime segment sequence. All time segments from H initial abnormal runtime segment sequences are collected, merged into one set, and then fused: if two time segments overlap or their interval is less than or equal to a certain tolerance level (e.g., one sampling interval), they are merged into a larger continuous time segment. The final output abnormal runtime segment sequence is the set of time intervals within the preset operating window that require key attention and optimization.
[0090] For example, the time periods of the three sequences are combined: The final abnormal runtime sequence is {(8,13),(20,22)}, indicating that the process is very likely to experience abnormal recovery effects during the 8th to 13th hour and the 20th to 22nd hour of the preset running window, requiring subsequent optimization and control.
[0091] In this embodiment of the invention, a sliding window and an index theory predictor are used to generate theoretical expected values for each node based on historical normal evolution patterns. By comparing the deviation with the original predicted values, local abnormal fluctuations in the prediction sequence can be sensitively captured, avoiding the limitations of a single fixed threshold method. By requiring at least two indicators to exceed the standard simultaneously to determine an abnormal node, the false alarm rate is reduced. Furthermore, by filtering based on the duration ratio of consecutive abnormal nodes, scattered noise is eliminated, focusing on continuous abnormal periods with practical process intervention value. Finally, the abnormal period results from multiple prediction branches are merged, enhancing the robustness of anomaly identification and enabling the early identification of high-risk periods within future operating windows where water quality may exceed standards, recovery rates may decrease, or energy consumption may increase, providing clear time targets for precise optimization and control of the S400 step.
[0092] S400: The abnormal runtime sequence is taken as the time period to be optimized within the preset runtime window. During the wastewater reuse process, the wastewater reuse process is optimized and controlled in combination with the preset monitoring-feedback mechanism.
[0093] In this embodiment of the invention, abnormal operating time sequences are designated as optimization periods within a preset operating window. During wastewater reuse treatment, a preset monitoring-feedback mechanism is used to optimize and control the wastewater reuse process within these optimization periods. In actual wastewater reuse treatment operation, influent water quality, quantity, temperature, and other conditions often exhibit periodic or non-periodic fluctuations, leading to mismatches between process parameters and real-time operating conditions during certain periods. This results in problems such as accelerated membrane fouling, decreased coagulation efficiency, and excessive or insufficient aeration. Applying a uniform control strategy to all periods not only increases operating energy consumption and reagent costs but may also cause unnecessary disturbances during stable periods. Therefore, targeted optimization only for identified abnormal operating time sequences can avoid ineffective control, concentrate resources to address performance shortcomings during critical periods, and improve the targeting and economy of process control.
[0094] Specifically, within the preset operating window, based on the determined abnormal operating period sequence, real-time monitoring data such as influent flow rate, turbidity, COD, ammonia nitrogen, membrane pressure difference, and dissolved oxygen are collected during that period. Parameter adaptive adjustment is performed through a preset monitoring-feedback mechanism. This includes: establishing correlation rules between monitoring data and control parameters; triggering corresponding actuators to adjust when monitored values deviate from set thresholds or exceed the rate of change; implementing refined control for different process units: the coagulation unit dynamically adjusts the coagulant dosage based on changes in influent turbidity and colloidal charge; the membrane separation unit adjusts the membrane flux setpoint or increases the backwashing frequency based on the increase in transmembrane pressure difference and flux decay rate; the biochemical unit optimizes aeration intensity and intermittent aeration cycle based on dissolved oxygen and load fluctuations; continuous closed-loop feedback is maintained during the control process, with monitoring values rereading after each adjustment. If the values still exceed the tolerance range, step-by-step fine-tuning is performed until the indicators return to the normal range. For non-abnormal periods, the original operating parameters remain unchanged, and no additional control is performed.
[0095] For example, the preset operating window is 24 hours, and the identified abnormal operating sequence sequences are (8,13) and (20,22). During the period from 08:00 to 13:00, which coincides with the morning peak and the continuous midday influent load period, the online turbidity meter shows that the influent turbidity gradually increases from the normal 5 NTU to 15 NTU, and the COD concentration increases from 40 mg / L to 85 mg / L. The monitoring-feedback mechanism automatically calculates the PAC dosage coefficient and dynamically adjusts the coagulant dosage from 2.5 mg / L to 3.8 mg / L; at the same time, the rate of increase of transmembrane pressure difference increases from 0.1 kPa / min to 0.25 kPa / min, the backwashing frequency of the ultrafiltration membrane is increased from once every 30 minutes to once every 15 minutes, and the aeration intensity is increased from 35 m³ / min according to the decreasing trend of dissolved oxygen. 3 / h increased to 55m 3 / h. During the period from 20:00 to 22:00, which is the second peak of evening water consumption, the turbidity of the influent rises again to 12 NTU. The PAC dosage is automatically adjusted to 3.2 mg / L, and the backwashing frequency is adjusted to once every 20 minutes. No adjustments are made during other periods, thereby saving energy and chemicals.
[0096] In this embodiment of the invention, by initiating monitoring-feedback optimization and control within the abnormal operating period sequence, process parameters such as coagulant dosage, aeration intensity, and backwashing frequency can always match the real-time operating conditions, thereby improving the effluent quality compliance rate; at the same time, unnecessary adjustments are avoided during non-abnormal periods, effectively reducing power consumption and reagent costs during operation.
[0097] Through the specific implementation methods described above, the embodiments of the present invention achieve the following technical effects: This invention provides a method and system for optimizing wastewater reuse processes based on recovery effect prediction. First, it predicts a multi-dimensional influent characteristic sequence within a preset operating window based on historical data. Then, it calls a recovery effect predictor to obtain multiple recovery effect index sequences. Next, it accurately identifies abnormal operating nodes and determines abnormal time periods through sliding window anomaly detection. Finally, it implements targeted optimization and control only during abnormal time periods, combining a monitoring-feedback mechanism. This achieves a shift from passive response to proactive prediction and precise intervention, improving the accuracy of recovery effect prediction and the sensitivity of anomaly identification. By guiding subsequent control through pre-prediction, it avoids blind adjustment and all-time intervention, effectively reducing chemical and energy consumption and backwashing frequency. Simultaneously, it enhances the adaptability to influent fluctuations and operational deviations, ensuring stable effluent quality and comprehensively improving the economy, stability, and intelligence of the wastewater reuse process.
[0098] Example 2, as Figure 2 As shown, this invention provides a wastewater reuse process optimization system based on recycling effect prediction, the system comprising: The influent feature prediction module 11 is used to predict and obtain a sequence of predicted multidimensional influent feature data within a preset operating window based on historical wastewater reuse operation monitoring data. The recycling effect prediction module 12 is used to perform a volatility analysis of the recycling effect prediction based on the predicted multidimensional influent characteristic data sequence. Based on the prediction volatility coefficient, it calls the pre-trained recycling effect predictor and, combined with the expected operating parameters of wastewater reuse, predicts a series of predicted recycling effect indicators within a preset operating window. The abnormal period identification module 13 is used to perform sliding window anomaly detection on each predicted recovery effect index sequence. Based on the deviation between the predicted value and the theoretical expected value of each node, it identifies the abnormal running nodes within the preset running window and determines the abnormal running period sequence. The process optimization and control module 14 is used to take the abnormal running segment sequence as the time period to be optimized within the preset running window. During the wastewater reuse process, it combines the preset monitoring-feedback mechanism to perform optimization and control of the wastewater reuse process within the time period to be optimized.
[0099] In one embodiment, the influent feature prediction module 11 is further configured to: Based on historical wastewater reuse operation monitoring data, historical multidimensional influent characteristic data sequences are collected within the historical operation window. These historical multidimensional influent characteristic data sequences include historical influent water quality characteristic sequences, historical influent water volume sequences, and historical influent temperature sequences. The historical operation window and the preset operation window have the same time span. A long short-term memory network was trained until convergence by collecting a sample training set, and a multi-dimensional water inflow feature time series prediction model was generated. The historical influent water quality characteristic sequence, historical influent water volume sequence, and historical influent temperature sequence are input into the multidimensional influent characteristic time series prediction model to obtain the predicted multidimensional influent characteristic data sequence within the preset operating window.
[0100] In one embodiment, the recycling effect prediction module 12 is further configured to: Among them, the volatility analysis of recovery effect prediction based on the predicted multidimensional influent characteristic data sequence includes: Parametric fluctuation analysis was performed on the predicted influent water quality characteristic sequence, predicted influent water quantity sequence, and predicted influent temperature sequence in the predicted multidimensional influent characteristic data sequence, and the coefficient of variation of water quality characteristic, influent water quantity, and influent temperature were calculated. The coefficients of variation of water quality characteristics, influent flow rate, and influent temperature are weighted and fused to obtain the comprehensive variability, which is used as the predictive fluctuation coefficient for predicting the recovery effect.
[0101] The training steps for the recycling effect predictor include: Collect multidimensional influent characteristic data sequences of several samples, as well as operating parameters of several samples of wastewater reuse, and obtain several historical actual recovery effect index sequences as recovery effect index sequences of several samples. Among them, the operating parameters include the dosage of chemicals and the operating control parameters, and the operating control parameters include at least membrane flux, backwashing frequency and aeration intensity. Several multidimensional influent feature data sequences of several samples, several sample operating parameters, and several historical actual recovery effect index sequences are used as training data, divided into J equal parts, and cross-selected with replacement to obtain J sample training sets. Using J sample training sets, the Long Short-Term Memory Network is trained under supervision until convergence, generating J recycling effect prediction branches, which are then combined to obtain the recycling effect predictor.
[0102] The sample recovery effect index sequence includes actual recovery effect indicators at several time points within a historical time window, including effluent quality, recovery rate, and energy consumption cost.
[0103] Specifically, based on the predicted fluctuation coefficient, a pre-trained recycling effect predictor is invoked, and combined with the expected operating parameters of wastewater reuse, several predicted recycling effect index sequences within a preset operating window are predicted, including: The optimal number of branches H is obtained by multiplying the ratio of the predicted volatility coefficient to the maximum historical volatility coefficient within the historical time period by J and rounding up. H is less than or equal to J. If H is less than 3, then H is equal to 3. H recovery effect prediction branches are randomly selected from J recovery effect prediction branches. Based on the predicted multidimensional influent characteristic data sequence and expected operating parameters, H predicted recovery effect index sequences within the preset operating window are obtained as several predicted recovery effect index sequences.
[0104] In one embodiment, the abnormal time period identification module 13 is further configured to: The first predicted recovery performance index sequence is randomly selected from several predicted recovery performance index sequences; A sliding window anomaly detection is performed on the first predicted recovery effect index sequence. Abnormal running nodes are identified based on the deviation between the predicted value and the theoretical expected value of each node, and the first initial abnormal running segment sequence is determined. Traverse the remaining sequences that are not selected from several predicted recovery effect index sequences, perform the same sliding window anomaly detection on each sequence, determine the corresponding initial anomaly runtime sequence, and obtain several initial anomaly runtime sequence sequences. Merge several initial abnormal runtime segment sequences to generate the final abnormal runtime segment sequence.
[0105] Specifically, a sliding window anomaly detection is performed on the first predicted recovery effect index sequence. Based on the deviation between the predicted value and the theoretical expected value of each node, abnormal running nodes are identified, and the first initial abnormal running segment sequence is determined, including: The first L consecutive data points of the current node are used as the reference sliding window, where L is not less than 5; Based on the predicted recovery performance index segment within the baseline sliding window, the theoretical expected value of the recovery performance index of the current node is predicted, and the relative deviation between the predicted value and the theoretical expected value of the current node is calculated. If the relative deviation exceeds the preset deviation threshold, the current node is marked as an abnormal running node. Traverse all nodes within the preset running window according to the baseline sliding window, and output the first abnormal running node set of the first predicted recovery effect index sequence within the preset running window. Obtain the time period consisting of all consecutive abnormal running nodes in the first abnormal running node set, and calculate the proportion of the time span of each consecutive abnormal running node time period to the total duration of the preset running window. When the proportion exceeds the preset abnormal proportion threshold, the consecutive abnormal running node time period is determined as an initial abnormal running segment. Summarize all consecutive time periods that meet the preset abnormal proportion threshold, and output the first initial abnormal running segment sequence corresponding to the first predicted recovery effect index sequence.
[0106] Specifically, based on the predicted recovery performance index segment within the baseline sliding window, the theoretical expected value of the recovery performance index for the current node is predicted, and the relative deviation between the predicted value and the theoretical expected value for the current node is calculated. If the relative deviation exceeds a preset deviation threshold, the current node is marked as an abnormal operating node, including: Using the node span of the baseline sliding window as a constraint, several sample recovery effect index segments are collected, and the historical actual recovery effect index of the next adjacent node of different sample recovery effect index segments is collected as the theoretical expected value of the sample, thus obtaining several sample theoretical expected values. Using several sample recovery performance index fragments and several sample theoretical expected values, a long short-term memory network is trained until convergence to generate an index theoretical predictor; Using an index theory predictor, the theoretical expected value of the recovery performance index for the current node is predicted based on the predicted recovery performance index segment within the baseline sliding window. Calculate the relative deviation between the predicted value and the theoretical expected value of the current node to obtain the deviation of effluent water quality characteristics, recovery rate deviation and energy consumption cost deviation; If at least two of the following deviations—the deviation in effluent water quality characteristics, the deviation in recovery rate, and the deviation in energy consumption cost—exceed their respective preset deviation thresholds, the current node will be marked as an abnormal operating node.
[0107] It should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for optimizing wastewater reuse processes based on predicting recycling effects, characterized in that, The methods include: Based on historical wastewater reuse operation monitoring data, predictive multidimensional influent characteristic data sequences are obtained within a preset operation window; Based on the predicted multidimensional influent characteristic data sequence, the fluctuation analysis of the predicted recycling effect is performed. Based on the predicted fluctuation coefficient, the pre-trained recycling effect predictor is called. Combined with the expected operating parameters of wastewater reuse, several predicted recycling effect index sequences within the preset operating window are predicted. Sliding window anomaly detection is performed on each of the predicted recovery effect index sequences. Based on the deviation between the predicted value and the theoretical expected value of each node, abnormal running nodes within the preset running window are identified, and abnormal running segment sequences are determined. The abnormal runtime sequence is taken as the time period to be optimized within the preset runtime window. During the wastewater reuse process, the wastewater reuse process is optimized and controlled in the time period to be optimized by combining the preset monitoring-feedback mechanism.
2. The wastewater reuse process optimization method based on recycling effect prediction according to claim 1, characterized in that, Based on historical wastewater reuse operation monitoring data, a predicted multidimensional influent characteristic data sequence is obtained within a preset operation window, including: Based on historical wastewater reuse operation monitoring data, a historical multidimensional influent characteristic data sequence is collected within the historical operation window. The historical multidimensional influent characteristic data sequence includes a historical influent water quality characteristic sequence, a historical influent water volume sequence, and a historical influent temperature sequence. The historical operation window and the preset operation window have the same time span. A long short-term memory network was trained until convergence by collecting a sample training set, and a multi-dimensional water inflow feature time series prediction model was generated. The historical influent water quality characteristic sequence, historical influent water volume sequence, and historical influent temperature sequence are input into the multidimensional influent characteristic time series prediction model to obtain the predicted multidimensional influent characteristic data sequence within the preset operating window.
3. The wastewater reuse process optimization method based on recycling effect prediction according to claim 1, characterized in that, Based on the predicted multidimensional influent characteristic data sequence, a volatility analysis of the recovery effect prediction is performed, including: Parameter fluctuation analysis was performed on the predicted influent water quality characteristic sequence, the predicted influent water quantity sequence, and the predicted influent temperature sequence in the predicted multidimensional influent characteristic data sequence, and the coefficient of variation of water quality characteristic, influent water quantity, and influent temperature were calculated. The coefficients of variation of the water quality characteristics, the coefficients of variation of the influent flow rate, and the coefficients of variation of the influent temperature are weighted and fused to obtain the comprehensive variability, which is used as the prediction fluctuation coefficient for predicting the recovery effect.
4. The wastewater reuse process optimization method based on recycling effect prediction according to claim 1, characterized in that, The training steps for the recycling effect predictor include: Collect multidimensional influent characteristic data sequences of several samples, as well as operating parameters of several samples of wastewater reuse, and obtain several historical actual recovery effect index sequences as recovery effect index sequences of several samples. Among them, the operating parameters include the dosage of reagents and the operating control parameters, which include at least membrane flux, backwashing frequency and aeration intensity. The multidimensional influent feature data sequence of several samples, the operating parameters of several samples, and the historical actual recovery effect index sequence are used as training data, divided into J equal parts, and cross-selected with replacement to obtain J sample training sets. Using the J sample training sets, supervised training of the Long Short-Term Memory Network is performed until convergence, generating J recycling effect prediction branches, which are then combined to obtain the recycling effect predictor.
5. The wastewater reuse process optimization method based on recycling effect prediction according to claim 4, characterized in that, The sample recovery effect index sequence includes actual recovery effect indicators at several time points within a historical time window, including effluent quality, recovery rate, and energy consumption cost.
6. The wastewater reuse process optimization method based on recycling effect prediction according to claim 4, characterized in that, Based on the predicted fluctuation coefficient, a pre-trained recycling effect predictor is invoked. Combined with the expected operating parameters for wastewater reuse, several predicted recycling effect index sequences within the preset operating window are predicted, including: The optimal number of branches H is obtained by multiplying the ratio of the predicted volatility coefficient to the maximum historical volatility coefficient within the historical time period by J and rounding up, where H is less than or equal to J, and if H is less than 3, then H is equal to 3. H recovery effect prediction branches are randomly selected from the J recovery effect prediction branches. Based on the predicted multidimensional influent feature data sequence and expected operating parameters, H predicted recovery effect index sequences within the preset operating window are predicted as several predicted recovery effect index sequences.
7. The wastewater reuse process optimization method based on recycling effect prediction according to claim 1, characterized in that, For each of the predicted recovery performance index sequences, a sliding window anomaly detection is performed. Based on the deviation between the predicted value and the theoretical expected value of each node, abnormal running nodes within the preset running window are identified, and the abnormal running segment sequence is determined, including: Randomly select a first predicted recovery effect index sequence from the plurality of predicted recovery effect index sequences; A sliding window anomaly detection is performed on the first predicted recovery effect index sequence. Abnormal running nodes are identified based on the deviation between the predicted value and the theoretical expected value of each node, and the first initial abnormal running segment sequence is determined. Traverse the remaining sequences that were not selected from the several predicted recovery effect index sequences, perform the same sliding window anomaly detection on each sequence, determine the corresponding initial anomaly runtime segment sequence, and obtain several initial anomaly runtime segment sequences. The initial abnormal runtime segment sequences are merged to generate the final abnormal runtime segment sequence.
8. The wastewater reuse process optimization method based on recycling effect prediction according to claim 7, characterized in that, A sliding window anomaly detection is performed on the first predicted recovery performance index sequence. Based on the deviation between the predicted value and the theoretical expected value of each node, abnormal running nodes are identified to determine the first initial abnormal running segment sequence, including: The first L consecutive data points of the current node are used as the reference sliding window, where L is not less than 5; Based on the predicted recovery performance index segment within the baseline sliding window, the theoretical expected value of the recovery performance index of the current node is predicted, and the relative deviation between the predicted value of the current node and the theoretical expected value is calculated. If the relative deviation exceeds a preset deviation threshold, the current node is marked as an abnormal running node. Traverse all nodes within the preset running window according to the baseline sliding window, and output the first abnormal running node set within the preset running window of the first predicted recovery effect index sequence. Obtain the time period consisting of all consecutive abnormal running nodes in the first abnormal running node set, and calculate the proportion of the time span of each consecutive abnormal running node time period to the total duration of the preset running window. When the proportion exceeds the preset abnormal proportion threshold, the consecutive abnormal running node time period is determined as an initial abnormal running segment. Summarize all consecutive time periods that meet the preset abnormal proportion threshold, and output the first initial abnormal running segment sequence corresponding to the first predicted recovery effect index sequence.
9. The wastewater reuse process optimization method based on recycling effect prediction according to claim 8, characterized in that, Based on the predicted recovery performance index segment within the baseline sliding window, the theoretical expected value of the recovery performance index for the current node is predicted, and the relative deviation between the predicted value and the theoretical expected value for the current node is calculated. If the relative deviation exceeds a preset deviation threshold, the current node is marked as an abnormal operating node, including: Using the node span of the benchmark sliding window as a constraint, several sample recovery effect index segments are collected, and the historical actual recovery effect index of the next adjacent node of different sample recovery effect index segments is collected as the theoretical expected value of the sample, thus obtaining several sample theoretical expected values. Using the aforementioned sample recovery effect index fragments and several sample theoretical expected values, a long short-term memory network is trained until convergence to generate an index theoretical predictor; Using the aforementioned indicator theory predictor, based on the predicted recovery effect indicator segment within the baseline sliding window, the theoretical expected value of the recovery effect indicator for the current node is predicted. Calculate the relative deviation between the predicted value of the current node and the theoretical expected value to obtain the deviation of effluent water quality characteristics, recovery rate deviation and energy consumption cost deviation; If at least two of the effluent water quality characteristic deviation, recovery rate deviation, and energy consumption cost deviation exceed their respective preset deviation thresholds, the current node will be marked as an abnormal operating node.
10. A wastewater reuse process optimization system based on recycling effect prediction, characterized in that, The system is used to implement the wastewater reuse process optimization method based on recycling effect prediction as described in any one of claims 1-9, the system comprising: The influent characteristic prediction module is used to predict and obtain a sequence of multidimensional influent characteristic data within a preset operating window based on historical wastewater reuse operation monitoring data. The recycling effect prediction module is used to perform a volatility analysis of the recycling effect prediction based on the predicted multidimensional influent feature data sequence, call the pre-trained recycling effect predictor based on the prediction volatility coefficient, and combine the expected operating parameters of wastewater reuse to predict a series of predicted recycling effect indicators within the preset operating window. An abnormal time period identification module is used to perform sliding window anomaly detection on each of the predicted recovery effect index sequences. Based on the deviation between the predicted value and the theoretical expected value of each node, it identifies the abnormal running nodes within the preset running window and determines the abnormal running time sequence. The process optimization and control module is used to take the abnormal running segment sequence as the time period to be optimized within the preset running window, and to perform optimization and control of the wastewater reuse process within the time period to be optimized in the wastewater reuse treatment process in combination with the preset monitoring-feedback mechanism.