Sewage plant total nitrogen concentration real-time prediction and process regulation and control method based on optimization integration algorithm
By optimizing the integrated algorithm, the real-time prediction and process control of total nitrogen concentration in sewage treatment plants were improved, solving the problems of long detection cycle, response delay and resource waste, and achieving efficient and accurate sewage treatment.
Patent Information
- Application Number
- CN202510562306.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-09-05
AI Technical Summary
Existing sewage treatment plants have problems in total nitrogen concentration detection, prediction and control, such as long detection cycle, complex operation, missing data, accuracy fluctuation, response delay, resource waste and systematic lag, making it difficult to achieve real-time process optimization.
By adopting a method based on optimization integration algorithm, through dynamic feature selection, adaptive compensation, deep-integrated model and collaborative optimization mechanism, an intelligent closed-loop control system is constructed to improve data quality and prediction stability and realize advanced control of process parameters.
It achieves high-precision total nitrogen concentration prediction and process control, reduces energy consumption and drug consumption, forms coordinated optimization of the entire process, and improves the real-time performance and efficiency of sewage treatment.
Smart Images

Figure CN120595735A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of environmental monitoring and treatment, and in particular to a method for real-time prediction and process control of total nitrogen concentration in a sewage treatment plant based on an optimized integrated algorithm. Background Art
[0002] Total nitrogen concentration is a core indicator for assessing the severity of water pollution. Exceeding the standard can lead to serious environmental problems such as eutrophication and ecosystem imbalance. Wastewater treatment plants are key facilities for reducing nitrogen pollution and require precise control of effluent total nitrogen concentration to meet increasingly stringent environmental standards. However, existing technology systems still have significant shortcomings in detection, prediction, and control, hindering the optimization of wastewater treatment processes.
[0003] Traditional total nitrogen testing relies on laboratory chemical analysis methods, which suffer from long testing cycles and complex operations, making it difficult to meet the needs of real-time process control. While online monitoring equipment can increase data collection frequency, it is limited by inherent flaws such as sensor susceptibility to interference and high maintenance costs. In practice, data loss and accuracy fluctuations are common, leading to a lack of reliable basis for real-time decision-making.
[0004] Existing prediction methods, primarily based on traditional empirical models or single machine learning algorithms, struggle to effectively characterize the complex dynamic characteristics of wastewater treatment processes. Specifically, single models are poorly adaptable to sudden changes in water quality and load fluctuations, have limited ability to extract features from high-dimensional data, and exhibit significantly reduced generalization across operating conditions. These models also suffer from insufficient predictive stability in dynamic environments and struggle to capture implicit correlations between process parameters, limiting the reliability of prediction results.
[0005] Current control methods often rely on manual experience or local feedback mechanisms, resulting in systemic lags and resource waste. Key parameter adjustments lack foresight, and operations like aeration volume and carbon source addition often suffer from redundant energy and chemical consumption due to delayed response. Multiple control units operate independently, and insufficient collaborative optimization can easily lead to system fluctuations. Furthermore, the disconnect between prediction and control leads to a failure to form a closed-loop optimization mechanism, hindering the effectiveness of real-time process adjustments.
[0006] While existing research attempts to combine machine learning and automatic control technologies, it still faces bottlenecks such as data-model disconnection and low system integration. Predictive model outputs are not deeply coupled with real-time control strategies, and process optimization remains limited to offline analysis. Detection, prediction, and control processes operate in isolation, lacking a framework for collaborative optimization across the entire process. There is an urgent need to build an intelligent decision-making system that integrates prediction and control through innovative algorithmic architecture and system integration to overcome the limitations of current technology. Summary of the Invention
[0007] To address the aforementioned issues with existing technologies, the present invention provides a method for real-time prediction and process control of total nitrogen concentration in sewage treatment plants based on an optimized integrated algorithm. This method improves data quality through dynamic feature selection and adaptive compensation, enhances prediction stability through the integration of a deep-integrated model, and utilizes a collaborative optimization mechanism to proactively control process parameters, forming an intelligent closed-loop control system that addresses control lag and enables precise and efficient sewage treatment operations.
[0008] The technical solutions of the present invention are as follows:
[0009] The present invention provides a real-time prediction method for total nitrogen concentration in a sewage treatment plant based on an optimized integrated algorithm, comprising the following steps:
[0010] Step S1: monitor and collect sewage treatment plant data, reconstruct the data into an original feature vector set according to the time series after cleaning, and use the total nitrogen concentration as the label vector;
[0011] Step S2: perform data correlation analysis on the original feature vector set, filter out feature vectors that are strongly correlated with the label vector, and perform data conversion on the filtered feature data to form a normalized data set;
[0012] Step S3: performing stratification and classification processing on the normalized data set, and using multiple machine learning algorithms to construct total nitrogen concentration prediction models, respectively, optimizing the parameters of each model using the training set, and evaluating the performance of each model using the validation set;
[0013] Step S4: Use each single machine learning model to predict total nitrogen concentration, perform basic performance evaluation and two-dimensional comparative analysis, perform comprehensive performance evaluation based on data characteristics, and determine the optimal total nitrogen concentration prediction model;
[0014] Step S5: using multiple optimization algorithms to optimize the hyperparameters of the total nitrogen concentration prediction model, achieving the global optimal model through iterative search, and improving the prediction accuracy;
[0015] Step S6: Use multiple optimization integrated models to predict total nitrogen concentration, compare the predicted values with the actual values to perform basic performance evaluation, and perform scenario-specific analysis. Combined with data characteristics, a comprehensive evaluation is performed to determine the optimal total nitrogen concentration prediction optimization model;
[0016] Step S7: Containerize and deploy the optimization model to achieve real-time data interaction and rolling prediction. Dynamically update the model through online learning, and implement closed-loop control to stabilize water quality and reduce costs.
[0017] Preferably, the sewage treatment plant data in step S1 include: water quality parameters and meteorological data; wherein the water quality parameters include: sewage discharge volume, pH value, chemical oxygen demand (COD) concentration, chemical oxygen demand (COD) discharge volume, ammonia nitrogen concentration, ammonia nitrogen discharge volume, total nitrogen concentration, total nitrogen discharge volume, total phosphorus concentration, and total phosphorus discharge volume; wherein the meteorological data include: local daily maximum temperature, minimum temperature, and average temperature;
[0018] The method for cleaning the sewage treatment plant data includes: detecting and repairing missing values by a sliding window method, and eliminating outliers based on a combination of box plot and 3σ criterion detection;
[0019] The method for forming the original feature vector set includes: extracting 72-hour dynamic correlation features using time series analysis, and constructing a feature vector containing trend terms, period terms and residual terms through a long short-term memory network (LSTM) encoder.
[0020] Preferably, the method of data correlation analysis in step S2 is: using the Pearson correlation coefficient r to calculate the linear correlation strength between the candidate feature vector and the label vector, and determining the vector with r greater than 0.2 as the feature vector of the model;
[0021] The specific method of data conversion is: using Z-Score standardization or Min-Max normalization method to eliminate dimensional differences, ensure the convergence and prediction stability of the model, and obtain a normalized data set.
[0022] Furthermore, step S2 also includes: selecting features that have a strong correlation with the label, and in order to reduce the feature dimension, remove noise, and enhance the model interpretability, screening out feature vectors with a Pearson correlation coefficient r greater than 0.2 with the label vector, and determining them as feature vectors of the total nitrogen concentration prediction model.
[0023] Preferably, the multiple machine learning algorithms in step S3 include: random forest RF, support vector machine SVM, Gaussian process regression GPR and k-nearest neighbor algorithm KNN;
[0024] The hierarchical processing is to segment the data according to different strategies, including:
[0025] (1) Random split: randomly allocate training set, validation set, and test set in a ratio of 8:1:1;
[0026] (2) Time series segmentation: strictly follow the principle of time continuity and divide the time series data into training set, validation set and test set in a ratio of 8:1:1;
[0027] The classification process is to filter specific condition data from the standard data set to form a subset, including:
[0028] (1) Extreme working condition data subsets: rainstorm period data subset, equipment failure data subset;
[0029] (2) Scenario data subset: Spatiotemporal division based on the inlet load fluctuation cycle and seasonal water quality change factors, including: high load period scenario subset, low temperature period scenario subset, and rainy season scenario subset;
[0030] The random segmentation dataset and the time series segmentation dataset are used to train each model on the training set, and the cross-validation method is used to optimize the hyperparameters, and the optimal parameters are finally determined. The performance of each model is evaluated through the validation set.
[0031] Furthermore, in step S3, the time series segmentation of the data set strictly follows the principle of time continuity to ensure that the timestamp of the test set data is after that of the training set.
[0032] Furthermore, in step S3, the following hyperparameters are tuned, including: the number and depth of decision trees of random forest, the kernel function and regularization parameters of support vector machine, the kernel function configuration of Gaussian process regression, and the neighborhood parameters of k-nearest neighbor algorithm.
[0033] Furthermore, in step S3, the validation set is used to monitor the overfitting risk of the model in real time, and indicators such as mean square error and determination coefficient are used to perform multi-dimensional performance evaluation.
[0034] Preferably, the specific method of step S4 includes:
[0035] (1) The total nitrogen concentration of the randomly segmented dataset and the time series segmented dataset were predicted based on each single machine learning model;
[0036] (2) Determine performance evaluation indicators, including: determination coefficient R 2 , mean square error MSE, mean absolute error MAE and root mean square error RMSE;
[0037] (3) Based on the performance evaluation indicators, performance evaluation is performed on different data sets, including: for randomly segmented data sets, box plot analysis of the error distribution characteristics of each model is used, and the Kruskal-Wallis test is used to compare the significant differences in prediction errors of different models; for time series segmented data sets, the rolling window prediction method is introduced to verify the temporal generalization ability of the model, and the dynamic time warping (DTW) algorithm is used to quantify the morphological differences between the predicted curve and the true curve;
[0038] (4) Conduct a two-dimensional comparative analysis: Under the same segmentation mode, compare the prediction accuracy, computational efficiency, and error stability of each model horizontally, use the entropy weight TOPSIS method to conduct a comprehensive evaluation of multiple indicators, and select the optimal model; under different segmentation modes, analyze the performance attenuation rate of the same model in random segmentation and time series segmentation scenarios vertically, and analyze the impact of the time-varying characteristics of water quality parameters on model robustness by combining feature importance ranking;
[0039] (5) Based on the above performance evaluation and comparative analysis, the model with excellent performance in the random segmentation dataset and a performance attenuation rate of ≤15% in the time series segmentation scenario was selected as the optimal total nitrogen concentration prediction model.
[0040] Furthermore, in step S4, on the same data set, the total nitrogen concentration is predicted by each model, the predicted value of the model is compared with the true value, and the error distribution is analyzed; on the differently segmented data sets, the prediction accuracy of the same model is compared, and a scenario-specific analysis is performed.
[0041] Furthermore, in step S4, based on the performance evaluation index, for the same data set, the predicted values of different models are compared with the true values, and the performance of different models in terms of prediction accuracy, computational efficiency and error control is compared; based on the performance evaluation index, for data sets formed by different segmentation strategies, the predicted values of different models are compared with the true values, and a scenario-specific analysis of the prediction model is performed;
[0042] Preferably, the optimization algorithm in step S5 includes: particle swarm optimization algorithm PSO, differential evolution algorithm DE, ant colony algorithm ACO and adaptive genetic algorithm AGA.
[0043] Furthermore, step S5 further includes searching the hyperparameter space for the selected optimal total nitrogen concentration prediction model. Each algorithm employs a differentiated search strategy: PSO explores parameter combinations through a particle velocity-position update mechanism, DE generates candidate solutions using differential mutation, ACO dynamically adjusts the search path based on pheromone concentration, and AGA employs adaptive crossover / mutation probabilities to improve global optimization efficiency. During the algorithm iteration process, the RMSE on the validation set is used as the fitness function, and Pareto front analysis is used to balance model accuracy and complexity.
[0044] Furthermore, step S5 also includes establishing a multi-algorithm collaborative optimization framework: 1) assigning differentiated hyperparameter search ranges to each optimization algorithm during the initialization phase to avoid local optimality traps; 2) running the four algorithms simultaneously on a parallel computing platform, exchanging the dominant solution set after each iteration using an elite retention strategy; and 3) implementing an early stopping mechanism to monitor and verify the convergence of the loss curve. Finally, a multi-objective decision-making method is used to select the optimal parameter configuration from the set of non-inferior solutions.
[0045] Furthermore, step S5 further includes: after completing hyperparameter optimization, loading the optimal configuration into the prediction model to construct a total nitrogen concentration prediction optimization system. This system can receive preprocessed water quality feature vectors (including LSTM-encoded 72-hour trend-cycle features) in real time, integrate the optimized algorithm core to perform multi-step rolling prediction, and output a high-precision total nitrogen concentration prediction curve for a period of time in the future. The confidence interval width of the prediction result is controlled within ±0.3 mg / L, providing a quantitative basis for process control such as aeration volume and carbon source addition.
[0046] Preferably, the specific method of step S6 includes:
[0047] Using the same dataset, different optimized ensemble models were used to predict total nitrogen concentration. The error distribution characteristics between the predicted and measured values were compared, including the response characteristics of each model under extreme conditions and key nodes of water quality mutations, to analyze the generalization capabilities of each model.
[0048] Compare the prediction accuracy differences of different data subsets after the classification processing of the same optimized integrated model, and complete the scenario-specific analysis by combining the correlation analysis of process parameters;
[0049] The optimization integration model includes: PSO-SVM, DE-SVM, ACO-SVM, and AGA-SVM optimization models.
[0050] Preferably, the performance evaluation index in step S6 includes: determination coefficient R 2 , mean square error MSE, mean absolute error MAE and root mean square error RMSE;
[0051] The specific methods of the comprehensive assessment include:
[0052] (1) Based on the evaluation indicators, a two-layer evaluation framework is constructed, which includes static accuracy evaluation and dynamic stability verification. In terms of prediction accuracy, the focus is on analyzing the error distribution concentration of the model within the confidence interval. In terms of computational efficiency, the model training time and prediction response delay are recorded simultaneously. In terms of error control, an error accumulation warning mechanism based on a sliding window is established.
[0053] (2) The optimal model for total nitrogen concentration prediction was selected through comprehensive evaluation of prediction accuracy, computational efficiency, and error control. At the same time, the operating characteristics of the data subsets were mined based on cluster analysis. By analyzing the performance fluctuation patterns of the same model in different data subsets, a model-scenario adaptation matrix containing weight distribution coefficients was established to determine the prediction model with specific adaptability in each scenario.
[0054] Furthermore, step S6 also includes analyzing the error distribution characteristics of each model in the peak total nitrogen concentration range (>15 mg / L) using kernel density estimation, and quantifying the model's response delay to sudden influent load events using a dynamic time warping algorithm. For a subset of extreme operating condition data (such as periods of heavy rain and equipment failure), Shapley values are used to analyze feature contributions and assess the model's generalization ability to process abnormalities.
[0055] Furthermore, step S6 also includes: calculating the normalized root mean square error (nRMSE) and relative deviation (RPD) of the models in each subset, and combining the correlation between the model prediction accuracy and process parameters (such as SRT, DO, MLSS) with the grey correlation analysis to identify the key control factors affecting the scenario adaptability.
[0056] Furthermore, in step S7, under the premise of ensuring that the effluent stably reaches the Class A standard, the coordinated optimization goal of reducing the amount of carbon source added and reducing the aeration energy consumption is achieved.
[0057] The present invention also provides a nitrogen removal process control method based on the real-time prediction method of total nitrogen concentration, comprising the following steps:
[0058] The optimal total nitrogen concentration prediction optimization model described in step S6 is deployed in a containerized manner to the sewage treatment plant monitoring platform, connected to the data acquisition and monitoring control system SCADA for real-time data interaction, and a rolling prediction is performed based on a sliding time window, dynamically integrating the latest water quality parameters to update the prediction results;
[0059] Based on dynamically collected influent water quality parameters, process operating status, and environmental factor data, a coordinated control scheme for reagent dosage, aeration intensity, and recirculation ratio is generated using a multi-objective optimization algorithm. This scheme is then verified through virtual simulation using a digital twin system to predict the risk of effluent compliance after process parameter adjustments.
[0060] A three-level early warning mechanism is established. When the predicted total nitrogen concentration exceeds the preset threshold, the control instruction generation module is automatically triggered, and the differentiated optimization plan is pushed to the operation interface for manual review and confirmation. At the same time, actual operation data is fed back to the model for incremental training through the online learning mechanism;
[0061] An adaptive feedback loop based on model predictive control (MPC) is established to dynamically calibrate the total nitrogen predicted value with the actual measured value of the online analyzer. The process control parameters are corrected in real time based on the error analysis results, forming a full-process intelligent decision-making system of "data collection-rolling prediction-scheme generation-virtual verification-early warning push-closed-loop optimization" to achieve coordinated optimization of reducing carbon source dosage and reducing aeration energy consumption.
[0062] The beneficial technical effects of the present invention are:
[0063] 1. Data-driven high-precision prediction and dynamic adaptability enhancement: Through the sliding window method, outlier detection, LSTM time series feature extraction technology, dynamic feature selection mechanism and adaptive data compensation system, improve data quality and real-time performance, overcome data loss and noise interference; adopt a deep-integrated hybrid modeling framework, integrate multiple machine learning algorithms (such as RF, SVM) and global optimization algorithms (such as PSO, DE), realize dynamic optimization of hyperparameters and cross-condition transfer learning, and enhance the model's robustness and generalization ability to water quality mutations and load fluctuations. Combined with a multi-dimensional evaluation system (static accuracy R 2 , dynamic error accumulation warning) and scenario adaptation matrix to achieve specialized model optimization under different working conditions, significantly improving the model's robustness, generalization ability, prediction accuracy and stability to water quality mutations and load fluctuations.
[0064] 2. Closed-loop intelligent control and resource collaborative optimization: A multi-objective optimization algorithm generates a coordinated control scheme for aeration intensity, carbon source dosage, and recirculation ratio. Through a feedforward-feedback mechanism and digital twin virtual verification, dynamic matching and proactive adjustment of process parameters are achieved, reducing redundant energy and chemical consumption. Relying on containerized deployment and real-time interaction with the SCADA system, a closed-loop "monitoring-prediction-control-optimization" management and control system is constructed. Combining a three-level early warning mechanism with a model predictive control (MPC) feedback loop, prediction errors are dynamically calibrated and online learning updates are triggered, forming an adaptive, full-process intelligent decision-making network that achieves efficient resource utilization while stabilizing effluent quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Figure 1 This is a schematic diagram of a method for real-time prediction and process control of total nitrogen concentration in a sewage treatment plant based on an optimized integrated algorithm proposed in the present invention. DETAILED DESCRIPTION
[0066] The present invention is described in detail below with reference to the accompanying drawings and embodiments. It is apparent that the embodiments described are only a portion of the embodiments of the present invention, rather than all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are intended to fall within the scope of protection of the present invention.
[0067] Example 1:
[0068] This embodiment provides a method for predicting total nitrogen concentration in a sewage treatment plant based on an optimized integrated algorithm. The specific process is as follows: Figure 1 As shown, the following steps are included:
[0069] Step S1: monitor and collect sewage treatment plant data, reconstruct the data into an original feature vector set according to the time series after cleaning, and use the total nitrogen concentration as the label vector;
[0070] This step is the starting link of the technical solution of the present invention, which mainly involves obtaining sewage plant monitoring data and performing preliminary cleaning and preprocessing on the data.
[0071] In this embodiment, sensors or other data acquisition equipment are used to continuously monitor data at the sewage treatment plant's main sewage outlet, including sewage discharge volume, pH value, chemical oxygen demand (COD) concentration, chemical oxygen demand (COD) emissions, ammonia nitrogen concentration, ammonia nitrogen emissions, total nitrogen concentration, total nitrogen emissions, total phosphorus concentration, total phosphorus emissions, and the local daily maximum, minimum, and average temperatures. The data should be arranged in a time series format to clearly reflect the trend of data changes over time. At the same time, the data collection frequency should be reasonably set according to actual needs to ensure that the complex characteristics of sewage plant operation data can be fully captured.
[0072] The original data obtained were tested for missing values and interpolated using the sliding window method, and outliers were identified and eliminated using a combination of boxplot analysis and the 3σ criterion.
[0073] Then, the dynamic correlation features of the 72-hour span are extracted through time series analysis methods, and a long short-term memory network (LSTM) encoder is used to construct a multi-dimensional feature vector containing trend terms, cycle terms, and residual terms.
[0074] Finally, the total nitrogen concentration of the effluent was used as the supervised learning label, and the preprocessed water quality parameters, meteorological data and derived features were integrated into a structured data set, providing a high-quality data foundation with time-series dynamic characteristics for subsequent model training.
[0075] Step S2: perform data correlation analysis on the original feature vector set, filter out feature vectors that are strongly correlated with the label vector, and perform data conversion on the filtered feature data to form a normalized data set;
[0076] This step mainly involves correlation analysis and data preprocessing of the monitoring data, and further determining performance evaluation indicators.
[0077] In this embodiment, in order to reduce model complexity and save computing time, features with low correlation with the label vector are eliminated through correlation analysis, and features that contribute more to model training are selected as input feature vectors. The linear correlation strength between the candidate feature vectors and the label vector is calculated using the Pearson correlation coefficient. The Pearson correlation coefficient is defined as the ratio of the covariance between two vectors to the product of their standard deviations, usually represented by r, and is calculated by the following formula:
[0078]
[0079] Where: x i is the value of the candidate eigenvector, y iis the value of the label vector, and They represent the sample means of the two respectively.
[0080] Vectors with a strong correlation with the label were selected. In order to reduce the feature dimension, remove noise, and enhance the model interpretability, vectors with a Pearson correlation coefficient r greater than 0.2 with the label vector were screened out and determined as the feature vectors of the total nitrogen concentration prediction model.
[0081] Furthermore, the filtered feature vector data is transformed using the Z-Score normalization or Min-Max normalization method. The formula is as follows:
[0082]
[0083] Where: X new The data values of each feature vector after standard normalization processing, X, X max , and X min are the data value, maximum value and minimum value of each eigenvector respectively.
[0084] This is to eliminate the impact of different dimensions and orders of magnitude on data analysis, ensure the convergence and prediction stability of the model, and obtain the final normalized data set that needs to be processed by the machine model.
[0085] Step S3: performing stratification and classification processing on the normalized data, and using multiple machine learning algorithms to construct total nitrogen concentration prediction models, respectively, optimizing the parameters of each model through the training set, and evaluating the performance of each model through the validation set;
[0086] This step involves different data set segmentation strategies, which classify the data set into data subsets and use the monitoring data to train multiple machine learning models to obtain a total nitrogen concentration prediction model.
[0087] In this example, a structured dataset was first partitioned using both random and time-series partitioning. The first group employed a random partitioning strategy, randomly assigning training, validation, and test sets in an 8:1:1 ratio. The second group strictly adhered to the principle of temporal continuity, partitioning the time-series data in the same proportions to ensure that the test set data timestamps were after the training set. Both partitioning groups used cross-validation to optimize hyperparameters on the training set, including the number and depth of decision trees for random forests, the kernel function and regularization parameters for support vector machines, the kernel function configuration for Gaussian process regression, and the neighborhood parameters for the k-nearest neighbor algorithm. Ultimately, the optimal parameter combination for each model was determined.
[0088] Table 1 Optimal parameters of support vector machine model
[0089] Hyperparameters Parameter value Kernel Function linear Regularization parameter 100 Parameters in the epsilon-insensitive loss function 0.01 Training sample ratio 80%
[0090] Furthermore, four prediction models were constructed based on the optimized parameters: random forest (RF), support vector machine (SVM), Gaussian process regression (GPR), and k-nearest neighbor (KNN). During model training, a feature vector containing dynamic water quality characteristics, meteorological factors, and a composite trend-cycle-residual feature generated by an LSTM encoder was used as input. Time series data on effluent total nitrogen concentration served as supervisory labels for iterative learning. A validation set was used to monitor the risk of model overfitting in real time, and multi-dimensional performance evaluation was performed using metrics such as mean squared error and coefficient of determination.
[0091] After model training, the prediction strengths of each model are combined through ensemble learning to generate robust total nitrogen concentration predictions. Based on real-time input of water quality parameters and feature vectors, this model can output predicted total nitrogen concentration values within a future time window, providing a basis for proactive decision-making in process control.
[0092] In addition, specific conditional data can be filtered from the normalized data set to form a data subset, which can be labeled with different labels, including:
[0093] (1) Extreme operating condition data subset (e.g., rainstorm period, equipment failure period);
[0094] (2) Scenario data subsets: Based on the periodicity of the inflow load (daily / monthly scale) and the seasonal water quality variation pattern (temperature, rainfall), the original normalized dataset is divided into typical scenario subsets such as high load period, low temperature period, and rainy season.
[0095] Step S4: Use each single machine learning model to predict total nitrogen concentration, perform basic performance evaluation and two-dimensional comparative analysis, perform comprehensive performance evaluation based on data characteristics, and determine the optimal total nitrogen concentration prediction model;
[0096] This step compares the prediction accuracy of different models based on performance evaluation indicators and selects the best total nitrogen concentration prediction model.
[0097] In this example, two data segmentation strategies (random segmentation and time series segmentation) were first used to generate a comparative experimental group, and the prediction performance of the random forest (RF), support vector machine (SVM), Gaussian process regression (GPR) and k-nearest neighbor (KNN) models was systematically evaluated. For each model, the coefficient of determination (R) was calculated on the two data sets. 2 ), mean square error (MSE), mean absolute error (MAE), and root mean square error (RMSE), whose mathematical expressions are defined as follows:
[0098]
[0099] Where: n is the length of sample data, y i is the actual value, is its corresponding predicted value, is the data average.
[0100] In the evaluation of the randomly segmented dataset, boxplots were used to analyze the error distribution characteristics of each model, and the Kruskal-Wallis test was used to compare the significance of the prediction errors of different models. For the time series segmented dataset, a rolling window prediction method was introduced to verify the temporal generalization ability of the model, and the dynamic time warping (DTW) algorithm was used to quantify the morphological differences between the predicted and true curves.
[0101] Furthermore, through two-dimensional comparative analysis: 1) Under the same segmentation mode, the prediction accuracy of each model (R 2 , RMSE), computational efficiency (single inference time) and error stability (MAE standard deviation), the entropy-weighted TOPSIS method was used to conduct a comprehensive evaluation of multiple indicators to screen the optimal model; 2) under different segmentation modes, the performance decay rate of the same model in independent and identically distributed (random segmentation) and time extrapolation (time series segmentation) scenarios was longitudinally analyzed, and the influence of the time-varying characteristics of water quality parameters on the robustness of the model was analyzed in combination with the feature importance ranking.
[0102] Furthermore, based on the comprehensive evaluation results of model accuracy, timeliness and scenario adaptability, the model with excellent performance in random segmentation data sets and a performance decay rate of ≤15% in time series segmentation scenarios was selected as the core prediction unit to provide decision support for process control with both accuracy and engineering applicability.
[0103] Table 2 Comparison of evaluation indicators of various prediction models on random split datasets
[0104]
[0105] Step S5: using multiple optimization algorithms to optimize the hyperparameters of the total nitrogen concentration prediction model, achieving the global optimal model through iterative search, and improving the prediction accuracy;
[0106] This step involves further optimization of the optimal total nitrogen concentration prediction model to improve its performance.
[0107] This embodiment includes two core links:
[0108] Four metaheuristic algorithms—particle swarm optimization (PSO), differential evolution (DE), ant colony algorithm (ACO), and adaptive genetic algorithm (AGA)—were used to search the hyperparameter space for the optimal total nitrogen concentration prediction model selected in step 4. Each algorithm employed a differentiated search strategy: PSO explored parameter combinations through a particle velocity-position update mechanism, DE generated candidate solutions using differential mutation, ACO dynamically adjusted the search path based on pheromone concentration, and AGA employed adaptive crossover / mutation probabilities to improve global optimization efficiency. During the algorithm iterations, the RMSE on the validation set was used as the fitness function, and Pareto front analysis was used to balance model accuracy and complexity.
[0109] Furthermore, a multi-algorithm collaborative optimization framework was established: 1) During the initialization phase, different hyperparameter search ranges were assigned to each optimization algorithm to avoid local optimality traps; 2) Four algorithms were run simultaneously on a parallel computing platform, and the dominant solution set was exchanged after each iteration using an elite retention strategy; 3) An early stopping mechanism was implemented to monitor the convergence of the validation loss curve, terminating training when the loss decreased by less than 1% after 20 consecutive optimization rounds. Finally, a multi-objective decision-making method was used to select the optimal parameter configuration from the non-inferior solution set, reducing the model's MAE on the test set by 15%-22%.
[0110] After completing hyperparameter optimization, the optimal configuration was loaded into the prediction model to build a total nitrogen concentration prediction optimization system. This system receives preprocessed water quality feature vectors (including LSTM-encoded 72-hour trend-cycle features) in real time, integrates the optimized algorithm core for multi-step rolling prediction, and outputs a high-precision total nitrogen concentration prediction curve for a period of time. The confidence interval width of the prediction result is controlled within ±0.3mg / L, providing a quantitative basis for process control such as aeration volume and carbon source addition.
[0111] Table 3 Optimal parameters and optimization time of each support vector basis model after optimization
[0112]
[0113]
[0114] Step S6: Use multiple optimization integrated models to predict total nitrogen concentration, compare the predicted values with the actual values to perform basic performance evaluation, and perform scenario-specific analysis. Combined with data characteristics, a comprehensive evaluation is performed to determine the optimal total nitrogen concentration prediction optimization model;
[0115] This step compares the prediction accuracy of different optimization models based on performance evaluation indicators and selects the best total nitrogen concentration prediction model.
[0116] This embodiment includes the following core processes:
[0117] The predictive performance of optimization models, including PSO-SVM, DE-SVM, ACO-SVM, and AGA-SVM, was compared on a unified test set. Kernel density estimation was used to analyze the error distribution characteristics of each model in the peak total nitrogen concentration range (>15 mg / L). The dynamic time warping algorithm was used to quantify the model's response delay to sudden influent load events. For a subset of data under extreme operating conditions (such as periods of heavy rain and equipment failure), Shapley values were used to analyze feature contributions and evaluate the model's generalization ability to process abnormalities.
[0118] Furthermore, based on the periodicity of influent load (daily / monthly) and seasonal water quality variations (temperature and rainfall), the original normalized dataset was divided into subsets representing typical scenarios, such as high-load period, low-temperature period, and rainy season. By calculating the normalized root mean square error (nRMSE) and relative deviation (RPD) of the model within each subset, and combining it with grey correlation analysis to analyze the correlation between model prediction accuracy and process parameters (such as SRT, DO, and MLSS), key control factors influencing scenario adaptability were identified.
[0119] A three-dimensional evaluation system was further constructed: 1) At the static accuracy level, the error distribution concentration within the 95% confidence interval of each model was calculated (Jaccard similarity coefficient ≥ 0.85); 2) At the dynamic stability level, the sliding window coefficient of variation (window width = 24 hours) was used to monitor the cumulative trend of prediction errors. When the CV value of three consecutive windows was > 20%, a regulatory warning was triggered; 3) At the computational efficiency level, the end-to-end latency from data input to output of the model was recorded (required to be < 2 seconds). Based on the entropy-weighted TOPSIS comprehensive evaluation method, accuracy, stability, and timeliness were assigned weights of 0.5:0.3:0.2, and models with the top 20% overall score were selected.
[0120] Furthermore, fuzzy C-means clustering is used to identify the operating condition feature patterns of the data subset, and a model-scenario adaptation matrix is established: for each cluster center scenario, the Z-score standardized value of each model performance indicator is calculated, and the Sigmoid function is used to generate the scenario adaptation score (0-1 range), and the model with adaptation > 0.7 and feature contribution matching > 80% is selected as the recommended predictor for the scenario. For example, in the low temperature and high ammonia nitrogen scenario, PSO-SVM shows significant advantages due to its higher weight allocation to temperature-sensitive features, and its scenario adaptation reaches 0.82. The matrix is dynamically updated through an online learning mechanism to ensure that the model selection keeps pace with real-time operating condition changes.
[0121] Table 4 Comparison of evaluation indicators of various prediction optimization models on random split datasets
[0122]
[0123] Step S7: Containerize and deploy the optimization model to achieve real-time data interaction and rolling prediction. Dynamically update the model through online learning, and implement closed-loop control to stabilize water quality and reduce costs.
[0124] This step involves applying the optimized prediction model to the sewage plant monitoring system to achieve real-time prediction of total nitrogen concentration.
[0125] In this embodiment, the prediction results of the SVM model based on multiple optimization algorithms are compared with the true value of the total nitrogen concentration under the same data set. According to the performance evaluation index, the prediction accuracy, computational efficiency and error index of different models are compared, and the specificity analysis of the scene is performed to determine the optimal total nitrogen concentration prediction optimization model based on the characteristics of the data.
[0126] Example 2:
[0127] This embodiment provides a method for intelligent real-time nitrogen removal process control using the above-mentioned total nitrogen concentration prediction method, including the following specific implementation steps:
[0128] The optimized total nitrogen prediction model is deployed to the production environment using containerization technology. Real-time data interaction is achieved with the monitoring platform through standardized interfaces. Rolling predictions are performed based on a sliding time window mechanism, and the latest water quality parameters are dynamically integrated to update the prediction results.
[0129] Furthermore, a collaborative process parameter optimization module was constructed, integrating a multi-objective optimization algorithm to generate control schemes encompassing key parameters such as reagent dosing and aeration control. This was then verified through virtual simulation using a digital twin system to assess the risk of effluent compliance under different control strategies. A hierarchical early warning mechanism was established. When predicted values approach or exceed standard limits, a multi-level response process was automatically triggered, pushing differentiated control suggestions to the operator terminal for manual review. Actual operating data was also fed back to the online learning module for dynamic model updates.
[0130] Ultimately, a closed-loop control system is formed, dynamically comparing model predictions with online monitoring data to calibrate process control parameters in real time. The system establishes a comprehensive intelligent decision-making process encompassing "data perception - trend prediction - solution generation - virtual verification - decision execution - and feedback on results." While ensuring consistent effluent quality, it achieves the dual optimization goals of operating costs and energy consumption, driving the transformation of the wastewater treatment process toward a smart operation and maintenance model.
[0131] Although the embodiments of the present invention have been disclosed as above, they are not limited to the applications listed in the description and implementation methods. They can be fully applied to various fields suitable for the present invention. For those familiar with the art, for those of ordinary skill in the art, various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. Therefore, without departing from the general concept defined by the claims and the scope of equivalents, the present invention is not limited to specific details.
Claims
1. A real-time prediction method for total nitrogen concentration in a sewage treatment plant based on an optimized integrated algorithm, characterized in that: The following steps are involved: Step S1: monitor and collect sewage treatment plant data, reconstruct the data into an original feature vector set according to the time series after cleaning, and use the total nitrogen concentration as the label vector; Step S2: perform data correlation analysis on the original feature vector set, filter out feature vectors that are strongly correlated with the label vector, and perform data conversion on the filtered feature data to form a normalized data set; Step S3: performing stratification and classification processing on the normalized data set, and using multiple machine learning algorithms to construct total nitrogen concentration prediction models, respectively, optimizing the parameters of each model using the training set, and evaluating the performance of each model using the validation set; Step S4: Use each single machine learning model to predict total nitrogen concentration, perform basic performance evaluation and two-dimensional comparative analysis, perform comprehensive performance evaluation based on data characteristics, and determine the optimal total nitrogen concentration prediction model; Step S5: using multiple optimization algorithms to optimize the hyperparameters of the total nitrogen concentration prediction model, achieving the global optimal model through iterative search, and improving the prediction accuracy; Step S6: Use multiple optimization integrated models to predict total nitrogen concentration, compare the predicted values with the actual values to perform basic performance evaluation, and perform scenario-specific analysis. Combined with data characteristics, a comprehensive evaluation is performed to determine the optimal total nitrogen concentration prediction optimization model; Step S7: Containerize and deploy the optimization model to achieve real-time data interaction and rolling prediction. Dynamically update the model through online learning, and implement closed-loop control to stabilize water quality and reduce costs.
2. The method for real-time prediction of total nitrogen concentration according to claim 1, wherein: The sewage treatment plant data in step S1 include: water quality parameters and meteorological data; wherein the water quality parameters include: sewage discharge volume, pH value, chemical oxygen demand (COD) concentration, chemical oxygen demand (COD) discharge volume, ammonia nitrogen concentration, ammonia nitrogen discharge volume, total nitrogen concentration, total nitrogen discharge volume, total phosphorus concentration, and total phosphorus discharge volume; wherein the meteorological data include: local daily maximum temperature, minimum temperature, and average temperature; The method for cleaning the sewage treatment plant data includes: detecting and repairing missing values by a sliding window method, and eliminating outliers based on a combination of box plot and 3σ criterion detection; The method for forming the original feature vector set includes: extracting 72-hour dynamic correlation features using time series analysis, and constructing a feature vector containing trend terms, period terms and residual terms through a long short-term memory network (LSTM) encoder.
3. The method for real-time prediction of total nitrogen concentration according to claim 1, wherein: The method of data correlation analysis in step S2 is: using the Pearson correlation coefficient r to calculate the linear correlation strength between the candidate feature vector and the label vector, and the vector with r greater than 0.2 is determined as the feature vector of the model; The specific method of data conversion is: using Z-Score standardization or Min-Max normalization method to eliminate dimensional differences, ensure the convergence and prediction stability of the model, and obtain a normalized data set.
4. The method for real-time prediction of total nitrogen concentration according to claim 1, wherein: The multiple machine learning algorithms in step S3 include: random forest RF, support vector machine SVM, Gaussian process regression GPR and k-nearest neighbor algorithm KNN; The hierarchical processing is to segment the data according to different strategies, including: (1) Random split: randomly allocate training set, validation set, and test set in a ratio of 8:1:1; (2) Time series segmentation: strictly follow the principle of time continuity and divide the time series data into training set, validation set and test set in a ratio of 8:1:1; The classification process is to filter specific condition data from the standard data set to form a subset, including: (1) Extreme working condition data subsets: rainstorm period data subset, equipment failure data subset; (2) Scenario data subset: Spatiotemporal division based on the inlet load fluctuation cycle and seasonal water quality change factors, including: high load period scenario subset, low temperature period scenario subset, and rainy season scenario subset; The random segmentation dataset and the time series segmentation dataset are used to train each model on the training set, and the cross-validation method is used to optimize the hyperparameters, and the optimal parameters are finally determined. The performance of each model is evaluated through the validation set.
5. The method for real-time prediction of total nitrogen concentration according to claim 4, wherein: The specific method of step S4 includes: (1) The total nitrogen concentration of the randomly segmented dataset and the time series segmented dataset were predicted based on each single machine learning model; (2) Determine performance evaluation indicators, including: determination coefficient R 2 , mean square error MSE, mean absolute error MAE and root mean square error RMSE; (3) Based on the performance evaluation indicators, performance evaluation is performed on different data sets, including: for randomly segmented data sets, box plot analysis of the error distribution characteristics of each model is used, and the Kruskal-Wallis test is used to compare the significant differences in prediction errors of different models; for time series segmented data sets, the rolling window prediction method is introduced to verify the temporal generalization ability of the model, and the dynamic time warping (DTW) algorithm is used to quantify the morphological differences between the predicted curve and the true curve; (4) Conduct a two-dimensional comparative analysis: Under the same segmentation mode, compare the prediction accuracy, computational efficiency, and error stability of each model horizontally, use the entropy weight TOPSIS method to conduct a comprehensive evaluation of multiple indicators, and select the optimal model; under different segmentation modes, analyze the performance attenuation rate of the same model in random segmentation and time series segmentation scenarios vertically, and analyze the impact of the time-varying characteristics of water quality parameters on model robustness by combining feature importance ranking; (5) Based on the above performance evaluation and comparative analysis, the model with excellent performance in the random segmentation dataset and a performance attenuation rate of ≤15% in the time series segmentation scenario was selected as the optimal total nitrogen concentration prediction model.
6. The method for real-time prediction of total nitrogen concentration according to claim 1, wherein: The optimization algorithms in step S5 include: particle swarm optimization algorithm PSO, differential evolution algorithm DE, ant colony algorithm ACO and adaptive genetic algorithm AGA.
7. The method for real-time prediction of total nitrogen concentration according to claim 1, wherein: The specific method of step S6 includes: Using the same dataset, different optimized ensemble models were used to predict total nitrogen concentration. The error distribution characteristics between the predicted and measured values were compared, including the response characteristics of each model under extreme conditions and key nodes of water quality mutations, to analyze the generalization capabilities of each model. Compare the prediction accuracy differences of different data subsets after the classification processing of the same optimized integrated model, and complete the scenario-specific analysis by combining the correlation analysis of process parameters; The optimization integration model includes: PSO-SVM, DE-SVM, ACO-SVM, and AGA-SVM optimization models.
8. The method for real-time prediction of total nitrogen concentration according to claim 7, wherein: The comprehensive evaluation index in step S6 includes: determination coefficient R 2 , mean square error MSE, mean absolute error MAE and root mean square error RMSE; The specific methods of the comprehensive assessment include: (1) Based on the evaluation indicators, a two-layer evaluation framework is constructed, which includes static accuracy evaluation and dynamic stability verification. In terms of prediction accuracy, the focus is on analyzing the error distribution concentration of the model within the confidence interval. In terms of computational efficiency, the model training time and prediction response delay are recorded simultaneously. In terms of error control, an error accumulation warning mechanism based on a sliding window is established. (2) The optimal model for total nitrogen concentration prediction was selected through comprehensive evaluation of prediction accuracy, computational efficiency, and error control. At the same time, the operating characteristics of the data subsets were mined based on cluster analysis. By analyzing the performance fluctuation patterns of the same model in different data subsets, a model-scenario adaptation matrix containing weight distribution coefficients was established to determine the prediction model with specific adaptability in each scenario.
9. A nitrogen removal process control method based on the real-time prediction method for total nitrogen concentration according to any one of claims 1 to 8, characterized in that: The following steps are involved: The optimal total nitrogen concentration prediction optimization model described in step S6 is deployed in a containerized manner to the sewage treatment plant monitoring platform, connected to the data acquisition and monitoring control system SCADA for real-time data interaction, and a rolling prediction is performed based on a sliding time window, dynamically integrating the latest water quality parameters to update the prediction results; Based on dynamically collected influent water quality parameters, process operating status, and environmental factor data, a coordinated control scheme for reagent dosage, aeration intensity, and recirculation ratio is generated using a multi-objective optimization algorithm. This scheme is then verified through virtual simulation using a digital twin system to predict the risk of effluent compliance after process parameter adjustments. A three-level early warning mechanism is established. When the predicted total nitrogen concentration exceeds the preset threshold, the control instruction generation module is automatically triggered, and the differentiated optimization plan is pushed to the operation interface for manual review and confirmation. At the same time, actual operation data is fed back to the model for incremental training through the online learning mechanism; An adaptive feedback loop based on model predictive control (MPC) is established to dynamically calibrate the total nitrogen predicted value with the actual measured value of the online analyzer. The process control parameters are corrected in real time based on the error analysis results, forming a full-process intelligent decision-making system of "data collection-rolling prediction-scheme generation-virtual verification-early warning push-closed-loop optimization", thereby achieving coordinated optimization of reducing carbon source dosage and reducing aeration energy consumption.
10. The nitrogen removal process control method according to claim 9, characterized in that: The three-level early warning mechanism is applied to the sewage treatment plant monitoring platform. When the predicted value exceeds the threshold, different levels of control instructions are triggered. After manual review, they are automatically sent to the aeration device and dosing pump through the industrial control bus to perform real-time operations.
Citation Information
Cited By
Sewage denitrification dosing method and system based on machine learning and storage medium
CN120877905A
Load prediction method and system based on virtual power plant operation
CN120893637A
Sewage treatment model dynamic optimization method based on time sequence rolling prediction and related equipment
CN121094244A
A wastewater treatment model dynamic optimization method based on time series rolling prediction and related equipment
CN121094244B
Shield tunneling machine tunneling speed intelligent prediction method based on multi-algorithm collaborative optimization
CN121257346A