Sewage plant total phosphorus concentration prediction method and control system based on multiple machine learning models

By combining multiple machine learning models and swarm intelligence optimization algorithms, a high-performance total phosphorus concentration prediction system was constructed, which solved the complexity and real-time problems of total phosphorus concentration prediction in sewage treatment plants and realized intelligent management and efficient operation of sewage treatment plants.

CN120595875APending Publication Date: 2025-09-05NORTH CHINA INST OF AEROSPACE ENG
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510562308.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

The prediction of total phosphorus concentration in sewage treatment plants is difficult to meet the requirements of high precision and real-time performance. Traditional methods have difficulty capturing complex nonlinear relationships, and the generalization performance of a single learning model is insufficient. It cannot be deeply integrated with industrial control systems and lacks online self-learning capabilities.

Method used

By adopting a variety of machine learning models (such as random forest, support vector regression, Gaussian process regression, and k-nearest neighbor) combined with swarm intelligence optimization algorithms, and through intelligent data preprocessing and multi-model collaborative training, a high-performance prediction model is constructed, which is then integrated with the sewage plant PLC system to achieve real-time data collection and dynamic regulation.

Benefits of technology

It improves the accuracy and stability of total phosphorus concentration prediction, realizes intelligent management of sewage treatment process, dynamically adjusts phosphorus removal process, meets real-time control needs, and improves the operating efficiency and accuracy of sewage treatment plants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120595875A_ABST
    Figure CN120595875A_ABST
Patent Text Reader

Abstract

The invention discloses a sewage plant total phosphorus concentration prediction method based on multiple machine learning models and a control system, and belongs to the field of environmental monitoring and treatment. The method comprises the following steps: automatically collecting detection data of a sewage plant, and generating a time sequence data set through intelligent preprocessing; according to the method, a total phosphorus concentration prediction model is constructed by adopting multiple machine learning algorithms, a reference prediction model is automatically selected through evaluation indexes, hyper-parameter tuning is performed by utilizing a swarm intelligence optimization method, and an optimized high-performance prediction model is obtained. The model is deployed to a real-time monitoring system, and through integration with a PLC and monitoring hardware, high-frequency prediction and dynamic regulation and control closed loop are realized; and continuously optimizing model parameters and a regulation and control strategy by returning deviation information to form a'prediction-control-optimization 'closed loop. According to the method, the effluent total phosphorus concentration prediction precision and regulation efficiency are remarkably improved, the agent adding cost is reduced, and the intelligent level of a sewage treatment system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of environmental monitoring and governance technology, and in particular to a method and control system for predicting total phosphorus concentration in a sewage treatment plant based on multiple machine learning models. Background Art

[0002] With the rapid development of urbanization and industrialization, sewage treatment has become a key link in environmental protection and water resource management. As core facilities, the effluent quality of sewage treatment plants directly affects the safety of the water environment. As a major indicator of eutrophication, the concentration control of total phosphorus (TP) is particularly important. However, the sewage treatment process is highly nonlinear and time-varying. Traditional methods based on experience or simple mathematical models are difficult to meet increasingly stringent emission standards and the need for efficient operation. In addition, due to the limitations of testing costs and project cycles, the lack of time series data on key water quality parameters in sewage treatment plants is common. How to build an accurate prediction model based on an incomplete data set has become a difficult technical challenge for the industry.

[0003] Current mainstream total phosphorus concentration prediction methods rely primarily on manual experience or simple statistical models, which struggle to capture complex nonlinear relationships, resulting in low prediction accuracy and difficulty supporting real-time optimization and control of process parameters. Notably, with the development of the Internet of Things (IoT) and big data technologies, online monitoring systems for wastewater treatment plants can collect large amounts of operational data in real time, providing a foundation for building data-driven intelligent prediction models.

[0004] In recent years, the application of machine learning technology in the field of sewage treatment has gradually attracted attention. Compared with traditional methods, machine learning algorithms can achieve more accurate water quality predictions and stronger adaptability by deeply mining the complex nonlinear correlation features in historical operation data. However, a single learning model has the defect of insufficient generalization performance when facing the diversity and uncertainty in the sewage treatment process, and the selection of model hyperparameters has a significant impact on performance. More importantly, the real-time prediction and control system for industrial sites needs to be deeply integrated with industrial control systems (such as PLC) and online monitoring hardware to build a closed-loop optimization mechanism of "high-frequency prediction-dynamic control-effect feedback".

[0005] This places the following technical requirements on the prediction model: it must meet real-time response speed constraints while ensuring prediction stability and accuracy under complex operating conditions. Furthermore, it must possess online self-learning capabilities, enabling the autonomous evolution of model parameters by continuously absorbing new sample data, thereby dynamically improving prediction accuracy and the effectiveness of control strategies. Therefore, developing a total phosphorus concentration prediction method based on the advantages of multiple machine learning models and establishing a matching intelligent control system not only has innovative value in terms of wastewater treatment process control theory, but also has important engineering practical significance for promoting the intelligent upgrading of the wastewater treatment industry. Summary of the Invention

[0006] To address the aforementioned shortcomings of existing technologies, the present invention provides a method and control system for predicting total phosphorus concentration in sewage treatment plants based on multiple machine learning models. Through intelligent data preprocessing and the coordinated application of multiple machine learning algorithms, the present invention effectively fills in missing data and builds a high-performance prediction model. Furthermore, based on the predicted total phosphorus concentration, it enables the pre-determined dosing strategy and optimizes the phosphorus removal process, providing technical support for the intelligent management and efficient operation of sewage treatment plants.

[0007] The specific technical solutions of the present invention are as follows:

[0008] A method for predicting total phosphorus concentration in a sewage treatment plant based on multiple machine learning models includes the following steps:

[0009] Step S1: monitor and obtain sewage treatment plant data, wherein the water inlet monitoring data is used as a feature vector and the water outlet monitoring data is used as a label vector;

[0010] Step S2: performing intelligent preprocessing on the sewage treatment plant data, standardizing the data, and dividing the data set into a training set, a validation set, and a test set in chronological order;

[0011] Step S3: constructing multiple total phosphorus concentration prediction models through simultaneous training of multiple machine learning algorithms, and then determining the predicted value within a preset time period in the future through the prediction models;

[0012] Step S4: Using the same training set, validation set, and test set, the accuracy of each model is compared using multiple performance evaluation indicators, and the adaptability of each model to the nonlinear characteristics of the sewage treatment field is comprehensively considered to automatically determine the optimal model as the benchmark prediction model for total phosphorus concentration;

[0013] Step S5: deeply tune the specific hyperparameters of the benchmark prediction model using a swarm intelligence optimization algorithm, monitor the core performance indicators of the model in real time, and evaluate the prediction accuracy, computational efficiency, and robustness by comparing the error changes before and after optimization, and automatically determine the optimized high-performance prediction model for total phosphorus concentration;

[0014] Step S6: Integrate the high-performance prediction model with the sewage treatment plant programmable logic controller (PLC) and the sewage treatment plant real-time monitoring hardware to dynamically collect real-time data and predict the total phosphorus concentration in the effluent.

[0015] Preferably, the water inlet monitoring data in step S1 includes monitoring time, sewage discharge, pH value, COD concentration, COD discharge, ammonia nitrogen concentration, ammonia nitrogen discharge, total nitrogen concentration, total nitrogen discharge, total phosphorus concentration, and total phosphorus discharge, and the water outlet monitoring data is total phosphorus concentration.

[0016] Furthermore, in step S1, the automated data acquisition module of the sewage treatment plant's real-time monitoring system is used to obtain the sewage treatment plant's inlet and outlet monitoring data in real time to ensure the dynamic update of the data set, wherein the historical monitoring data is used for model training, and the real-time data is used to predict the total phosphorus concentration at the outlet based on the trained model;

[0017] Preferably, the method for intelligent preprocessing of data in step S2 is: identifying and eliminating outliers through a box plot method, using a long short-term memory network LSTM combined with a sliding window to fill in missing time series data, and using linear interpolation to supplement missing values ​​at the beginning of the sequence that cannot meet the sliding window length to ensure the continuity of the prediction;

[0018] The method for normalizing the data is: using Z-Score normalization or Min-Max normalization method to normalize the data after outlier processing, selecting an appropriate data normalization method according to the distribution of characteristic data, eliminating dimensional differences, and generating a high-quality data set reflecting the temporal changes of pollutants in the sewage treatment plant;

[0019] The method for dividing the data set into a training set, a validation set and a test set is: dividing the high-quality data set into a training set, a validation set and a test set in a ratio of 8:1:1 in chronological order.

[0020] Furthermore, when filling missing values ​​in step S2, the sliding window mechanism is combined with the long-term dependency characteristics of the network to fully utilize the temporal characteristics and potential laws of the time series, providing a more accurate basis for data analysis and modeling, and accurately filling missing data using the long-term dependency memory characteristics of the long short-term memory network (LSTM) model.

[0021] Furthermore, during the normalization process in step S2, an appropriate data normalization method is selected based on the distribution of the characteristic data to eliminate the dimensional effects caused by unit differences between the characteristics, ensure data quality, and generate a high-quality data set reflecting the temporal changes of pollutants in the sewage treatment plant;

[0022] Furthermore, in step S2, the segmented data is used to generate a time-series dataset, the training set is used for model training to optimize model parameters, the validation set is used to adjust the model's hyperparameters and monitor model performance, and the test set is used to independently evaluate the generalization performance of the model;

[0023] Preferably, the specific method of step S3 includes:

[0024] Multiple prediction models are constructed for the training set using multiple machine learning algorithms to simulate the nonlinear mapping relationship between the feature vector and the total phosphorus concentration label vector of the outlet, respectively. The multiple machine learning algorithms include: random forest RF, support vector regression SVM, Gaussian process regression GPR and k-nearest neighbor KNN;

[0025] Cross-validation techniques were used to initially optimize the hyperparameters of each model. Performance feedback from the validation set was then used to further fine-tune each model. This fine-tuning included adjusting the number and depth of decision trees for the RF model, the kernel function type and regularization parameter for the SVM model, the kernel function and noise configuration for the GPR model, and the k value and distance metric for the KNN model.

[0026] Each prediction model generates a total phosphorus concentration prediction value based on the test set data through training and learning.

[0027] Furthermore, in step S3, during the training process, the model gradually learns the mapping relationship between input features and output labels based on the training set. Ultimately, each trained model generates a predicted value based on the input data, thereby determining the trend of total phosphorus concentration changes within a preset time period in the future.

[0028] Preferably, the specific method of step S4 includes:

[0029] Calculate the difference between the predicted value and the true value of each prediction model in step S3 based on the test set data;

[0030] A variety of performance evaluation indicators are used to comprehensively measure the prediction model, and the determination coefficient R is used to 2 Evaluate the explanatory variables of the model, use the mean square error (MSE) and root mean square error (RMSE) to evaluate the fluctuation of the sum of squares of the prediction errors, and use the mean absolute error (MAE) to measure the error magnitude;

[0031] The model with the best performance is automatically screened according to the evaluation index as the total phosphorus concentration benchmark prediction model.

[0032] Furthermore, in step S4, in order to adapt to the complexity of sewage treatment data and the nonlinear characteristics of pollutant changes, the system automatically selects the model with the best performance as the total phosphorus concentration benchmark prediction model, providing a basis for subsequent hyperparameter optimization and system deployment;

[0033] Preferably, the swarm intelligence optimization algorithm in step S5 includes: particle swarm optimization algorithm PSO, artificial bee colony algorithm ABC, firefly algorithm FA and firework algorithm FWA;

[0034] The core performance indicators include: determination coefficient R 2 , mean square error MSE, root mean square error RMSE, mean absolute error MAE;

[0035] The computational efficiency is quantified by the actual running time of the model processing a large sample data set, and the robustness is verified through repeated experimental analysis: the model is run multiple times in different data scenarios, and the stability of its output results is analyzed. The test results are presented in the form of mean ± standard deviation, which intuitively reflects the fluctuation range of the indicator.

[0036] Furthermore, the computational efficiency in step S5 is used to evaluate the time-consuming performance of the model on a large sample data set; the robustness of the model is analyzed by repeated running, and its stability performance in different data scenarios is deeply studied, and the index fluctuations of each model in multiple tests are recorded, and the performance of the optimized prediction model in prediction accuracy, computational efficiency and error control is compared, and the changes in model performance before and after optimization are recorded; the error changes of the optimized model and the original benchmark model are compared to verify the significance of the optimization algorithm in improving the model prediction accuracy, reducing error fluctuations and enhancing the model robustness, and finally determine the optimized high-precision model to provide a better solution for actual engineering applications.

[0037] Preferably, the total phosphorus concentration of the effluent predicted in step S6 is used to control the dosage of the phosphorus removal agent.

[0038] The present invention also provides a sewage treatment plant total phosphorus concentration prediction system based on multiple machine learning models, including the following modules:

[0039] Automatic data acquisition module: monitors and obtains sewage treatment plant data, where the inlet monitoring data is used as the feature vector and the outlet monitoring data is used as the label vector;

[0040] Data preprocessing and standardization module: intelligently preprocess the sewage plant data, standardize the data, and divide the data set into training set, validation set and test set in chronological order;

[0041] Multiple machine learning model parallel training and construction modules: Multiple total phosphorus concentration prediction models are constructed through simultaneous training of multiple machine learning algorithms, and then the prediction values ​​within a preset time period in the future are determined by the prediction models;

[0042] Multiple machine learning model evaluation and optimization modules: Using the same training, validation, and test sets, the module uses multiple performance evaluation metrics to compare the accuracy of each model. By comprehensively considering their adaptability to the nonlinear characteristics of wastewater treatment, the module automatically determines the optimal model as the benchmark prediction model for total phosphorus concentration.

[0043] Total phosphorus concentration prediction model intelligent optimization module: This module uses a swarm intelligence optimization algorithm to deeply tune specific hyperparameters of the baseline prediction model, monitors the model's core performance indicators in real time, and compares the error changes before and after optimization to evaluate prediction accuracy, computational efficiency, and robustness. It automatically determines the optimized high-performance total phosphorus concentration prediction model.

[0044] Intelligent management and dynamic deployment module of the total phosphorus concentration prediction model: The high-performance total phosphorus concentration prediction model is deployed to the sewage plant real-time monitoring system, and integrated with the sewage plant programmable logic controller (PLC) and the sewage plant real-time monitoring hardware to dynamically collect real-time data and predict the effluent total phosphorus concentration.

[0045] The present invention further provides a method for controlling the total phosphorus concentration in a sewage treatment plant using a total phosphorus prediction model based on multiple machine learning models, comprising the following steps:

[0046] Integrating the high-performance prediction model in step S5 with the sewage treatment plant programmable logic controller (PLC) and the sewage treatment plant real-time monitoring hardware;

[0047] Dynamically collect real-time data using a set detection frequency, and predict the total phosphorus concentration in the effluent based on the high-performance prediction model;

[0048] When the predicted results deviate from the preset emission standards, the dosage of phosphorus removal agents is dynamically and intelligently adjusted according to the change in the predicted concentration. At the same time, the predicted deviation is fed back to the optimization module, and the model weights and parameters are dynamically recalibrated through deviation back calculation to continuously improve the prediction accuracy and regulation effect until the long-term operation stability requirements are met, realizing the "prediction-control-optimization" closed-loop management.

[0049] Preferably, the monitoring frequency includes minute level and hour level.

[0050] The beneficial technical effects of the present invention are:

[0051] (1) The present invention adopts intelligent data cleaning and feature reconstruction technology to establish a data filling and repair mechanism (such as box plot outlier identification and long short-term memory network missing value filling), which effectively solves the data missing problem caused by equipment aging, environmental fluctuations or improper human operation, and improves the integrity and utilization of data.

[0052] (2) The present invention combines a variety of machine learning algorithms (such as random forest model RF, support vector machine model SVM, Gaussian process regression model GPR, k-nearest neighbor model KNN) and swarm intelligence optimization methods (such as particle swarm optimization algorithm PSO, artificial bee colony algorithm ABC, firefly algorithm FA and fireworks algorithm FWA) to construct a high-precision and robust total phosphorus concentration prediction model, which can accurately capture the complex nonlinear relationship in the sewage treatment process and improve the generalization ability of the prediction model.

[0053] (3) Furthermore, the present invention deeply integrates the high-performance prediction model with the programmable logic controller (PLC) and online monitoring hardware of the sewage treatment plant, adopts the set monitoring frequency to dynamically collect real-time data, predicts the total phosphorus concentration of the effluent in advance based on the prediction model, formulates the dosing strategy in advance according to the predicted value, and dynamically adjusts the amount of phosphorus removal agent; at the same time, by feeding back the deviation information between the actual value and the predicted value to the optimization module, the model parameters are automatically adjusted and the prediction performance is continuously optimized, realizing the "prediction-control-optimization" closed-loop management. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 This is a schematic diagram of a method for predicting total phosphorus concentration in a sewage treatment plant based on multiple machine learning models proposed in the present invention.

[0055] Figure 2 This is a schematic diagram of a sewage treatment plant total phosphorus concentration prediction and control system based on multiple machine learning models proposed in the present invention. DETAILED DESCRIPTION

[0056] The present invention is described in detail below with reference to the accompanying drawings and embodiments. It is apparent that the embodiments described are only a portion of the embodiments of the present invention, rather than all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are intended to fall within the scope of protection of the present invention.

[0057] Example 1:

[0058] This embodiment provides a method for predicting total phosphorus concentration in a sewage plant based on multiple machine learning models. The specific process is as follows: Figure 1 As shown, the following steps are included:

[0059] Step S1, monitoring and acquiring sewage treatment plant data, wherein the water inlet monitoring data is used as a feature vector and the water outlet monitoring data is used as a label vector;

[0060] In this embodiment, 1095 sets of data from a municipal sewage treatment center from January 2021 to December 2023 are obtained through automated sensors or other automated data acquisition equipment. The indicators continuously monitored at the water inlet of the sewage treatment plant are: monitoring time, sewage discharge, pH value, COD concentration, COD discharge, ammonia nitrogen concentration, ammonia nitrogen discharge, total nitrogen concentration, total nitrogen discharge, total phosphorus concentration, total phosphorus discharge, and total phosphorus concentration at the outlet. The data are arranged in a time series to clearly reflect the trend of data changes over time. At the same time, the data collection frequency is reasonably set according to actual needs to ensure that the complex characteristics of the sewage plant operation data can be fully captured.

[0061] The acquired monitoring data of the water inlet is used as the feature vector (i.e., input variable), and the total phosphorus concentration of the outlet is used as the label vector.

[0062] Step S2, performing intelligent preprocessing on the sewage treatment plant data, standardizing the data, and dividing the data set into a training set, a validation set, and a test set in chronological order;

[0063] In this embodiment, the outliers in the original detection data are automatically identified and removed using the box plot method. The data of each feature vector is processed separately. Taking "sewage discharge" as an example, the operation steps are as follows:

[0064] (1) Data sorting: sort all the “wastewater discharge volume” data and arrange them in ascending order;

[0065] (2) Calculate the lower quartile (Q1), median (Q2), and upper quartile (Q3): Assume the number of data is n. If n is an odd number, the median is the value of the (n+1) / 2th data; if n is an even number, the median is the average of the n / 2th and (n / 2+1)th data. Q1 is the median of the first half of the data, and Q3 is the median of the second half of the data.

[0066] (3) Determine the interquartile range: (IQR), IQR = Q3 - Q1;

[0067] (4) Calculate the upper and lower limits: the lower limit is Q1-1.5\times IQR, and the upper limit is Q3+1.5\times IQR;

[0068] (5) Drawing and identification: Draw a box plot, and the data points outside the upper and lower limits are outliers. For sewage discharge, the range of values ​​is large, most of which are between 3000-4500m 3 If the abnormal value is not between 3000-4500m 3 However, when doing so, it is important to consider the actual business situation and data background to avoid blind operations, which may cause deviations in the analysis results.

[0069] Similarly, the pH value usually fluctuates around 7.0, and the data outside the range of 5.0-6.0 are eliminated in the same way as above.

[0070] Furthermore, a deep learning method based on the Long Short-Term Memory (LSTM) network is used to fill in missing values.

[0071] (1) For the missing values ​​at the beginning of the sequence, linear interpolation is used to fill in the missing values ​​because the sliding window length requirement cannot be met. Linear interpolation is to calculate the missing values ​​by performing a linear fit between the two points based on the trend of the known data before and after the missing value.

[0072] (2) For missing values ​​in other locations, a deep learning method based on long short-term memory network (LSTM) is used to fill them.

[0073] ① First, determine the length of the sliding window, for example, set it to 5, and divide the time series data into data windows according to this window length. Each data window contains data from the past 5 consecutive time steps, which serves as the input of the model.

[0074] ② Build a two-layer LSTM network model. The first LSTM layer is responsible for initially extracting features from the input data window and learning short-term time series dependencies. It uses a certain number of neurons, for example, 64. It processes the input data and passes the learned features to the second layer. Building on the foundation of the first layer, the second LSTM layer further explores long-term dependencies in the data, capturing more complex and long-term time series patterns. It also uses a certain number of neurons, for example, 32. These two LSTM layers work together to fully utilize the temporal characteristics and underlying patterns of time series.

[0075] ③ Input the divided data windows into the constructed two-layer LSTM network for training. During training, use an appropriate loss function, such as mean squared error (MSE), to measure the difference between the model's predicted values ​​and the true values. Using the backpropagation algorithm, the model parameters are continuously adjusted and optimized to reduce the loss function, allowing the model to better learn patterns and regularities in the data. A validation set is used during training to monitor model performance and prevent overfitting.

[0076] ④ After the model training is completed, the data window containing missing values ​​is input into the trained two-layer LSTM network model. The model predicts the missing values ​​according to the learned time series rules and outputs the complete data window after filling, thereby completing the accurate filling of missing values ​​in the entire time series data, providing a complete and accurate data foundation for subsequent data analysis and modeling.

[0077] ⑤After a complete data processing and missing value filling process, a data set containing 1095 sets of valid data was finally constructed.

[0078] Furthermore, the padded data is preprocessed by standardization: the Min-Max normalization method is used to linearly normalize the data of each feature vector to the interval [0,1]. The specific formula is:

[0079]

[0080] Where: Xnew is the data value of each eigenvector after standard normalization processing, X, Xmax, and Xmin are the data value, maximum value, and minimum value of each eigenvector respectively.

[0081] Through normalization processing, the dimensional impact caused by unit differences between features is eliminated, thereby improving data quality.

[0082] Furthermore, the 1,095 sets of valid data after preprocessing and standardization were split into a training set: validation set: test set ratio of 8:1:1, resulting in a dataset with time series characteristics. The training set is used to train the model, continuously adjusting the model parameters to enable the model to effectively learn the patterns in the data. During the model training process, the validation set is used to adjust the model's hyperparameters, such as the learning rate, number of network layers, and number of neurons, and to monitor the model's performance to prevent overfitting. The test set is used to independently evaluate the model's generalization performance and verify its predictive ability and accuracy when faced with new data.

[0083] Step S3, constructing multiple total phosphorus concentration prediction models through simultaneous training of multiple machine learning algorithms, and then determining the predicted value within a preset time period in the future through the prediction models;

[0084] In this embodiment, the following steps are used to construct and apply the total phosphorus concentration prediction model:

[0085] (1) Constructing a total phosphorus concentration prediction model

[0086] The following four machine learning models were used to construct total phosphorus concentration prediction models, and the key parameters of each model were defined:

[0087] ① Random Forest Model (RF)

[0088] Parameters include: number of decision trees, maximum depth, minimum number of sample splits, minimum number of leaf node samples, and number of feature selections.

[0089] ②Support Vector Machine Model (SVM)

[0090] Parameters include: hyperparameters (such as penalty coefficient C), kernel function type (such as RBF, linear kernel), regularization parameters, epsilon insensitive loss function parameters, and training sample ratio.

[0091] ③Gaussian process regression model (GPR)

[0092] Parameters include: kernel function type (such as RBF kernel), noise level, and number of optimizer restarts.

[0093] ④K-nearest neighbor model (KNN)

[0094] Parameters include: k value, distance measurement method (such as Euclidean distance, Manhattan distance), weight type (such as uniform weight, distance weighted), and number of nearest neighbors.

[0095] (2) Model training and parameter optimization

[0096] ① Training set training: The above model is trained on the training set to learn the correlation between input features and total phosphorus concentration at the outlet.

[0097] ② Cross-validation parameter adjustment: Use cross-validation methods (such as K-fold cross-validation) to adjust the basic hyperparameters of each model to avoid overfitting and optimize generalization performance.

[0098] ③ Parameter iterative optimization: By adjusting model parameters multiple times (such as the number of decision trees for random forests, the selection of kernel functions for SVMs, etc.), the prediction accuracy and stability of the model can be gradually improved.

[0099] (3) Model application and prediction

[0100] After the training is completed, four total phosphorus concentration prediction models are obtained. Each model can predict the total phosphorus concentration at the outlet within a preset time period in the future based on the input feature vectors, such as sewage discharge, COD concentration, ammonia nitrogen concentration, ammonia nitrogen discharge, etc.

[0101] Step S4: Under the same data set, the accuracy of each model is compared using the set evaluation index, and the adaptability of each model to the nonlinear characteristics of the sewage treatment field is comprehensively considered to automatically determine the optimal model as the total phosphorus concentration benchmark prediction model;

[0102] In this embodiment, the difference between the predicted value and the true value of each prediction model is calculated based on the test set data and a detailed comparative analysis is performed.

[0103] The performance comparison of each model is shown in Table 1. Observe the data and analyze it:

[0104] The poor performance of the RF and GPR models indicates their limitations in dealing with this data-specific issue, primarily in that they fail to adequately capture trends in areas of high phosphorus concentration.

[0105] In contrast, the SVR model has advantages in all performance indicators of the validation set and the test set due to its excellent nonlinear fitting ability and strong robustness, and its modeling of nonlinear relationships is the most accurate.

[0106] The KNN model is too sensitive to data distribution and cannot provide reliable prediction capabilities in extreme data samples, resulting in its performance being unsuitable for actual scenarios.

[0107] Table 1 Comparison of evaluation indicators of various prediction models on the dataset

[0108]

[0109] Furthermore, the system automatically selects the SVR model as the baseline prediction model, and further optimizes the selection of kernel functions on this basis, and even adapts to more complex feature relationships through multi-kernel combination methods.

[0110] Step S5, combining a swarm intelligence optimization algorithm to deeply tune specific hyperparameters of the benchmark prediction model, monitor the core performance indicators of the model in real time, and evaluate the prediction accuracy, computational efficiency, and robustness by comparing the error changes before and after optimization, and automatically determine the optimized high-performance prediction model for total phosphorus concentration;

[0111] In this example, four optimization algorithms (PSO, ABC, FA, and FWA) were used to tune the hyperparameters of the SVR model, resulting in four optimized prediction models (PSO-SVR, ABC-SVR, FA-SVR, and FWA-SVR). Specifically, these algorithms were used to explore and adjust the key parameters of the SVR model to improve the model's predictive performance and generalization capabilities.

[0112] In order to reduce the uncertainty in the model learning process, each optimized prediction model was run 30 times and the coefficient of determination (R 2 ), mean absolute error (MAE), root mean square error (RMSE), and time consumption to analyze the model's prediction accuracy and time consumption. Model uncertainty is expressed as "mean ± standard deviation," quantifying the range of fluctuation in the model's predictive ability.

[0113] The model performance indicators and uncertainty analysis results are shown in Table 2.

[0114] Table 2 Model performance indicators and uncertainty analysis results

[0115] Model Name <![CDATA[R 2 ]]> MAE (mg / L) RMSE (mg / L) Prediction time (seconds) PSO-SVR 0.9987±0.0029 0.0007±0.0007 0.0009±0.0009 5.0263±4.3303 ABC-SVR 0.9996±0.0005 0.0005±0.0003 0.0006±0.0003 5.5217±3.5084 FA-SVR 0.9989±0.0020 0.0007±0.0006 0.0009±0.0008 13.201±5.9765 FWA-SVR 0.9999±0.0000 0.0002±0.0000 0.0003±0.0000 100.205±23.4949

[0116] Further data analysis revealed that the FWA-SVR model performed best overall and was recommended as the final prediction model for real-time monitoring of total phosphorus concentration at sewage treatment plant outlets. Its prediction accuracy and stability reached high standards. If operational efficiency is also a concern, ABC-SVR is the preferred model, offering the optimal balance between efficiency and accuracy.

[0117] Furthermore, the FWA-SVR model was compared with the traditional mechanism prediction model and the single SVR model, and the results are shown in Table 3.

[0118] Table 3 Performance indicators and uncertainty analysis results of different prediction models

[0119] method <![CDATA[R 2 ]]> RMSE (mg / L) MAE (mg / L) MRE (mg / L) Prediction time (seconds) Mechanistic prediction model 0.6~0.75 0.5~1.5 0.4~1.0 15%~25% 300~1200 SVR model 0.9265 0.0071 0.0092 10.80% 565.81 FWA-SVR model 0.9999 0.0002 0.0003 0.29% 83.20

[0120] During the sewage treatment process, there is a significant non-steady-state nonlinear coupling relationship between the monitoring data at the water inlet and the total phosphorus concentration at the outlet. The traditional mechanism model is constructed based on a simplified mass balance equation and relies on expert experience to preset fixed parameters (such as reaction rate constant, etc.). The FWA-SVR model uses FWA to perform adaptive global optimization of the SVR hyperparameters (such as penalty factor C, kernel width γ), breaking through the limitations of traditional empirical parameter adjustment. Compared with the traditional mechanism model, the R 2 The value increased by 25% (from 0.75 to 0.9999), the RMSE decreased by 99.96% (from 0.5 mg / L to 0.0002 mg / L), and the key error metric, MRE (0.29%), was optimized by nearly two orders of magnitude compared to the traditional model (25%). Furthermore, the model's prediction time was reduced from 1200 seconds using the traditional method to 83.2 seconds, a 13.42-fold increase in efficiency, enabling high-precision, real-time dynamic control.

[0121] While the single SVR model outperforms the traditional mechanism model in terms of accuracy, its kernel function parameters (such as the penalty factor C and kernel width γ) rely on empirical settings, making the model prone to falling into local optimal solutions. Compared with the single SVR model, FWA-SVR further reduces the RMSE by 97.2% (from 0.0071 to 0.0002 mg / L) while maintaining the same generalization ability. The computational time is only 14.7% of that of the grid search parameter-adjusted SVR (from 565.81 seconds to 83.2 seconds), achieving a synergistic leap in accuracy and efficiency.

[0122] Step S6, integrating the high-performance prediction model with the sewage treatment plant programmable logic controller (PLC) and the sewage treatment plant real-time monitoring hardware to dynamically collect real-time data and predict the effluent total phosphorus concentration;

[0123] In this embodiment, the optimized total phosphorus concentration prediction model is integrated into the monitoring system of the sewage treatment plant to achieve real-time prediction of the total phosphorus concentration.

[0124] During deployment, the system integrates with the PLC and online monitoring hardware (such as total phosphorus concentration sensors and flow meters) to ensure real-time access to inlet and outlet data. It predicts the total phosphorus concentration at the outlet at a set frequency (e.g., every 5 minutes or every hour). Based on the predictions, the dosage of the phosphorus removal agent is dynamically adjusted.

[0125] Furthermore, the actual total phosphorus concentration at the outlet is collected in real time and compared with the predicted value to calculate the deviation. This deviation information is fed back to the optimization module, which automatically adjusts model parameters (such as hyperparameters and weights) to continuously improve prediction performance. Through iterative optimization, the model's prediction accuracy and robustness are gradually improved, achieving an intelligent "prediction-control-optimization" closed-loop management.

[0126] Example 2:

[0127] This embodiment provides a sewage treatment plant total phosphorus concentration prediction system based on multiple machine learning models, such as Figure 2 As shown, this system includes the following modules:

[0128] (1) Data acquisition module: real-time acquisition of the inlet and outlet monitoring data of the sewage treatment plant. The inlet monitoring data includes monitoring time, sewage discharge, pH value, COD concentration, COD discharge, ammonia nitrogen concentration, ammonia nitrogen discharge, total nitrogen concentration, total nitrogen discharge, total phosphorus concentration, and total phosphorus discharge. The outlet monitoring data includes total phosphorus concentration.

[0129] The historical monitoring data of the sewage treatment plant are used for model training. The inlet monitoring data are used as the feature vector of the model, and the outlet monitoring data, namely the total phosphorus concentration, are used as the label vector of the model.

[0130] (2) Data preprocessing and standardization module: The historical monitoring data of the sewage treatment plant, i.e., the original monitoring data obtained, is intelligently preprocessed. The data outliers are identified and eliminated through the box plot method. The missing values ​​are filled by the deep learning method based on the long short-term memory network (LSTM). The sliding window mechanism is combined with the long-term dependency characteristics of the network to fully utilize the temporal characteristics and potential laws of the time series to provide a more accurate basis for data analysis and modeling. The long-term dependency memory characteristics of the LSTM model are used to accurately fill in the missing data. For the missing values ​​at the beginning of the sequence, linear interpolation is used to supplement them because they cannot meet the sliding window length to ensure the continuity of the prediction.

[0131] Use Z-Score normalization or Min-Max normalization to normalize the data after removing outliers. Select an appropriate data normalization method based on the distribution of feature data to eliminate the dimensional effects caused by unit differences between features, ensure data quality, and generate a high-quality dataset that reflects the temporal changes of pollutants in the sewage treatment plant.

[0132] The cleaned and preprocessed high-quality dataset is divided into a training set, a validation set, and a test set in chronological order at a ratio of 8:1:1. The training set is used for model training to optimize model parameters, the validation set is used to adjust the model's hyperparameters and monitor model performance, and the test set is used to independently evaluate the model's generalization performance.

[0133] (3) Parallel training and model building of multiple machine learning models: Based on the segmented training set, multiple machine learning algorithms such as random forest (RF), support vector regression (SVM), Gaussian process regression (GPR) and k-nearest neighbor (KNN) were used to build multiple prediction models to simulate the nonlinear mapping relationship between characteristic variables and outlet total phosphorus concentration labels;

[0134] Cross-validation techniques were used to initially optimize the hyperparameters of each model. Performance feedback from the validation set was then used to further fine-tune each model. This included adjusting the number and depth of decision trees for the RF model, the kernel function type and regularization parameters for the SVM model, the kernel function and noise configuration for the GPR model, and the k value and distance metric for the KNN model.

[0135] During the training process, the model gradually learns the mapping relationship between input features and output labels based on the training set. Finally, each trained model generates a predicted value based on the input data, thereby determining the changing trend of total phosphorus concentration in the future preset time period.

[0136] (4) Multiple machine learning model evaluation and optimization module: Multiple total phosphorus concentration prediction models constructed by simultaneous training of multiple machine learning algorithms, namely, random forest (RF), support vector regression (SVM), Gaussian process regression (GPR), and k-nearest neighbor (KNN) models, are applied to the test set respectively, and the difference between the predicted value and the true value is calculated based on the test set data;

[0137] A variety of performance evaluation indicators are used to comprehensively measure the performance of the prediction model, including the coefficient of determination (R 2 ) is used to evaluate the model's ability to explain variables, the mean square error (MSE) and root mean square error (RMSE) are used to evaluate the fluctuation of the sum of squares of the prediction errors, and the mean absolute error (MAE) is used to measure the error margin to ensure the comprehensiveness and scientificity of the evaluation results;

[0138] In order to adapt to the complexity of data in the sewage treatment field and the nonlinear characteristics of pollutant changes, the system automatically selects the model with the best performance as the benchmark prediction model for total phosphorus concentration, providing a basis for subsequent hyperparameter optimization and system deployment.

[0139] (5) Intelligent optimization module for total phosphorus concentration prediction model: The total phosphorus concentration benchmark prediction model is deeply optimized for key hyperparameters using a variety of swarm intelligence optimization algorithms, including particle swarm optimization (PSO), artificial bee colony algorithm (ABC), firefly algorithm (FA), and firework algorithm (FWA);

[0140] During the parameter tuning process, it is necessary to monitor the core performance indicators of the model in real time and simultaneously evaluate the computational efficiency and robustness. The core performance indicators include: coefficient of determination (R 2 ), mean square error (MSE), root mean square error (RMSE), mean absolute error (MAE);

[0141] Computational efficiency is quantified by recording the actual running time of the model processing a large sample data set. Robustness is verified through repeated experimental analysis: the model is run multiple times in different data scenarios and the stability of its output is analyzed. The test results are presented in the form of "mean ± standard deviation" to intuitively reflect the fluctuation range of the indicator.

[0142] Compare the performance of the optimized prediction models in terms of prediction accuracy, computational efficiency, and error control, and record the changes in model performance before and after optimization;

[0143] By comparing the error changes between the optimized model and the original benchmark model, the optimization algorithm was verified to be significant in improving model prediction accuracy, reducing error fluctuations and enhancing model robustness. Finally, the optimized high-performance prediction model for total phosphorus concentration was determined, providing a better solution for practical engineering applications.

[0144] (6) Intelligent management and dynamic deployment module of total phosphorus concentration prediction model: The optimized high-performance prediction model of total phosphorus concentration is deployed to the real-time monitoring system of the sewage treatment plant to realize real-time prediction of total phosphorus concentration in sewage effluent. The system is deeply integrated with the programmable logic controller (PLC) control unit of the sewage treatment plant and the sewage treatment plant online monitoring hardware equipment. It uses the set monitoring frequency (such as minute level, hour level) to dynamically collect real-time data and predict the total phosphorus concentration in effluent in advance based on the high-performance prediction model of total phosphorus concentration;

[0145] When the predicted results deviate from the preset emission standards, the system intelligently adjusts the dosage of the phosphorus removal agent according to the predicted concentration range, thereby achieving closed-loop control of the total phosphorus concentration. At the same time, the deviation between the actual total phosphorus concentration measured at the outlet and the predicted value is fed back to the optimization module, and the model weights and parameters are dynamically recalibrated through deviation backcalculation to continuously improve the prediction accuracy and regulation effect until the long-term operation stability requirements are met, realizing a closed-loop system of "prediction-control-optimization".

[0146] Although the embodiments of the present invention have been disclosed as above, they are not limited to the applications listed in the description and implementation methods. They can be fully applied to various fields suitable for the present invention. For those familiar with the art, for those of ordinary skill in the art, various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. Therefore, without departing from the general concept defined by the claims and the scope of equivalents, the present invention is not limited to specific details.

Claims

1. A method for predicting total phosphorus concentration in a sewage treatment plant based on multiple machine learning models, characterized in that: The following steps are involved: Step S1: monitor and obtain sewage treatment plant data, wherein the water inlet monitoring data is used as a feature vector and the water outlet monitoring data is used as a label vector; Step S2: performing intelligent preprocessing on the sewage treatment plant data, standardizing the data, and dividing the data set into a training set, a validation set, and a test set in chronological order; Step S3: constructing multiple total phosphorus concentration prediction models through simultaneous training of multiple machine learning algorithms, and then determining the predicted value within a preset time period in the future through the prediction models; Step S4: Using the same training set, validation set, and test set, the accuracy of each model is compared using multiple performance evaluation indicators, and the adaptability of each model to the nonlinear characteristics of the sewage treatment field is comprehensively considered to automatically determine the optimal model as the benchmark prediction model for total phosphorus concentration; Step S5: deeply tune the specific hyperparameters of the benchmark prediction model using a swarm intelligence optimization algorithm, monitor the core performance indicators of the model in real time, and evaluate the prediction accuracy, computational efficiency, and robustness by comparing the error changes before and after optimization, and automatically determine the optimized high-performance prediction model for total phosphorus concentration; Step S6: Integrate the high-performance prediction model with the sewage treatment plant programmable logic controller (PLC) and the sewage treatment plant real-time monitoring hardware to dynamically collect real-time data and predict the total phosphorus concentration in the effluent.

2. The method for predicting total phosphorus concentration in a sewage treatment plant according to claim 1, wherein: The water inlet monitoring data in step S1 includes monitoring time, sewage discharge, pH value, COD concentration, COD discharge, ammonia nitrogen concentration, ammonia nitrogen discharge, total nitrogen concentration, total nitrogen discharge, total phosphorus concentration, and total phosphorus discharge. The water outlet monitoring data is total phosphorus concentration.

3. The method for predicting total phosphorus concentration in a sewage treatment plant according to claim 1, wherein: The method for intelligent data preprocessing in step S2 is as follows: identifying and eliminating outliers through the box plot method, using the long short-term memory network LSTM combined with the sliding window to fill the missing time series data, and using linear interpolation to supplement the missing values ​​at the beginning of the sequence that cannot meet the sliding window length to ensure the continuity of the prediction; The method for normalizing the data is: using Z-Score normalization or Min-Max normalization method to normalize the data after outlier processing, selecting an appropriate data normalization method according to the distribution of characteristic data, eliminating dimensional differences, and generating a high-quality data set reflecting the temporal changes of pollutants in the sewage treatment plant; The method for dividing the data set into a training set, a validation set and a test set is: dividing the high-quality data set into a training set, a validation set and a test set in a ratio of 8:1:1 in chronological order.

4. The method for predicting total phosphorus concentration in a sewage treatment plant according to claim 1, wherein: The specific method of step S3 includes: Multiple prediction models are constructed for the training set using multiple machine learning algorithms to simulate the nonlinear mapping relationship between the feature vector and the total phosphorus concentration label vector of the outlet, respectively. The multiple machine learning algorithms include: random forest RF, support vector regression SVM, Gaussian process regression GPR and k-nearest neighbor KNN; Cross-validation techniques were used to initially optimize the hyperparameters of each model. Performance feedback from the validation set was then used to further fine-tune each model. This fine-tuning included adjusting the number and depth of decision trees for the RF model, the kernel function type and regularization parameter for the SVM model, the kernel function and noise configuration for the GPR model, and the k value and distance metric for the KNN model. Each prediction model generates a total phosphorus concentration prediction value based on the test set data through training and learning.

5. The method for predicting total phosphorus concentration in a sewage treatment plant according to claim 1, wherein: The specific method of step S4 includes: Calculate the difference between the predicted value and the true value of each prediction model in step S3 based on the test set data; A variety of performance evaluation indicators are used to comprehensively measure the prediction model, and the determination coefficient R is used to 2 Evaluate the explanatory variables of the model, use the mean square error (MSE) and root mean square error (RMSE) to evaluate the fluctuation of the sum of squares of the prediction errors, and use the mean absolute error (MAE) to measure the error magnitude; The model with the best performance is automatically screened according to the evaluation index as the total phosphorus concentration benchmark prediction model.

6. The method for predicting total phosphorus concentration in a sewage treatment plant according to claim 1, wherein: The swarm intelligence optimization algorithms in step S5 include: particle swarm optimization algorithm PSO, artificial bee colony algorithm ABC, firefly algorithm FA and firework algorithm FWA; The core performance indicators include: determination coefficient R 2 , mean square error MSE, root mean square error RMSE, mean absolute error MAE; The computational efficiency is quantified by the actual running time of the model processing a large sample data set, and the robustness is verified through repeated experimental analysis: the model is run multiple times in different data scenarios, and the stability of its output results is analyzed. The test results are presented in the form of mean ± standard deviation, which intuitively reflects the fluctuation range of the indicator.

7. The method for predicting total phosphorus concentration in a sewage treatment plant according to claim 1, wherein: The total phosphorus concentration of the effluent predicted in step S6 is used to control the dosage of the phosphorus removal agent.

8. A sewage treatment plant total phosphorus concentration prediction system based on multiple machine learning models, characterized in that: Includes the following modules: Automatic data acquisition module: monitors and obtains sewage treatment plant data, where the inlet monitoring data is used as the feature vector and the outlet monitoring data is used as the label vector; Data preprocessing and standardization module: intelligently preprocess the sewage plant data, standardize the data, and divide the data set into training set, validation set and test set in chronological order; Multiple machine learning model parallel training and construction modules: Multiple total phosphorus concentration prediction models are constructed through simultaneous training of multiple machine learning algorithms, and then the prediction values ​​within a preset time period in the future are determined by the prediction models; Multiple machine learning model evaluation and optimization modules: Using the same training, validation, and test sets, the module uses multiple performance evaluation metrics to compare the accuracy of each model. By comprehensively considering their adaptability to the nonlinear characteristics of wastewater treatment, the module automatically determines the optimal model as the benchmark prediction model for total phosphorus concentration. Total phosphorus concentration prediction model intelligent optimization module: This module uses a swarm intelligence optimization algorithm to deeply tune specific hyperparameters of the baseline prediction model, monitors the model's core performance indicators in real time, and compares the error changes before and after optimization to evaluate prediction accuracy, computational efficiency, and robustness. It automatically determines the optimized high-performance total phosphorus concentration prediction model. Intelligent management and dynamic deployment module of the total phosphorus concentration prediction model: The high-performance total phosphorus concentration prediction model is deployed to the sewage plant real-time monitoring system, and integrated with the sewage plant programmable logic controller (PLC) and the sewage plant real-time monitoring hardware to dynamically collect real-time data and predict the effluent total phosphorus concentration.

9. A method for controlling the total phosphorus concentration in a sewage treatment plant using the total phosphorus prediction model according to claim 1, characterized in that: The following steps are involved: Integrating the high-performance prediction model in step S5 with the sewage treatment plant programmable logic controller (PLC) and the sewage treatment plant real-time monitoring hardware; Dynamically collect real-time data using a set detection frequency, and predict the total phosphorus concentration in the effluent based on the high-performance prediction model; When the predicted results deviate from the preset emission standards, the dosage of phosphorus removal agents is dynamically and intelligently adjusted according to the change in the predicted concentration. At the same time, the predicted deviation is fed back to the optimization module, and the model weights and parameters are dynamically recalibrated through deviation back calculation to continuously improve the prediction accuracy and regulation effect until the long-term operation stability requirements are met, realizing the "prediction-control-optimization" closed-loop management.

10. The method according to claim 9, characterized in that The monitoring frequency includes minute level and hour level.

Citation Information

Cited By

  • Sewage treatment strategy adaptive optimization method and system based on reinforcement learning

    CN121020687A

  • Recursive extreme learning machine-based effluent total phosphorus concentration prediction method in sewage treatment process

    CN121191632A

  • Municipal sludge yield prediction method and system based on multi-dimensional data driving

    CN121303449A

  • Suspension chain cleaning control method based on interval optimization

    CN121523016A