Water quality data prediction method based on hybrid model and intelligent monitoring system
By constructing a combined combination model of feature extraction-decomposition-optimization, combined with intelligent cleaning strategies and improved NPDAO algorithm, the problem of data incompleteness of existing water quality data prediction methods and poor training effect of deep learning combination model is solved, and efficient, real-time and accurate water quality data prediction and intelligent monitoring are achieved.
Patent Information
- Application Number
- CN202510199126.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing data incompleteness or inaccuracy of existing water quality data prediction methods lead to large prediction deviations and poor interpretation, and the deep learning combination model is not trained when there is little data.
Using a water quality data prediction method based on a mixed model, a combined model with feature extraction-decomposition-optimization coupled combination model is constructed through model input feature hierarchical depth extraction, model parameter optimization and multi-model combination, and combined with intelligent cleaning strategies and improved NPDAO algorithms to optimize model parameters and data processing.
It significantly improves the accuracy, stability and reliability of the water quality prediction model, reduces the risk of environmental pollution, optimizes resource allocation, and realizes intelligent monitoring and management of water quality data.
Smart Images

Figure CN119990461A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of water quality data prediction and monitoring in a plant area, and aims to provide a water quality data prediction method and an intelligent monitoring system. Background Art
[0002] As water shortage and water pollution become increasingly serious worldwide, water quality data monitoring in factories is particularly important, not only related to water treatment and discharge, but also of great significance for protecting the environment and public health and promoting the use of recycled water. Accurate and real-time water quality data prediction can provide a scientific basis for factories to optimize water treatment processes, ensure that the treated water quality meets national or local emission standards, and reduce harm to the environment and human health.
[0003] Although a variety of water quality data predictions have been developed, there are some limitations to water quality data prediction methods. Traditional methods such as multivariate regression analysis rely on historical data, but the prediction bias is large and the interpretability is poor due to incomplete or inaccurate data. Although time series and neural network methods are used in water quality prediction, model complexity and data quality and quantity issues may limit their accuracy. In contrast, hybrid prediction models enhance the robustness and generalization ability of the model by combining the advantages of multiple models, thereby improving prediction accuracy. In addition, the use of optimization algorithms to adjust model parameters can further improve the accuracy and robustness of hybrid models in water quality data prediction.
[0004] In order to improve the accuracy and reliability of water quality data prediction, overcome the limitations of existing methods and deal with problems such as insufficient data from newly built collection points, solve the problem of poor training results due to insufficient data in deep learning combination models, so as to be more suitable for plant water quality prediction and intelligent systems. Summary of the invention
[0005] Purpose of the invention: The purpose of this invention is to provide an efficient, real-time and accurate water quality data prediction method and system. Through the combination of model input feature layered deep extraction, model parameter optimization and multi-model combination, a combined model of feature extraction-decomposition-optimization coupling is constructed, thereby comprehensively improving the accuracy of the water quality prediction model; an intelligent decision-making and system maintenance module is established to monitor abnormal data, realize intelligent decision-making and regularly upgrade the system, aiming to solve the limitations of existing water quality data prediction methods and significantly enhance the accuracy, stability and reliability of predictions. It can not only significantly reduce environmental pollution risks and optimize resource allocation, but also realize intelligent monitoring and management of water quality data through real-time monitoring and early warning, system detection, optimization and improvement by intelligent means, thereby promoting environmental protection to a higher level of intelligent and refined development, and providing technical support for water quality supervision of chemical plants, specifying scientific and reasonable sewage discharge and treatment strategies, environmental protection and sustainable development of human society.
[0006] Technical solution: The present invention discloses a water quality data prediction method based on a hybrid model, comprising:
[0007] Step 1: Collect historical water quality data set D0. The number of D0 water quality types includes but is not limited to pH, temperature, turbidity, suspended solids, dissolved oxygen, ammonia nitrogen, total nitrogen, and total phosphorus;
[0008] Step 2: Use the intelligent cleaning strategy DCS to clean the water quality data set D0 to remove outliers and fill missing values in the water quality data, perform Min-Max standardization on the processed water quality data to unify the data scale, and use the principal component analysis method to select time series characteristic water quality data with strong correlation to form the water quality characteristic data set D1;
[0009] Step 3: The water quality characteristic dataset D1 is decomposed by the AMEMD composite method to obtain the water quality dataset D2 and divide it;
[0010] Step 4: Construct a water quality intelligent hybrid prediction model. The prediction results of the water quality intelligent hybrid prediction model are mixed by weighted average of the prediction results of the GT-GAN-CNN-Bi-LSTM-Attention model and the XGBoost model. The steps of the GT-GAN-CNN-Bi-LSTM-Attention model are as follows:
[0011] Step 4.1: Design the generative adversarial network model GT-GAN to expand the water quality dataset;
[0012] Step 4.2: Use convolutional neural network (CNN) to process the expanded water quality data and extract local features in the time series;
[0013] Step 4.3: The features extracted by CNN are further processed through the bidirectional long short-term memory network Bi-LSTM to capture the long-term dependencies of the time series;
[0014] Step 4.4: Introduce the attention mechanism to enable the model to identify and focus on the key features in the time series;
[0015] Step 5: Use the improved NPDAO algorithm to optimize the hyperparameter combination P0 of the water quality intelligent hybrid prediction model constructed in step 4, and train and predict the water quality data set divided in step 3. 2 , root mean square error RMSE and mean absolute error MAE to evaluate the model performance and obtain the optimal hyperparameter combination of the hybrid model;
[0016] Step 6: Use the GT-GAN-CNN-Bi-LSTM-Attention and XGBoost intelligent hybrid prediction model with the optimal parameter combination to predict real-time water quality data.
[0017] Furthermore, in step 2, the specific steps of the intelligent cleaning strategy DCS are as follows:
[0018] Step 2.1: Use the box plot method to preliminarily identify outliers in the D0 data set, and then use the Z-score method to verify the outliers. Calculate the Z-score of each data point. If the absolute value of the Z-score is greater than the threshold determined by the 3σ rule, the data point is marked as the final outlier. The Z-score formula is as follows:
[0019] For each data point d i , calculate its Z-score:
[0020]
[0021] Where: μ represents the mean value of the data set D0; σ represents the standard deviation;
[0022] Step 2.2: Mark the outliers identified in step 2.1 as missing values;
[0023] Step 2.3: Use k-NN to interpolate the missing values marked in step 2.2, and use the local similarity of the D0 data set to find neighbors similar to the missing data points through the Mahalanobis distance for interpolation.
[0024] Furthermore, the specific steps of the AMEMD composite decomposition method in step 3 are:
[0025] Step 3.1: Perform EMD decomposition on the water quality data D1 to obtain K IMF components;
[0026] Step 3.2: Calculate the Pearson correlation coefficient for the K IMF components obtained in step 3.1. The calculation steps are as follows:
[0027] Step 3.2.1: For K IMF components, first, for any two IMF components i and IMF j ,1≤i<j≤K, first calculate its covariance Cov(IMF i ,IMF j ), whose formula is:
[0028]
[0029] Where: t represents the index in the water quality data sequence t=1,2,3,…,N; N represents the length of the water quality data; IMFi,t ,IMF j,t IMF i ,IMF j The value at t; IMF i ,IMF j The mean of
[0030] Step 3.2.2: Calculate IMF i and IMF j The standard deviation of σIMF i ,σIMF j ;
[0031] Step 3.2.3: Calculate IMF using Pearson correlation coefficient formula i and IMF j The correlation coefficient p i,j , the formula is:
[0032]
[0033] Step 3.2.4: Repeat the above steps to calculate the Pearson correlation coefficient of the IMF components. correlation coefficients;
[0034] Step 3.2.5: Construct the correlation coefficient matrix R, which is represented as follows:
[0035]
[0036] Step 3.3: Aggregate the IMF components with strong correlation using the matrix R obtained in step 3.2. The steps are as follows:
[0037] Step 3.3.1: Create a Boolean array used of length K, with all initial values set to False, indicating that the K IMF components obtained by EMD decomposition are not aggregated;
[0038] Step 3.3.2: Considering the selection of the maximum and optimal Pearson correlation coefficient, a greedy algorithm is used to optimize to ensure that the maximum Pearson correlation coefficient and its corresponding index (i, j) are found from the matrix R at each step;
[0039] Step 3.3.3: Set used[i], used[j] to True, indicating that the two components have been aggregated;
[0040] Step 3.3.4: IMF i ,IMF jAggregation is performed and the aggregated results are stored in a new array D2. The rows and columns related to i and j in the matrix R are set to 0 to prevent these components from being selected again.
[0041] Step 3.3.5: Repeat steps 3.3.2 to 3.3.4 until all IMF components are aggregated or there are no aggregateable components that meet the stopping condition, and obtain the AMEMD composite decomposition data set D2. The specific stopping condition is:
[0042]
[0043] Where: Υ is a very small threshold used to determine whether it is large enough for aggregation.
[0044] Furthermore, in step 5, the hyperparameter P0 includes the number of XGBoost trees, estimators, and learning rate; the number of filters, filter size, and pool size in the convolutional layer of CNN; the number of LSTM units of Bi-LSTM, dropout rate, and learning rate; and the number of hidden units and dropout rate in the Attention mechanism.
[0045] Furthermore, in step 5, the specific steps of optimizing the improved NPDAO algorithm are as follows:
[0046] Step 5.1: Input the parameters of the NPDOA algorithm, and at the same time input the hyperparameter group P0 to be optimized of the hybrid model of the GT-GAN-CNN-Bi-LSTM-Attention model and the XGBoost model, and use the Tent chaotic map to initialize P0;
[0047] Step 5.2: Use the attractor trend strategy to approach the excellent hybrid model parameter group P, couple the interference strategy to increase randomness to promote exploration, and the information projection strategy to control the information flow, so as to achieve the transition from extensive exploration to in-depth utilization;
[0048] Step 5.3: Use the competitive update mechanism to update P;
[0049] Step 5.4: Determine whether the maximum number of evaluations has been reached. If yes, proceed to the next step. If no, return to step 5.2.
[0050] Step 5.5: Select the optimal parameter set P through the memory strategy max , output P max and the corresponding target value;
[0051] Furthermore, the improved NPDOA algorithm uses Tent chaotic map initialization as follows:
[0052] Generating chaotic sequences based on Tent The process is as follows:
[0053]
[0054] Where: k is the number of parameter groups P0 of the hybrid model to be optimized; t is the number of evaluations, and u is a random number
[0055] Combined with chaotic sequence The process of further generating the parameter group of the hybrid model is as follows:
[0056]
[0057] in: The worst and best parameter sets of the neural hybrid model, respectively.
[0058] Furthermore, the competitive update mechanism is adopted in step 5.3, and the steps are as follows:
[0059] Step 5.3.1: Considering that a single evaluation index cannot more comprehensively evaluate the current parameter combination P = {ξ1,ξ2,...,ξ O}, so MAE, MAPE, RMSE multiple evaluation indicators are selected to form a multi-objective optimization objective function f() to measure the parameter combination P = {ξ1,ξ2,...,ξ O}, where O is the number of groups;
[0060] Step 5.3.2: According to f(), select the top 10% as an elite hybrid model parameter group individual set P elite ={ξ e1 ,ξ e2 ,...,ξ eQ}, Q is the number of elite individuals;
[0061] Step 5.3.3: For each non-elite mixture model parameter group individual ξ h , h>Q, randomly select elite individuals ξ eq , As a competitor, if f(ξ h )<f(ξ eq ), then ξ h Replacement eq .
[0062] Furthermore, in step 5.5, the optimal parameter group is selected by the memory strategy, and the steps are as follows:
[0063] Step 5.5.1: Construct a memory bank M to store the historical optimal hybrid model parameter set, P best is the current optimal hybrid model parameter set, P M,best is the optimal hybrid model parameter set in the memory bank;
[0064] Step 5.5.2: Update the memory M, if f(P best )<f(P M,best ), then M=P best ;
[0065] Step 5.5.3: Use M to search based on the following:
[0066] P new =Search(P,M) (8)
[0067] Where: P is the current hybrid model parameter group, Search is the search guidance operation, P new is the newly generated hybrid model parameter group.
[0068] The present invention also discloses a water quality data intelligent monitoring system based on a hybrid model, comprising: a data acquisition module, a data and storage module, a data transmission module, and a prediction and monitoring cloud platform;
[0069] The data acquisition module collects historical water quality data sets and stores them in the data and storage module, and transmits them to the prediction and monitoring cloud platform through the data transmission module;
[0070] The prediction and monitoring cloud platform is provided with the water quality data prediction method based on the hybrid model as above;
[0071] The intelligent decision-making and system maintenance module is used to make decisions and perform maintenance based on the predicted water quality data obtained by the prediction and monitoring cloud platform.
[0072] Preferably, it also includes an intelligent decision-making and system maintenance module, which displays the prediction results in real time, intelligently analyzes historical data to check whether the water quality data meets national standards, and provides intelligent decision-making and optimization suggestions.
[0073] Beneficial effects:
[0074] The prediction method proposed in the present invention adopts the intelligent cleaning strategy DCS, which can effectively process the outliers and missing values in the water quality data, significantly improve the quality of the water quality data, reduce the impact of outliers and missing values, and thus improve the prediction accuracy. In addition, the AMEMD decomposition is used to effectively reduce the complexity of the water quality data, extract components of different frequencies and scales of the water quality data, simplify the data analysis and processing process, and facilitate further analysis and processing of water quality prediction, thereby further improving the water quality prediction accuracy.
[0075] The present invention proposes a hybrid prediction model of GT-GAN-CNN-Bi-LSTM-Attention and XGBoost. In the hybrid model, GT-GAN is used to expand and improve the quality of a water quality data set, providing rich samples for the training of a CNN-Bi-LSTM-Attention deep combination model, thereby improving the model training effect; the combination of CNN and Bi-LSTM effectively extracts local and global features of a time series, while the attention mechanism enables the model to focus on key features and improves prediction accuracy; the integrated learning of XGBoost further improves the prediction performance; finally, by weighted averaging the prediction results of the two models, an accurate and stable water quality data is output; by combining GT-GAN, CNN, Bi-LSTM, Attention and XGBoost, the water quality data can be more fully understood through the synergistic effect, so that the model can cope with various complex data distributions and changes, enhance the applicability of the model, and thus improve the accuracy and robustness of water quality data prediction.
[0076] The present invention uses an improved NPDOA optimization algorithm to obtain the optimal parameter combination of the model for the hybrid prediction model. The NPDOA algorithm combines chaotic mapping, competitive update mechanism and memory strategy, enhances global search capability and improves diversity; accelerates the update and propagation of excellent parameter combinations and improves convergence speed; retains the information of the historical optimal parameter combination and enhances stability and robustness. This further improves the accuracy of water quality data prediction.
[0077] The system proposed in this invention can display the prediction results in real time, conduct in-depth analysis based on these results, provide intelligent decision-making solutions, and ensure the continuous updating and maintenance of the system through fault diagnosis and online update technology. The system provides users with an efficient, stable and easy-to-maintain platform through real-time and accurate prediction, intelligent decision support, fault diagnosis and online update. BRIEF DESCRIPTION OF THE DRAWINGS
[0078] Figure 1 The figure shows a water quality data intelligent monitoring system diagram proposed by the present invention;
[0079] Figure 2 Shown is a flow chart of water quality data prediction proposed by the present invention;
[0080] Figure 3 The figure shows the GT-GAN-CNN-Bi-LSTM-Attention and XGBoost hybrid model proposed in the present invention;
[0081] Figure 4 Shown is a specific implementation flow chart of the present invention. DETAILED DESCRIPTION
[0082] The technical solution of the present invention is further described in detail below in conjunction with the accompanying drawings.
[0083] like Figure 1 As shown, an embodiment of the present invention is a hybrid model water quality data intelligent monitoring system, including: a data acquisition module, a data and storage module, a data transmission module, a prediction and monitoring cloud platform, and an intelligent decision-making and system maintenance module. The data acquisition module collects historical water quality data sets and stores them in the data and storage module, and transmits them to the prediction and monitoring cloud platform through the data transmission module; the prediction and monitoring cloud platform is provided with the above-mentioned hybrid model-based water quality data prediction method; the intelligent decision-making and system maintenance module is used to make decisions and perform maintenance according to the predicted water quality data obtained by the prediction and monitoring cloud platform.
[0084] The intelligent decision-making and system maintenance module specifically includes: real-time display of prediction results and intelligent analysis of historical data to check whether the water quality data meets national standards, and provide intelligent decision-making and optimization suggestions; secondly, this module regularly upgrades system functions and performance according to demand and technological development, and introduces the latest monitoring technology and prediction methods to improve the overall performance of the system.
[0085] like Figure 3 As shown, the prediction flow chart proposed by the present invention, a water quality data prediction method of a hybrid model, specifically includes the following steps:
[0086] Step 1: Collect historical water quality data sets, including but not limited to pH, temperature, turbidity, suspended solids, dissolved oxygen, ammonia nitrogen, total nitrogen, and total phosphorus.
[0087] Step 2: Use the intelligent cleaning strategy DCS to clean the water quality data set to remove outliers and fill missing values in the water quality data. Then, perform Min-Max standardization on the processed water quality data to unify the data scale. Then, use the principal component analysis method to select time series characteristic water quality data with strong correlation to form a water quality characteristic data set.
[0088] The specific steps of the intelligent cleaning strategy DCS are as follows:
[0089] Step 2.1: Use the box plot method to preliminarily identify outliers in the D0 data set, and then use the Z-score method to further verify these preliminarily identified outliers. Specifically, calculate the Z-score of each data point. If the absolute value of the Z-score is greater than the threshold determined by the 3σ rule, these data points are marked as final outliers. The Z-score formula is as follows:
[0090] For each data point d i , calculate its Z-score:
[0091]
[0092] Where: μ represents the mean value of the data set D0; σ represents the standard deviation.
[0093] Step 2.2: Mark the outliers identified in step 2.1 as missing values.
[0094] Step 2.3: Use k-NN to interpolate the missing values marked in step 2.2. This method uses the local similarity of the D0 data set and uses the Mahalanobis distance to find neighbors similar to the missing data points for interpolation, thereby better retaining the local characteristics of the water quality data.
[0095] Step 3: Obtain water quality data through the AMEMD composite decomposition method and divide it.
[0096] The specific steps of the AMEMD composite decomposition method are:
[0097] Step 3.1: Perform EMD decomposition on the water quality data D1 to obtain K IMF components.
[0098] Step 3.2: Calculate the Pearson correlation coefficient for the K IMF components obtained in step 3.1. The Pearson correlation coefficient mainly measures the strength of the linear relationship between two IMF components. The calculation steps are as follows:
[0099] a: For K IMF components, first, for any two IMF components i and IMF j (1≤i<j≤K), first calculate its covariance Cov(IMF i ,IMF j ), whose formula is:
[0100]
[0101] Where: t represents the index in the water quality data sequence t=1,2,3,…,N; N represents the length of the water quality data; IMF i,t ,IMF j,t IMF i ,IMF j The value at t; IMF i ,IMF j The mean of .
[0102] b: Calculate IMF i and IMF j The standard deviation of σIMF i ,σIMF j .
[0103] c: Calculate IMF based on Pearson correlation coefficient formula i and IMF j The correlation coefficient p i,j , the formula is:
[0104]
[0105] d: Repeat the above steps to calculate the Pearson correlation coefficient of the IMF components. A correlation coefficient.
[0106] e: Construct the correlation coefficient matrix R, where R is expressed as follows:
[0107]
[0108] Step 3.3: Aggregate the IMF components with strong correlation using the matrix R obtained in step 3.2. The steps are as follows:
[0109] a: Create a Boolean array used of length K, with all initial values being False, indicating that the K IMF components obtained by EMD decomposition are not aggregated.
[0110] b: Considering the selection of the maximum and optimal Pearson correlation coefficient, a greedy algorithm is used for optimization to ensure that the maximum Pearson correlation coefficient and its corresponding index (i, j) are found from the matrix R obtained in step 3.2 at each step.
[0111] c: Set used[i], used[j] to True, indicating that these two components are aggregated.
[0112] d: IMF i ,IMF j Aggregation is performed and the aggregated results are stored in a new array D2. The rows and columns related to i and j in the matrix R are set to 0 to prevent these components from being selected again.
[0113] e: Repeat steps bd until all IMF components are aggregated or there are no aggregateable components that meet the stopping condition, and obtain the AMEMD composite decomposition data set D2. The specific stopping condition is:
[0114]
[0115] Where: Υ is a very small threshold used to determine whether it is large enough for aggregation.
[0116] Step 4: Build a water quality intelligent hybrid prediction model, which is a mixture of GT-GAN-CNN-Bi-LSTM-Attention and XGBoost by weighted averaging.
[0117] GT-GAN-CNN-Bi-LSTM-Attention model, combined Figure 3 As shown, the specific steps are as follows:
[0118] Step 4.1: Considering that the data obtained by the newly constructed data collection points are incomplete and limited in quantity, and the prediction effect of the CNN-Bi-LSTM-Attention deep combination model is better trained with a large amount of water quality data, the GT-GAN generative adversarial network is designed to expand the water quality data set, solve the problem of imperfect data, increase sample diversity, enable the model to better learn the laws of water quality data, and effectively improve the prediction accuracy and robustness of the model.
[0119] Step 4.2: Use the convolutional neural network (CNN) to process the expanded water quality data and extract local features in the time series.
[0120] Step 4.3: The features extracted by CNN are further processed by a bidirectional long short-term memory network Bi-LSTM to capture the long-term dependencies of the time series.
[0121] Step 4.4: Introduce the attention mechanism to enable the model to identify and focus on the key features in the time series, further improving the accuracy of water quality prediction.
[0122] Step 5: In order to improve the prediction accuracy, the improved NPDAO algorithm is used to optimize the hyperparameter combination P0 of the water quality intelligent hybrid prediction model constructed in step 4 and train and predict the water quality data set divided in step 3. The determination coefficient R is used 2 , root mean square error RMSE, and mean absolute error MAE to evaluate the model performance and obtain the optimal hyperparameter combination of the hybrid model.
[0123] The hyperparameter combination P0 includes the number of XGBoost trees, estimators, and learning rate; the number of filters, filter size, and pool size in the convolutional layer of CNN; the number of LSTM units in Bi-LSTM, dropout rate, and learning rate; and the number of hidden units and dropout rate in the Attention mechanism.
[0124] The specific steps of the improved NPDAO algorithm optimization are as follows:
[0125] Step 5.1: Input the parameters of the NPDOA algorithm, and at the same time input the hyperparameter group P0 to be optimized for the GT-GAN-CNN-Bi-LSTM-Attention and XGBoost hybrid model, and use the Tent chaotic map to initialize P0.
[0126] Step 5.2: Use the attractor trend strategy to approach the excellent hybrid model parameter group P, couple the interference strategy to increase randomness to promote exploration, and the information projection strategy to control the information flow to achieve the transition from extensive exploration to in-depth utilization.
[0127] Step 5.3: Use the competitive update mechanism to update P.
[0128] Step 5.4: Determine whether the maximum number of evaluations has been reached. If yes, proceed to the next step; otherwise, return to step 5.2.
[0129] Step 5.5: Select the optimal parameter set P through the memory strategy max , output P max and the corresponding target value.
[0130] Step 6: Use the GT-GAN-CNN-Bi-LSTM-Attention and XGBoost intelligent hybrid prediction model with the optimal parameter combination to predict real-time water quality data.
[0131] Improve the NPDAO algorithm. The improvement strategy is as follows:
[0132] Strategy 1: Chaotic mapping strategy Tent has randomness, ergodicity and initial value sensitivity, which makes the algorithm converge faster. Chaotic sequence is generated based on Tent. The process is as follows:
[0133]
[0134] Where: k is the number of parameter groups P0 of the hybrid model to be optimized; t is the number of evaluations, and u is a random number
[0135] Combined with chaotic sequence The process of further generating the parameter group of the hybrid model is as follows:
[0136]
[0137] in: The worst and best parameter sets of the neural hybrid model, respectively.
[0138] Strategy 2: Introducing competitive updates can enable the NPDOA algorithm to more effectively maintain the diversity of parameter groups and improve the quality of optimization. It helps the NPDOA algorithm avoid premature convergence and maintain a balance between global search and local search. The steps are as follows:
[0139] a: Considering that a single evaluation index cannot more comprehensively evaluate the current parameter combination P = {ξ1,ξ2,...,ξ O}, so MAE, MAPE, RMSE multiple evaluation indicators are selected to form a multi-objective optimization objective function f() to measure the parameter combination P = {ξ1,ξ2,...,ξ O}, where O is the number of groups.
[0140] b: According to f(), select the top 10% as an elite hybrid model parameter group individual set P elite ={ξ e1 ,ξ e2 ,...,ξ eQ}, Q is the number of elite individuals.
[0141] c: For each non-elite mixed model parameter group individual ξ h (h>Q), randomly select elite individuals ξ eq , As a competitor, if f(ξ h )<f(ξ eq ), then ξ h Replacement eq .
[0142] Strategy 3: Introducing the memory strategy enables the NPDOA algorithm to effectively use the historical optimal parameter combination hybrid model parameter group to improve the algorithm's search efficiency and algorithm performance. The steps are as follows:
[0143] a: Construct a memory library M to store the historical optimal hybrid model parameter group, P best is the current optimal hybrid model parameter set, P M,best is the optimal set of hybrid model parameters in the memory bank.
[0144] b: Update the memory M, if f(P best )<f(P M,best ), then M=P best .
[0145] c: Use M to search based on the following:
[0146] P new =Search(P,M) (8)
[0147] Where: P is the current hybrid model parameter group, Search is the search guidance operation, P new is the newly generated hybrid model parameter group.
[0148] according to Figure 4 The specific implementation flow chart is as follows: First, the water quality in the data acquisition module is sent to the sensor sampling unit through the pipeline by a water pump, and the sensor converts the detection signal into an electrical signal and transmits it to the data and storage module; then, the data and storage module amplifies, filters and digitizes the electrical signal output by the sensor, and transmits the data to the prediction and monitoring cloud platform through the data transmission module; then, in the prediction and monitoring cloud platform, after the data enters the storage layer, it is cleaned by DCS and normalized by Min-Max through the data analysis layer, and the time series feature data with strong correlation is selected by principal component analysis, and the data set is decomposed and divided by AMEMD; a GT-GAN-CNN-Bi-LSTM-Attention and XGBoost hybrid prediction model is constructed and the improved NPDAO algorithm is used to optimize the hyperparameters of the model; real-time water quality prediction and intelligent analysis of the prediction results are carried out through the hybrid model with the optimal parameters, abnormal data monitoring and over-threshold alarm are realized, and the alarm mechanism autonomously alarms and records the alarm information.
[0149] The water quality prediction results are displayed in real time in intelligent decision-making and system maintenance, historical data are intelligently analyzed to check whether sewage indicators meet national standards, and intelligent analysis and intelligent decision-making are provided; finally, according to demand and technological development, the system functions and performance are regularly upgraded, and the latest monitoring technologies and prediction methods are introduced to improve system performance.
[0150] The above embodiments are only for illustrating the technical concept and features of the present invention, and their purpose is to enable people familiar with the technology to understand the content of the present invention and implement it accordingly, and they cannot be used to limit the protection scope of the present invention. Any equivalent transformation or modification made according to the spirit of the present invention should be included in the protection scope of the present invention.
Claims
1. A water quality data prediction method based on a hybrid model, characterized in that: include: Step 1: Collect historical water quality data set D0. The number of D0 water quality types includes but is not limited to pH, temperature, turbidity, suspended solids, dissolved oxygen, ammonia nitrogen, total nitrogen, and total phosphorus; Step 2: Use the intelligent cleaning strategy DCS to clean the water quality data set D0 to remove outliers and fill missing values in the water quality data, perform Min-Max standardization on the processed water quality data to unify the data scale, and use the principal component analysis method to select time series characteristic water quality data with strong correlation to form the water quality characteristic data set D1; Step 3: The water quality characteristic dataset D1 is decomposed by the AMEMD composite method to obtain the water quality dataset D2 and divide it; Step 4: Construct a water quality intelligent hybrid prediction model. The prediction results of the water quality intelligent hybrid prediction model are mixed by weighted average of the prediction results of the GT-GAN-CNN-Bi-LSTM-Attention model and the XGBoost model. The steps of the GT-GAN-CNN-Bi-LSTM-Attention model are as follows: Step 4.1: Design the generative adversarial network model GT-GAN to expand the water quality dataset; Step 4.2: Use convolutional neural network (CNN) to process the expanded water quality data and extract local features in the time series; Step 4.3: The features extracted by CNN are further processed through the bidirectional long short-term memory network Bi-LSTM to capture the long-term dependencies of the time series; Step 4.4: Introduce the attention mechanism to enable the model to identify and focus on the key features in the time series; Step 5: Use the improved NPDAO algorithm to optimize the hyperparameter combination P0 of the water quality intelligent hybrid prediction model constructed in step 4, and train and predict the water quality data set divided in step 3. 2 , root mean square error RMSE and mean absolute error MAE to evaluate the model performance and obtain the optimal hyperparameter combination of the hybrid model; Step 6: Use the GT-GAN-CNN-Bi-LSTM-Attention and XGBoost intelligent hybrid prediction model with the optimal parameter combination to predict real-time water quality data.
2. The water quality data prediction method based on the hybrid model according to claim 1 is characterized in that: In step 2, the specific steps of the intelligent cleaning strategy DCS are as follows: Step 2.1: Use the box plot method to preliminarily identify outliers in the D0 data set, and then use the Z-score method to verify the outliers. Calculate the Z-score of each data point. If the absolute value of the Z-score is greater than the threshold determined by the 3σ rule, the data point is marked as the final outlier. The Z-score formula is as follows: For each data point d i , calculate its Z-score: Where: μ represents the mean value of the data set D0; σ represents the standard deviation; Step 2.2: Mark the outliers identified in step 2.1 as missing values; Step 2.3: Use k-NN to interpolate the missing values marked in step 2.2, and use the local similarity of the D0 data set to find neighbors similar to the missing data points through the Mahalanobis distance for interpolation.
3. The water quality data prediction method based on the hybrid model according to claim 1 is characterized in that: The specific steps of the AMEMD composite decomposition method in step 3 are: Step 3.1: Perform EMD decomposition on the water quality data D1 to obtain K IMF components; Step 3.2: Calculate the Pearson correlation coefficient for the K IMF components obtained in step 3.
1. The calculation steps are as follows: Step 3.2.1: For K IMF components, first, for any two IMF components i and IMF j ,1≤i<j≤K, first calculate its covariance Cov(IMF i ,IMF j ), whose formula is: Where: t represents the index in the water quality data sequence t=1,2,3,…,N; N represents the length of the water quality data; IMF i,t ,IMF j,t IMF i ,IMF j The value at t; IMF i ,IMF j The mean of Step 3.2.2: Calculate IMF i and IMF j The standard deviation of σIMF i ,σIMF j ; Step 3.2.3: Calculate IMF using Pearson correlation coefficient formula i and IMF j The correlation coefficient p i,j , the formula is: Step 3.2.4: Repeat the above steps to calculate the Pearson correlation coefficient of the IMF components. correlation coefficients; Step 3.2.5: Construct the correlation coefficient matrix R, which is represented as follows: Step 3.3: Aggregate the IMF components with strong correlation using the matrix R obtained in step 3.
2. The steps are as follows: Step 3.3.1: Create a Boolean array used of length K, with all initial values set to False, indicating that the K IMF components obtained by EMD decomposition are not aggregated; Step 3.3.2: Considering the selection of the maximum and optimal Pearson correlation coefficient, a greedy algorithm is used to optimize to ensure that the maximum Pearson correlation coefficient and its corresponding index (i, j) are found from the matrix R at each step; Step 3.3.3: Set used[i], used[j] to True, indicating that the two components have been aggregated; Step 3.3.4: IMF i ,IMF j Aggregation is performed and the aggregated results are stored in a new array D2. The rows and columns related to i and j in the matrix R are set to 0 to prevent these components from being selected again. Step 3.3.5: Repeat steps 3.3.2 to 3.3.4 until all IMF components are aggregated or there are no aggregateable components that meet the stopping condition, and obtain the AMEMD composite decomposition data set D2. The specific stopping condition is: Where: Υ is a very small threshold used to determine whether it is large enough for aggregation.
4. The water quality data prediction method based on the hybrid model according to claim 1 is characterized in that: In step 5, the hyperparameter P0 includes the number of XGBoost trees, estimators, and learning rate; the number of filters, filter size, and pool size in the convolutional layer of CNN; the number of LSTM units, dropout rate, and learning rate of each LSTM layer of Bi-LSTM; and the number of hidden units, and dropout rate in the Attention mechanism.
5. The water quality data prediction method based on the hybrid model according to claim 1 is characterized in that: In step 5, the specific steps of improving the NPDAO algorithm optimization are as follows: Step 5.1: Input the parameters of the NPDOA algorithm, and at the same time input the hyperparameter group P0 to be optimized of the hybrid model of the GT-GAN-CNN-Bi-LSTM-Attention model and the XGBoost model, and use the Tent chaotic map to initialize P0; Step 5.2: Use the attractor trend strategy to approach the excellent hybrid model parameter group P, couple the interference strategy to increase randomness to promote exploration, and the information projection strategy to control the information flow, so as to achieve the transition from extensive exploration to in-depth utilization; Step 5.3: Use the competitive update mechanism to update P; Step 5.4: Determine whether the maximum number of evaluations has been reached. If yes, proceed to the next step. If no, return to step 5.
2. Step 5.5: Select the optimal parameter set P through the memory strategy max , output P max and the corresponding target value.
6. The water quality data prediction method based on the hybrid model according to claim 5 is characterized in that: The improved NPDOA algorithm uses Tent chaotic map initialization as follows: Generating chaotic sequences based on Tent The process is as follows: Where: k is the number of parameter groups P0 of the hybrid model to be optimized; t is the number of evaluations, and u is a random number Combined with chaotic sequence The process of further generating the parameter group of the hybrid model is as follows: in: The worst and best parameter sets of the neural hybrid model, respectively.
7. The water quality data prediction method based on the hybrid model according to claim 5 is characterized in that: The competitive update mechanism is adopted in step 5.3, and the steps are as follows: Step 5.3.1: Considering that a single evaluation index cannot more comprehensively evaluate the current parameter combination P = {ξ1,ξ2,...,ξ O }, so MAE, MAPE, RMSE multiple evaluation indicators are selected to form a multi-objective optimization objective function f() to measure the parameter combination P = {ξ1,ξ2,...,ξ O }, where O is the number of groups; Step 5.3.2: According to f(), select the top 10% as an elite hybrid model parameter group individual set P elite ={ξ e1 ,ξ e2 ,...,ξ eQ }, Q is the number of elite individuals; Step 5.3.3: For each non-elite mixture model parameter group individual ξ h , h>Q, randomly select elite individuals ξ eq , As a competitor, if f(ξ h )<f(ξ eq ), then ξ h Replacement eq .
8. The water quality data prediction method based on hybrid model according to claim 5 is characterized in that: In step 5.5, the optimal parameter group is selected by the memory strategy, and the steps are as follows: Step 5.5.1: Construct a memory bank M to store the historical optimal hybrid model parameter set, P best is the current optimal hybrid model parameter set, P M,best is the optimal hybrid model parameter set in the memory bank; Step 5.5.2: Update the memory M, if f(P best )<f(P M,best ), then M=P best ; Step 5.5.3: Use M to search based on the following: P new =Search(P,M) (8) Where: P is the current hybrid model parameter group, Search is the search guidance operation, P new is the newly generated hybrid model parameter group.
9. A water quality data intelligent monitoring system based on a hybrid model, characterized in that: include: Data collection module, data and storage module, data transmission module and prediction and monitoring cloud platform; The data acquisition module collects historical water quality data sets and stores them in the data and storage module, and transmits them to the prediction and monitoring cloud platform through the data transmission module; The prediction and monitoring cloud platform is provided with a water quality data prediction method based on a hybrid model as described in any one of claims 1 to 8; The intelligent decision-making and system maintenance module is used to make decisions and perform maintenance based on the predicted water quality data obtained by the prediction and monitoring cloud platform.
10. The water quality data intelligent monitoring system based on hybrid model according to claim 9 is characterized in that: It also includes an intelligent decision-making and system maintenance module, which displays the prediction results in real time, intelligently analyzes historical data to check whether the water quality data meets national standards, and provides intelligent decision-making and optimization suggestions.
Citation Information
Cited By
Oxygen concentration regulation and control method based on hybrid intelligent model
CN120949841A
River management-oriented multi-parameter water quality intelligent prediction method and system
CN121393643A