A real-time water quality prediction method and system based on time series multi-scale feature fusion
By combining GRU and sparse attention mechanism, the multi-scale feature fusion is solved, and the problem of short-term fluctuations and long-term trends in traditional water quality prediction methods is difficult to take into account, achieving high-precision and efficient real-time prediction of water quality, which is suitable for water quality monitoring systems.
Patent Information
- Application Number
- CN202510584696.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-05-08
AI Technical Summary
Traditional water quality prediction methods are difficult to accurately identify short-term fluctuations and long-term trends in water quality parameters at the same time, resulting in insufficient prediction accuracy and real-time performance. The existing models have problems such as high computational complexity and insufficient global feature capture in long-term series data.
Short-term dynamic features are extracted using gated recurrent unit network (GRU), combined with sparse attention mechanism Informer model extracts long-term global dependency features, and multi-scale feature fusion, whale optimization algorithm is used to optimize hyperparameters, combined with NB-IoT technology to achieve efficient data transmission and solar power supply, and build a real-time water quality prediction system based on time-sequence multi-scale feature fusion.
It significantly improves the complex time series listing ability and prediction accuracy of water quality prediction, reduces the complexity of long-sequence full-connection calculation, improves operating efficiency and real-time response capabilities, and meets the high-time requirements of water quality monitoring.
Smart Images

Figure CN120105021B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of water quality detection technology, and in particular to a real-time water quality prediction method and system based on time series multi-scale feature fusion. Background Art
[0002] In recent years, water quality monitoring and prediction have become important tools for water environment protection. However, water quality parameters are often affected by multiple environmental factors and exhibit significant time-dependence. Traditional prediction methods often use single models, such as long-short-term memory networks (LSTMs) or recurrent neural networks (GRUs). While these methods can capture the spatiotemporal characteristics and short-term dynamics of water quality parameters to a certain extent, they are limited in addressing long-term dependencies and capturing multi-scale features. Especially in long-term series data, a single model struggles to accurately identify both short-term fluctuations and long-term trends, limiting the accuracy and real-time nature of predictions.
[0003] Chinese patent publication number "CN117744703A," titled "A Wastewater Quality Prediction Method, Electronic Device, and Medium," collects historical wastewater quality data (COD and -N concentrations), uses Butterworth low-pass filtering for denoising, and constructs a prediction model consisting of an autoencoder, ResGRU, and decoder. The autoencoder extracts real-time uncertainty features, the ResGRU learns long-term dependencies, and the decoder outputs predicted values and adjusts model parameters. Ultimately, wastewater quality prediction is performed using the trained model. The Butterworth low-pass filter has limited processing capabilities for non-stationary signals and multi-frequency noise, which can result in loss of feature information or incomplete noise removal. Long-term dependencies are primarily learned through the ResGRU, which can suffer from high computational complexity and insufficient global feature capture. Furthermore, a single network structure makes it difficult to simultaneously model both short-term dynamic features and long-term dependencies. Summary of the Invention
[0004] The present invention solves the above technical problems by providing a method for real-time water quality prediction based on multi-scale feature fusion, comprising the following steps:
[0005] S10, obtaining water quality data: the water quality data includes pH, total nitrogen, total phosphorus, ammonia nitrogen, water temperature, humidity, chemical oxygen demand, conductivity, and dissolved oxygen;
[0006] S20: Preprocessing of water quality data: including missing value processing using mean filling method, outlier processing using Z-score method, denoising using variational mode decomposition, extracting more multi-scale features and removing residual noise using wavelet transform, standardization, and dividing the processed data into training set, test set and validation set;
[0007] S30: Establishing a network model: including a gated recurrent unit network and a sparse attention mechanism. The gated recurrent unit network is used to extract short-term dynamic features and time dependencies in water quality data, and the informer extracts global information and long-distance dependency features.
[0008] S40: Establish the best parameter combination: Use the whale optimization algorithm to optimize the hyperparameters in the network model, such as the learning rate, the number of hidden units in the gated recurrent unit network, the batch size, the number of attention heads and the sparsity factor of the sparse attention mechanism, to find the best parameter combination;
[0009] S50: Feature fusion: The output features of the gated recurrent unit network and the sparse attention mechanism are fused through feature concatenation, and the fused features are input into the fully connected layer to generate future water quality parameter predictions.
[0010] Furthermore, the real-time water quality prediction method based on multi-scale feature fusion further includes the following steps:
[0011] S60: Upload data to the cloud platform in real time and automatically generate detailed data logs.
[0012] Furthermore, the gated recurrent unit network includes an update gate and a reset gate. Through the gating mechanism, the gated recurrent unit network selectively retains or forgets the past state, which is used to extract the short-term time-dependent characteristic period of the original water quality data and is more computationally efficient and more suitable for real-time applications.
[0013] Furthermore, the calculation formula of the sparse attention mechanism is:
[0014]
[0015] Among them, Q′ is the filtered key query subset for sparse calculation, K is all key vectors, V is all value vectors, d k is the scaling factor, which is the dimension of the key vector and is used to avoid the inner product value being too large; calculate each query vector Q in the query matrix Q i The L2 norm of is used to filter out the most important query vectors; the formula is as follows:
[0016] ;
[0017] Where: Q i is the query vector in the query matrix; ||Q i || is Q i The L2 norm of
[0018] Calculate the attention score for Q′ and K:
[0019]
[0020] Multiply the sparse attention result with the value matrix V to get the final output.
[0021] Furthermore, the whale optimization algorithm includes:
[0022] Set the algorithm parameters and randomly generate initial individuals;
[0023] Calculate individual fitness and record the best individual;
[0024] Generate all parameters, including scaling factor a, enclosing coefficient A, distance factor C, random probability p, current iteration number t, and maximum iteration number G;
[0025] Update the control parameter a, and update A and C according to the formula;
[0026] Determine its update strategy according to random probability p;
[0027] If |A| < 1, the whale individual shrinks to the current best solution;
[0028] If |A| >= 1, the whale individual learns from other random positions;
[0029] The entire process is repeated until the maximum number of iterations is reached, and the optimal hyperparameter combination is returned. The learning rate in the fitness function is 0.0001 to ensure stability during network training. The Adam optimizer is selected to update the model parameters based on the gradient information of the loss, so that the model is gradually optimized. The optimal hyperparameters found by WOA are used to train the model on the training set and validation set.
[0030] In order to solve the above technical problems, the present invention further proposes a real-time water quality prediction system based on time series multi-scale feature fusion, which is used to execute the above-mentioned real-time water quality prediction method based on multi-scale feature fusion, comprising:
[0031] The main control module uses STM32F103C8T6 as the main control chip, which is responsible for coordinating and processing the data and tasks of each functional module;
[0032] The data acquisition module collects water quality parameters through a variety of sensors and uses a digital-to-analog converter to transmit the signals to the main control module for processing;
[0033] Data preprocessing module, used to clean, reduce noise and standardize the collected data;
[0034] The data prediction module, which integrates short-term and long-term features based on a network model, is used to analyze pre-processed data and generate high-precision water quality prediction results;
[0035] Storage module, used to record collected and predicted data and generate data logs;
[0036] Display module, used to display various water quality indicators and equipment operating status in real time;
[0037] Positioning module, used to provide accurate location information when water quality is abnormal;
[0038] Monitoring module, used to detect system status in real time, collect operating data, and perform abnormality analysis;
[0039] Data transmission module, used to achieve efficient data transmission and synchronization with the cloud, supporting remote access and management;
[0040] The power management module is used to ensure stable power supply of the system in different environments.
[0041] Furthermore, the data preprocessing module uses Python's data processing library to perform missing value completion and noise reduction preprocessing operations.
[0042] Furthermore, the storage module writes the collected data and prediction results into the database of the cloud platform in real time for expanding the historical data set and subsequent analysis.
[0043] Furthermore, the data transmission module adopts NB-IoT technology to achieve efficient and reliable transmission of various types of data within the system.
[0044] Furthermore, the power management module supports solar power supply, has intelligent management functions, and can efficiently collect and store solar energy.
[0045] Compared with the prior art, the present invention has the following beneficial effects:
[0046] 1. This paper designs a parallel structure based on a gated recurrent unit (GRU) and an informer with a sparse attention mechanism. The GRU captures short-term dynamic features, and the informer models long-term global dependencies. Through multi-scale feature fusion, it effectively solves the problem of balancing short-term changes and long-term trends in water quality data, significantly improving the representation ability and prediction accuracy of complex time series.
[0047] 2. This paper proposes a parallel architecture combining the GRU and the Informer with a sparse attention mechanism. This architecture utilizes the key time step selection strategy of the probabilistic sparse attention mechanism to significantly reduce the computational complexity of fully connected long sequences. The GRU reduces computational overhead through an efficient gating mechanism, while the Informer utilizes sparsification operations to model global long-term dependencies, significantly optimizing computational resource allocation. While ensuring high accuracy, this architecture improves operational efficiency and the ability to rapidly respond to real-time data, meeting the critical timeliness requirements of water quality monitoring.
[0048] 3. The present invention realizes the precise extraction and fusion of short-term dynamic and long-term global features through the main control chip, high-precision data acquisition module and the prediction model of the GRU-Informer parallel architecture. Combined with the positioning, monitoring, cloud transmission, real-time display module and power management module, the system has the capabilities of efficient collection, intelligent prediction, anomaly positioning, data visualization and stable power supply guarantee, which significantly improves the real-time, accuracy and adaptability of water quality monitoring. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without paying any creative work.
[0050] Figure 1 This is a flowchart of the steps of the real-time water quality prediction method based on multi-scale feature fusion according to the present invention;
[0051] Figure 2 This is a flow chart of the steps of data preprocessing according to the present invention;
[0052] Figure 3 This is a flow chart of the whale optimization algorithm described in the present invention;
[0053] Figure 4 This is a schematic diagram of the structure of the network model of the present invention;
[0054] Figure 5 This is a schematic flow chart of the water quality prediction model of the present invention;
[0055] Figure 6 This is a structural diagram of the real-time water quality prediction system based on multi-scale feature fusion described in the present invention.
[0056] Description of Figure Numbers:
[0057] 11. Main control module; 12. Data acquisition module; 13. Data preprocessing module; 14. Data prediction module; 15. Storage module; 16. Display module; 17. Positioning module; 18. Monitoring module; 19. Data transmission module; 20. Power management module. DETAILED DESCRIPTION
[0058] The present invention proposes a real-time water quality prediction method and system based on time series multi-scale feature fusion, aiming to.
[0059] The following is an explanation of the real-time water quality prediction method and system based on time series multi-scale feature fusion proposed by the present invention in a specific embodiment:
[0060] Example 1: A real-time water quality prediction method based on multi-scale feature fusion, such as Figure 1 As shown, the following steps are included:
[0061] S10, obtaining water quality data: the water quality data includes pH, total nitrogen, total phosphorus, ammonia nitrogen, water temperature, humidity, chemical oxygen demand, conductivity, and dissolved oxygen;
[0062] Specifically, water quality monitoring equipment is first deployed. This equipment integrates a variety of pluggable sensor probes that can detect key indicators in the water, including alkalinity, total nitrogen, total phosphorus, ammonia nitrogen, water temperature, humidity, chemical oxygen demand, conductivity, and dissolved oxygen. The equipment has excellent waterproof and dustproof properties, ensuring its stable operation in complex and harsh environments. The sensor probe adopts a pluggable design, which is easy to maintain and replace. At the same time, according to the requirements of different indicators, the probe can be placed at different depths of the water body to improve monitoring accuracy. The equipment can be fixed at locations such as the edge of the water or bridges to support long-term continuous monitoring. The data will be stored in the system's storage module in the offline state and automatically uploaded to the cloud after connecting to the network to ensure data integrity and traceability.
[0063] S20: Preprocessing of water quality data: including missing value processing using mean filling method, outlier processing using Z-score method, denoising using variational mode decomposition, extracting more multi-scale features and removing residual noise using wavelet transform, standardization, and dividing the processed data into training set, test set and validation set;
[0064] Specifically, if Figure 2 As shown in the figure, the collected water quality data are first processed for missing values and outliers and denoised. The missing values are filled with the mean value, that is, for columns with missing values in the data set, the mean value of the column is used to replace the missing values.
[0065] Outliers are identified using the Z-score method, which is a statistically standardized scoring method suitable for numerical data. It determines whether a value is an outlier by calculating the degree of deviation of each value from the mean.
[0066] The variational mode decomposition (VMD) module is further added to effectively reduce noise interference;
[0067] Then add wavelet transform to extract more multi-scale features and remove residual noise;
[0068] Finally, the data is standardized and divided into training set, test set and validation set.
[0069] Variational mode decomposition (VMD) is a non-recursive signal processing method that decomposes water quality time series data into a series of intrinsic mode functions with limited bandwidth. The specific steps are as follows: First, the input water quality data (such as dissolved oxygen, pH value, etc.) are initialized with parameters, including the mode number K, penalty factor and convergence threshold , ensuring that the number of decomposed modes matches the complexity of the data; then enter the iterative optimization stage, gradually update the spectral components of each mode, calculate the center frequency of the mode, and adjust the error term based on the Lagrange multiplier to ensure the accuracy and stability of the decomposition results; when the update amplitude of all modes is less than the threshold, stop the iteration and obtain IMFs corresponding to different frequency bands. These IMFs separate the multi-frequency components of water quality data, where the high-frequency part reflects short-term fluctuations and the low-frequency part represents the long-term trend, providing high-quality feature input for subsequent wavelet transform optimization and model modeling. This method uses iterative search to find the optimal solution of the variational mode, so that each mode can adaptively update its optimal center frequency and bandwidth. This feature enables variational mode decomposition to effectively deal with non-stationarity and noise in water quality data, thereby extracting the dynamic characteristics of different frequency components.
[0070] The wavelet transform achieves a multi-level representation of the signal by decomposing it into a series of components with different frequencies, and then into wavelet coefficients of different frequencies and temporal resolutions. First, a mother wavelet function (such as the db4 wavelet) suitable for the characteristics of water quality data is selected, and the number of decomposition layers, L, and the type of threshold function are set. Each IMF is subjected to layer-by-layer wavelet decomposition, decomposing the signal into detail components and approximate components. The detail components capture short-term fluctuations, while the approximate components preserve long-term trends. Then, soft or hard thresholding is applied to the high-frequency detail components to remove noise interference and retain valid features. Finally, the denoised detail and approximate components are reconstructed to obtain an optimized signal. This process, by combining multi-frequency decomposition using VMD with multi-scale analysis using wavelets, not only enhances the short-term dynamics and long-term trend characteristics of the water quality time series but also significantly improves the signal purity and the accuracy of the model input features.
[0071] S30: Establishing a network model: including a gated recurrent unit network and a sparse attention mechanism. The gated recurrent unit network is used to extract short-term dynamic features and time dependencies in water quality data, and the informer extracts global information and long-distance dependency features.
[0072] S40: Establish the best parameter combination: Use the whale optimization algorithm to optimize the hyperparameters in the network model, such as the learning rate, the number of hidden units in the gated recurrent unit network, the batch size, the number of attention heads and the sparsity factor of the sparse attention mechanism, to find the best parameter combination;
[0073] S50: Feature fusion: The output features of the gated recurrent unit network and the sparse attention mechanism are fused through feature concatenation, and the fused features are input into the fully connected layer to generate future water quality parameter predictions.
[0074] S60: Upload data to the cloud platform in real time and automatically generate detailed data logs.
[0075] Furthermore, the gated recurrent unit network includes an update gate and a reset gate. Through the gating mechanism, the gated recurrent unit network selectively retains or forgets the past state, which is used to extract the short-term time-dependent characteristic period of the original water quality data and is more computationally efficient and more suitable for real-time applications.
[0076] Furthermore, the calculation formula of the sparse attention mechanism is:
[0077]
[0078] Among them, Q′ is the filtered key query subset for sparse calculation, K is all key vectors, V is all value vectors, d k is the scaling factor, which is the dimension of the key vector and is used to avoid the inner product value being too large; calculate each query vector Q in the query matrix Q i The L2 norm of is used to filter out the most important query vectors; the formula is as follows:
[0079] ;
[0080] Where: Q i is the query vector in the query matrix; ||Q i || is Q i The L2 norm of
[0081] Calculate the attention score for Q′ and K:
[0082]
[0083] Multiply the sparse attention result with the value matrix V to get the final output.
[0084] The short-term dynamic features extracted by the gated recurrent unit network and the long-term dependency features captured by the informer are combined to achieve multi-scale information fusion through feature splicing. The fused high-dimensional features are embedded in the fully connected layer and further characterized by the deep learning network to accurately generate prediction results for future water quality parameters.
[0085] The structure of the long time series prediction model based on the attention mechanism that introduces the sparse attention mechanism is as follows Figure 5 First, input the time series data It is fed into the encoder module, where it is extracted through a multi-layer, multi-head sparse probabilistic self-attention mechanism (ProbSparse Attention), and combined with a distillation mechanism to gradually compress redundant information to form a refined feature map (feature concatenation map). Subsequently, the features output by the encoder are processed by the decoder. The decoder receives the prior input , which includes data at known time steps and the initial value of the fill In the decoder, a multi-head attention mechanism and a masked sparse attention mechanism are used to further extract key global and local features from the time series, effectively modeling long sequences. Ultimately, the decoder's output features are fed into a fully connected layer to generate predictions, accurately outputting parameter values for future time steps. This architecture reduces computational complexity through sparse attention while maintaining effective modeling of long-term dependencies.
[0086] Furthermore, if Figure 3 As shown, the whale optimization algorithm includes:
[0087] Set the algorithm parameters and randomly generate initial individuals;
[0088] Calculate individual fitness and record the best individual;
[0089] Generate all parameters, including scaling factor a, enclosing coefficient A, distance factor C, random probability p, current iteration number t, and maximum iteration number G;
[0090] Update the control parameter a, and update A and C according to the formula;
[0091] Determine its update strategy according to random probability p;
[0092] If |A| < 1, the whale individual shrinks to the current best solution;
[0093] If |A| >= 1, the whale individual learns from other random positions;
[0094] The entire process is repeated until the maximum number of iterations is reached, and the optimal hyperparameter combination is returned. The learning rate in the fitness function is 0.0001 to ensure stability during network training. The Adam optimizer is selected to update the model parameters based on the gradient information of the loss, so that the model is gradually optimized. The optimal hyperparameters found by WOA are used to train the model on the training set and validation set.
[0095] Example 2: A real-time water quality prediction method based on time series multi-scale feature fusion, comprising the following steps:
[0096] Step 1: Obtain water quality data: Deploy water quality monitoring equipment. This equipment integrates a variety of pluggable sensor probes that can detect key water indicators, including alkalinity, total nitrogen, total phosphorus, ammonia nitrogen, water temperature, humidity, chemical oxygen demand, conductivity, and dissolved oxygen. The equipment is waterproof and dustproof, ensuring stable operation in complex and harsh environments. The sensor probes are pluggable for easy maintenance and replacement. They can be placed at different depths in the water to improve monitoring accuracy based on the needs of different indicators. The equipment can be fixed at locations such as the edge of the water or on bridges, supporting long-term continuous monitoring. Data is stored in the system's storage module when offline and automatically uploaded to the cloud when connected to the network, ensuring data integrity and traceability.
[0097] Step 2: The cloud platform displays and stores data: The cloud platform displays alkalinity, total nitrogen, total phosphorus, ammonia nitrogen, water temperature, humidity, chemical oxygen demand, conductivity, and dissolved oxygen data in real time and stores them in the database.
[0098] Step 3: Data Preprocessing: The collected water quality data was first processed for missing values and outliers, and denoised. Missing values were imputed using the mean-filling method, which replaces missing values in columns with missing values in the dataset with the mean of that column. Outliers were identified using the Z-score method, a statistically standardized score method suitable for numerical data. This method calculates the degree of deviation of each value from the mean to determine whether it is an outlier. A variational mode decomposition (VMD) module was further added to effectively reduce noise interference. A wavelet transform was then used to extract more multi-scale features and remove residual noise. Finally, the data was normalized.
[0099] Step 4: Optimize network model parameters. Use the Whale Optimization Algorithm (WOA) to optimize hyperparameters in the network model, including the learning rate, number of hidden units in the gated recurrent unit network, batch size, number of attention heads in the sparse attention mechanism, and sparsity factor. Finding the optimal parameter combination is crucial. The Whale Optimization Algorithm simulates the bubble net hunting behavior of humpback whales, which includes three behaviors: swarming, bubble net feeding, and random prey search. These behaviors enable whale groups to efficiently explore the search space and find the optimal solution. During hyperparameter optimization, WOA continuously updates the whales' positions to find the optimal parameter combination to minimize the objective function.
[0100] Step 5: Train the water quality prediction network model. In order to improve the prediction accuracy, the gated recurrent unit network (GRU) and the long time series prediction model (Informer) that introduces the sparse attention mechanism are used to extract the multi-scale features of the water quality data in parallel. GRU selectively retains or forgets the pre-processed time series data by updating the gate and resetting the gate, efficiently extracting short-term dynamic features and time dependencies, and is suitable for the rapid processing of real-time data. At the same time, the Informer model models the global features in the sequence through the self-attention mechanism, and adopts the sparse self-attention mechanism to significantly reduce the computational complexity by screening the key time steps, while maintaining the ability to capture global features. Compared with the traditional fully connected self-attention mechanism, Complexity, sparse attention uses the probabilistic sparse strategy (ProbSparse Attention) to calculate the query vector Norm filters out the most representative subset of key queries, focuses on the interaction between key time steps and all key-value vectors, and only calculates the necessary attention scores, reducing the complexity to This mechanism retains the global dependency characteristics and enhances numerical stability through scaling factors, effectively processing long time series data, taking into account both efficiency and modeling accuracy, and is very suitable for long time series prediction tasks. The calculation formula of the sparse attention mechanism can be expressed as:
[0101]
[0102] Q′ is the filtered subset of key queries for sparse computation, K is all key vectors, V is all value vectors, and dk is a scaling factor, usually the dimension of the key vector, to avoid excessively large inner product values.
[0103] Calculate the L2 norm of each query vector Qi in the query matrix Q to filter out the most important query vectors. The formula is as follows:
[0104] ;
[0105] Where: Qi is the query vector in the query matrix. ||Qi|| is the L2 norm of Qi.
[0106] Calculate the attention score for Q′ and K:
[0107]
[0108] Multiply the sparse attention result with the value matrix V to get the final output.
[0109] The short-term dynamic features extracted by the gated recurrent unit network and the long-term dependency features captured by the informer are combined to achieve multi-scale information fusion through feature splicing. The fused high-dimensional features are embedded in the fully connected layer and further characterized by the deep learning network to accurately generate prediction results for future water quality parameters.
[0110] The sparse attention mechanism is introduced into the long time series prediction model structure based on the attention mechanism; first, the time series data is input It is fed into the encoder module, where it is extracted through a multi-layer, multi-head sparse probabilistic self-attention mechanism (ProbSparse Attention), and combined with a distillation mechanism to gradually compress redundant information to form a refined feature map (feature concatenation map). Subsequently, the features output by the encoder are processed by the decoder. The decoder receives the prior input , which includes data at known time steps and the initial value of the fill In the decoder, a multi-head attention mechanism and a masked sparse attention mechanism are used to further extract key global and local features from the time series, effectively modeling long sequences. Ultimately, the decoder's output features are fed into a fully connected layer to generate predictions, accurately outputting parameter values for future time steps. This architecture reduces computational complexity through sparse attention while maintaining effective modeling of long-term dependencies.
[0111] Step 6: Complete the water quality prediction. The test dataset includes not only real-time data collected on-site but also historical data stored in the cloud. Expanding the size and diversity of the dataset improves the model's generalization and characterization of water quality characteristics, resulting in more accurate and reliable predictions.
[0112] Specifically, as shown in Table 1, the deployed detection equipment collects water quality information in the Yitong River Basin and evaluates it through performance error.
[0113] Table 1. One-day water quality parameters of the Yitong River Basin. Unit: mg / L; Conductivity (ms / m)
[0114] The mean absolute error (MAE), mean square error (MSE), root mean square error (RMSE), mean absolute percentage error (MAPE), and coefficient of determination (R²) are used to evaluate the performance of the model. The calculation formulas are as follows:
[0115] ;
[0116] ;
[0117] ;
[0118] ;
[0119] ;
[0120] Where N is the data length (test set sample size), is the predicted value, is the true value, and Table 2 shows the prediction performance of the calculated DO index.
[0121] Table 2 DO index prediction performance:
[0122]
[0123] Example 3: A real-time water quality prediction system based on time series multi-scale feature fusion, used to execute the real-time water quality prediction method based on multi-scale feature fusion as described in Example 1, comprising:
[0124] The main control module 11 uses STM32F103C8T6 as the main control chip, which is responsible for coordinating and processing the data and tasks of each functional module;
[0125] The data acquisition module 12 collects water quality parameters through various sensors and transmits the signals to the main control module for processing using a digital-to-analog converter;
[0126] The data preprocessing module 13 is used to clean, reduce noise and standardize the collected data;
[0127] The data prediction module 14 is used to analyze the pre-processed data and generate high-precision water quality prediction results by fusing short-term and long-term features based on the network model;
[0128] The storage module 15 is used to record the collected and predicted data and generate a data log;
[0129] Display module 16, used to display various water quality indicators and equipment operating status in real time;
[0130] Positioning module 17, used to provide accurate location information when water quality is abnormal;
[0131] Monitoring module 18, used to detect system status in real time, collect operating data, and perform abnormality analysis;
[0132] Data transmission module 19, used to achieve efficient data transmission and synchronization with the cloud, and support remote access and management;
[0133] The power management module 20 is used to ensure stable power supply of the system in different environments.
[0134] Working principle: Figure 6 As shown, the main control module 11 serves as the core of the system, coordinating various tasks. Through the data acquisition module 12, the system acquires water quality data in real time and rapidly transmits this data to the storage module 15 via the data transmission module 13, ensuring the security and reliability of the information. Once the data is stored, it is immediately incorporated into the trained model on the cloud platform. During this process, the data preprocessing module 13 performs real-time denoising on the data, laying a solid foundation for subsequent analysis. After preprocessing, the data is rapidly fed into the prediction module, which uses advanced algorithms to perform real-time water quality predictions and generate accurate predicted values. These predictions are immediately stored in the storage module 15 for subsequent query and analysis. Simultaneously, the monitoring module 18 ensures that users can view system status and data dynamics in real time, providing timely decision-making support. The positioning module provides location information in emergency situations, ensuring that relevant personnel can quickly respond and handle emergencies. To ensure stable system operation, the power management module 20 is responsible for maintaining the device's power supply. As for the user interface, the display module 16, equipped with an OLED screen, displays water quality data in real time, allowing users to intuitively access key information.
[0135] Furthermore, the data preprocessing module 13 uses Python's data processing library to perform missing value completion and noise reduction preprocessing operations.
[0136] Furthermore, the storage module writes the data collected by 15 and the prediction results into the database of the cloud platform in real time for expanding the historical data set and subsequent analysis.
[0137] Furthermore, the data transmission module 19 uses NB-IoT technology to achieve efficient and reliable transmission of various types of data within the system.
[0138] Furthermore, the power management module 20 supports solar power supply, has intelligent management functions, and can efficiently collect and store solar energy.
[0139] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. A real-time water quality prediction method based on multi-scale feature fusion, characterized in that: The following steps are involved: S10: Obtaining water quality data: The water quality data includes pH, total nitrogen, total phosphorus, ammonia nitrogen, water temperature, humidity, chemical oxygen demand, conductivity, and dissolved oxygen; S20: Preprocessing of water quality data: including missing value processing using mean filling method, outlier processing using Z-score method, denoising using variational mode decomposition, extracting multi-scale features of each modal component using wavelet transform and removing residual noise, standardization, and dividing the processed data into training set, test set and validation set; S30: Establishing a network model: including a gated recurrent unit network and an informer, wherein the gated recurrent unit network is used to extract short-term dynamic features and time dependencies in water quality data, and the informer extracts global information and long-distance dependency features; S40: Establish the best parameter combination: Use the whale optimization algorithm to optimize the learning rate, number of hidden units in the gated recurrent unit network, batch size, number of attention heads of the informer, and sparsity factor hyperparameters in the network model to find the best parameter combination; S50: Feature fusion: The output features of the gated recurrent unit network and the informer are fused through feature concatenation, and the fused features are input into the fully connected layer to generate future water quality parameter predictions; The gated recurrent unit network includes an update gate and a reset gate. Through the gating mechanism, the gated recurrent unit network selectively retains or forgets the past state to extract the short-term time-dependent features of the original water quality data; The calculation formula of the Informer is: Among them, Q′ is the filtered key query subset for sparse calculation, K is all key vectors, V is all value vectors, d k is the scaling factor, which is the dimension of the key vector and is used to avoid the inner product value being too large; calculate each query vector Q in the query matrix Q i The L2 norm of is used to filter out the most important query vectors; the formula is as follows: ; Where: Q i is the i-th query vector in the query matrix; ||Q i || is Q i The L2 norm of Calculate the attention score for Q′ and K: Multiply the sparse attention result with the value matrix V to get the final output.
2. The real-time water quality prediction method based on multi-scale feature fusion according to claim 1 is characterized in that: The following steps are also included: S60: Upload data to the cloud platform in real time and automatically generate detailed data logs.
3. The real-time water quality prediction method based on multi-scale feature fusion according to claim 1 is characterized in that: The whale optimization algorithm includes: Set the algorithm parameters and randomly generate initial individuals; Calculate individual fitness and record the best individual; Generate all parameters, including scaling factor a, enclosing coefficient A, distance factor C, random probability p, current iteration number t, and maximum iteration number G; Update the control parameter a, and update A and C according to the formula; Determine its update strategy according to random probability p; If |A| < 1, the whale individual shrinks to the current best solution; If |A| >= 1, the whale individual learns from other random positions; The entire process is repeated until the maximum number of iterations is reached, and the optimal hyperparameter combination is returned. The learning rate in the fitness function is 0.0001 to ensure stability during network training. The Adam optimizer is selected to update the model parameters based on the gradient information of the loss, so that the model is gradually optimized. The optimal hyperparameters found by WOA are used to train the model on the training set and validation set.
4. A real-time water quality prediction system based on time series multi-scale feature fusion, used to execute the real-time water quality prediction method based on multi-scale feature fusion according to any one of claims 1 to 3, characterized in that: include: The main control module uses STM32F103C8T6 as the main control chip, which is responsible for coordinating and processing the data and tasks of each functional module; The data acquisition module collects water quality parameters through a variety of sensors and uses a digital-to-analog converter to transmit the signals to the main control module for processing; Data preprocessing module, used to clean, reduce noise and standardize the collected data; The data prediction module, which integrates short-term and long-term features based on a network model, is used to analyze pre-processed data and generate high-precision water quality prediction results; Storage module, used to record collected and predicted data and generate data logs; Display module, used to display various water quality indicators and equipment operating status in real time; Positioning module, used to provide accurate location information when water quality is abnormal; Monitoring module, used to detect system status in real time, collect operating data, and perform abnormality analysis; Data transmission module, used to achieve efficient data transmission and synchronization with the cloud, supporting remote access and management; The power management module is used to ensure stable power supply of the system in different environments.
5. The real-time water quality prediction system based on time series multi-scale feature fusion according to claim 4 is characterized in that: The data preprocessing module uses Python's data processing library to perform missing value completion and noise reduction preprocessing operations.
6. The real-time water quality prediction system based on time series multi-scale feature fusion according to claim 4 is characterized in that: The storage module writes the collected data and prediction results into the database of the cloud platform in real time for expanding the historical data set and subsequent analysis.
7. The real-time water quality prediction system based on time series multi-scale feature fusion according to claim 4 is characterized in that: The data transmission module uses NB-IoT technology to achieve efficient and reliable transmission of various types of data within the system.
8. The real-time water quality prediction system based on time series multi-scale feature fusion according to claim 4 is characterized in that: The power management module supports solar power supply, has intelligent management functions, and can efficiently collect and store solar energy.
Citation Information
Patent Citations
Sewage quality prediction method, electronic equipment and medium
CN117744703A
Sea cucumber culture water quality prediction method for optimizing GRU neural network based on whale algorithm
CN115859057A
TCN-LSTM-based underground space passenger flow prediction method and system
CN118172948A
Seawater parameter calibration system
CN119920364A