Water quality monitoring and predicting method and system based on multi-factor correlation analysis
By using multi-factor association analysis and conditional generative adversarial networks, the problem of underutilization of indicator correlation information in water quality monitoring was solved, achieving accuracy and timeliness in water quality prediction and ensuring the safety of the water environment.
Patent Information
- Application Number
- CN202511073753.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-01
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-08-01
AI Technical Summary
Existing water quality monitoring and prediction technologies have failed to fully explore the deep-seated, cross-timescale correlations between water quality indicators, resulting in insufficient ability of models to capture the evolution patterns of water quality. Furthermore, they have failed to distinguish the characteristics of indicators for joint prediction, leading to inaccurate prediction results.
Through multi-factor correlation analysis, water quality evaluation indicators were determined and data were collected. After preprocessing, correlation analysis was used to classify the indicator types, a data support set was constructed, and different types of indicator prediction models were used for classification modeling, including prediction of single-cause, multi-cause, and isolated indicators. Conditional generative adversarial networks were used for joint prediction.
It improves the accuracy and efficiency of water quality monitoring and forecasting, enabling timely early warning and ensuring water environment safety.
Smart Images

Figure CN120930934A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of water quality monitoring technology, and in particular to a water quality monitoring and prediction method and system based on multi-factor correlation analysis. Background Technology
[0002] Water quality is directly related to ecological security, drinking water health, and sustainable socio-economic development. With the intensification of industrial and agricultural activities and the acceleration of urbanization, water pollution problems in river basins are becoming increasingly complex and volatile, making water quality monitoring and early warning a core aspect of water resource management and protection. Traditional methods relying on manual sampling and laboratory analysis are time-consuming and costly, failing to meet the needs of large-scale, high-frequency, and real-time dynamic monitoring. Therefore, utilizing sensor networks, IoT technology, and big data analytics to construct intelligent water quality monitoring and prediction systems for early identification and warning of water pollution risks has become an urgent need and an important research direction in the field of water environment management. Accurately predicting future water quality trends is of great practical significance for scientifically formulating pollution prevention and control measures and ensuring water safety.
[0003] However, existing water quality monitoring and prediction technologies still face many challenges and limitations. First, existing methods often ignore or simplify the complex nonlinear and time-delayed relationships between water quality indicators. Water bodies are organic wholes; changes in indicators such as pH, dissolved oxygen, ammonia nitrogen, and total phosphorus are not isolated but rather mutually influential and causal. Existing technologies typically fail to consider this when making water quality predictions, making it difficult to fully explore and utilize the deep-seated, cross-timescale correlations between indicators, resulting in insufficient ability of models to capture water quality evolution patterns. Second, existing water quality monitoring and prediction models, such as those based on ARIMA, SVM, or general LSTM, mostly adopt a "one-size-fits-all" approach to treat all indicators, failing to distinguish indicator characteristics and making it difficult to integrate the effective information provided by interrelated indicators for joint prediction, which may lead to inaccurate water quality prediction results. Therefore, it is necessary to explore new water quality monitoring and prediction schemes. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this invention provides a water quality monitoring and prediction method and system based on multi-factor correlation analysis.
[0005] To achieve the above objectives, in a first aspect, this invention provides a water quality monitoring and prediction method based on multi-factor correlation analysis. The method includes the following steps: determining water quality evaluation indicators and collecting water quality evaluation indicator data for water bodies in a monitored watershed, thereby evaluating the water quality level and obtaining a water quality monitoring dataset for the water bodies in the monitored watershed; based on the water quality monitoring dataset, determining the correlation between each of the water quality evaluation indicators through correlation analysis, and then constructing data support sets for each of the water quality evaluation indicators; based on the data support sets of the water quality evaluation indicators, using an indicator prediction model to obtain the predicted values of the water quality evaluation indicators; predicting the future water quality level of the monitored water body based on the predicted values of each of the water quality evaluation indicators, and triggering an early warning when the warning conditions are met. This invention improves the accuracy of water quality monitoring and prediction through multi-factor correlation analysis and classification modeling.
[0006] Optionally, the steps of determining water quality evaluation indicators, collecting water quality evaluation indicator data of water bodies in the monitored watershed, evaluating water quality levels, and obtaining water quality monitoring datasets of water bodies in the monitored watershed include the following: Determine water quality evaluation indicators, including pH value, dissolved oxygen concentration, chemical oxygen demand, five-day biochemical oxygen demand, ammonia nitrogen, total nitrogen, total phosphorus, permanganate index, lead concentration, mercury concentration, and fluoride; Collect water quality evaluation index data of the water body and preprocess the water quality evaluation index data to obtain time series data of the water quality evaluation index data; After data preprocessing, the water quality level is evaluated using the current water quality evaluation index data, and a water quality monitoring dataset is constructed using all the time series data.
[0007] Furthermore, this method improves data quality and accuracy of water quality monitoring by collecting and preprocessing water quality evaluation index data, and provides a reliable data foundation for water quality prediction by constructing a water quality monitoring dataset.
[0008] Optionally, the step of determining the correlation between various water quality evaluation indicators based on the water quality monitoring dataset through correlation analysis, and then constructing data support sets for each water quality evaluation indicator, includes the following steps: Based on the water quality monitoring dataset, the comprehensive correlation between each water quality evaluation index is calculated, and then the water quality evaluation index is divided into single-factor indexes, multi-factor indexes, and isolated indexes. Based on the water quality monitoring dataset, corresponding data support sets are constructed for each of the single-factor indicators, multi-factor indicators, and isolated indicators.
[0009] Furthermore, this method uses correlation analysis to help to understand the interrelationships between various indicators, thereby classifying water quality evaluation indicators and constructing data support sets for each. This facilitates the application of different analysis methods for different types of indicators and provides a data foundation for indicator prediction, thus improving the accuracy of water quality prediction.
[0010] Optionally, the step of calculating the comprehensive correlation between the various water quality evaluation indicators based on the water quality monitoring dataset, and then classifying the water quality evaluation indicators into single-factor indicators, multi-factor indicators, and isolated indicators, includes the following steps: The mean of the maximum information coefficient and time-lag correlation between each of the water quality evaluation indicators is calculated and used as the comprehensive correlation between each of the water quality evaluation indicators. For any water quality evaluation index A, if there are other water quality evaluation index B whose overall correlation with it is greater than the correlation threshold, then B is an associated index of A. For any water quality evaluation indicator A, if it has one of the aforementioned related indicators, then A is the single-cause indicator; if it has multiple of the aforementioned related indicators, then A is the multi-cause indicator; if it does not have any of the aforementioned related indicators, then A is the isolated indicator.
[0011] Furthermore, this method calculates the comprehensive correlation between various water quality evaluation indicators, taking into account both linear and nonlinear correlations, as well as time lag relationships, making the correlation analysis more comprehensive and accurate, and thus enabling rapid and accurate classification of the indicators.
[0012] Optionally, constructing corresponding data support sets for each of the single-factor indicators, multi-factor indicators, and isolated indicators based on the water quality monitoring dataset includes the following steps: For any of the single-factor indicators, the time series data of the single-factor indicator and the time series data of its related indicators are extracted from the water quality monitoring dataset to construct the data support set of the single-factor indicator. For any of the multi-factor indicators, select multiple indicators from its associated indicators as its main associated indicators according to the magnitude of the comprehensive correlation. The time series data of the multifactor indicators and their main related indicators are extracted from the water quality monitoring dataset, and then a data support set for the multifactor indicators is constructed. For any of the isolated indicators, time series data are extracted from the water quality monitoring dataset to construct its data support set.
[0013] Furthermore, this method classifies different water quality assessment indicators and constructs separate data support sets for each, facilitating the application of different prediction methods for different types of indicators and providing a data foundation for indicator prediction, thereby improving the accuracy of water quality prediction. Specifically, for multi-factor indicators, joint prediction is achieved by selecting their main related indicators, which can improve prediction efficiency to a certain extent.
[0014] Optionally, the indicator prediction model includes a first type of indicator prediction model, a second type of indicator prediction model, and a third type of indicator prediction model; The step of obtaining the predicted values of the water quality evaluation indicators using an indicator prediction model based on the data support set of the water quality evaluation indicators includes the following steps: For the single-cause index, a first-class index prediction model is constructed using its data support set and conditional generative adversarial network, thereby realizing the prediction of the single-cause index; For the multifactor index, multivariate empirical mode decomposition is performed on each time series data in its data support set, and the mode decomposition set of the multifactor index is further obtained. The modality decomposition set of the multi-factor indicators is used to complete the training and validation of the conditional hierarchical generative adversarial network, thereby obtaining the second type of indicator prediction model, which is used to predict the multi-factor indicators. For the isolated indicator, a third type of indicator prediction model is constructed using its data support set and generative adversarial network, thereby realizing the prediction of the isolated indicator.
[0015] Furthermore, this method classifies and models different types of water quality evaluation indicators, which can give full play to the advantages of different models, improve the accuracy of water quality evaluation indicator prediction, and thus improve the accuracy of water quality prediction.
[0016] Optionally, for the multifactor index, performing multivariate empirical mode decomposition on each time series data in its data support set, and further obtaining the mode decomposition set of the multifactor index, includes the following steps: For the multifactor indicators, multivariate empirical mode decomposition is used to decompose the time series data of each water quality evaluation indicator in the same period of its data support set to obtain the IMF components and residual terms of the multifactor indicators and their main related indicators. The IMF components of the multifactor index and its main related indexes are divided into high-frequency components and low-frequency components by using sample entropy. The high-frequency components, low-frequency components, and residual terms of each major correlation indicator of the multi-factor index are fused into a high-frequency condition vector, a low-frequency condition vector, and a residual condition vector by using a weighted summation method. The mode decomposition set is constructed using the time series data, high-frequency components, low-frequency components, and residual terms of the multifactor index, as well as the high-frequency condition vector, the low-frequency condition vector, and the residual condition vector.
[0017] Furthermore, this method uses multivariate empirical mode decomposition (IMF) to decompose the time series data of various water quality evaluation indicators within the same time period in the multi-factor index data support set. This effectively extracts different frequency components of the data, providing richer information for subsequent analysis and prediction. Dividing the IMF components into high-frequency and low-frequency components using sample entropy helps distinguish different features in the data, providing a basis for subsequent conditional vector construction. Using a weighted summation method, the high-frequency, low-frequency components, and residual terms of each major correlated indicator of the multi-factor index are fused into high-frequency conditional vectors, low-frequency conditional vectors, and residual conditional vectors, respectively. This comprehensively considers the influence of major correlated indicators and improves the performance of the prediction model. The construction of the mode decomposition set provides comprehensive and rich data support for the training and validation of conditional hierarchical generative adversarial networks.
[0018] Optionally, the conditional hierarchical generative adversarial network includes a condition combination module, a generator module, and a four-level discriminant module. The four-level discriminant module includes three condition discriminators and a total discriminator. The condition discriminators include a first condition discriminator, a second condition discriminator, and a third condition discriminator. The first condition discriminator, the second condition discriminator, and the third condition discriminator each correspond to a condition. The conditional hierarchical generative adversarial network performs the following steps during runtime: The condition combination module extracts features from each input condition through a lightweight feature extraction network and concatenates them to obtain a condition fusion vector. The conditional fusion vector and random noise are input into the generator module to generate simulated samples; The conditions, the simulated samples, and the real samples are input into a condition discriminator corresponding to the conditions to determine the authenticity of the simulated samples and whether they meet the conditions. The conditional fusion vector, the simulated sample, and the real sample are input into the overall discriminator to determine the authenticity of the simulated sample and whether it conforms to the conditional fusion vector.
[0019] Furthermore, the conditional hierarchical generative adversarial network of this method can fully consider the influence of different conditions, improve the quality and accuracy of generated samples, and thus improve the accuracy of water quality prediction.
[0020] Optionally, the step of predicting the future water quality level of the tested water body based on the predicted values of each of the water quality evaluation indicators, and triggering an early warning when the early warning conditions are met, includes the following steps: Based on the predicted values of the water quality evaluation indicators, the water quality grade is predicted using the single-factor evaluation method. The water quality requirements for the water bodies in the monitored watershed are determined, and an early warning is triggered when the water quality evaluation indicators and the water quality level do not meet the water quality requirements.
[0021] Furthermore, this method uses a single-factor evaluation method to predict water quality levels, which can quickly obtain the prediction results of water quality levels. Based on the water quality prediction results and the water quality requirements of the water bodies in the monitored watershed, early warnings can be triggered in a timely manner, which is conducive to taking timely measures to deal with water quality problems and ensure water environment safety.
[0022] Secondly, this invention provides a water quality monitoring and prediction system based on multi-factor correlation analysis. The system includes a data acquisition device, a data output device, a processor, and a storage device. The storage device includes a computer-readable storage medium storing a computer program. The computer program includes program instructions, which, when executed by the processor, cause the processor to implement the water quality monitoring and prediction method based on multi-factor correlation analysis provided by this invention. This system improves the practicality of the method and facilitates its promotion. Attached Figure Description
[0023] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 This is a flowchart illustrating a water quality monitoring and prediction method based on multi-factor correlation analysis according to an embodiment of the present invention. Figure 2 This is a simplified structural diagram of a conditional hierarchical generative adversarial network according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the framework of a water quality monitoring and prediction system based on multi-factor correlation analysis according to an embodiment of the present invention. Detailed Implementation
[0025] Specific embodiments of the present invention will now be described in detail. It should be noted that the embodiments described herein are for illustrative purposes only and are not intended to limit the invention. In the following description, numerous specific details are set forth in order to provide a thorough understanding of the invention. However, it will be apparent to those skilled in the art that these specific details are not necessary to practice the invention. In other instances, well-known circuits, software, or methods have not been specifically described to avoid obscuring the invention.
[0026] Throughout this specification, references to "an embodiment," "an embodiment," "an example," or "an example" mean that a particular feature, structure, or characteristic described in connection with that embodiment or example is included in at least one embodiment of the invention. Therefore, the phrases "in an embodiment," "in an embodiment," "an example," or "an example" appearing in various places throughout the specification do not necessarily refer to the same embodiment or example. Furthermore, specific features, structures, or characteristics can be combined in one or more embodiments or examples in any suitable combination and / or sub-combination. Moreover, those skilled in the art will understand that the illustrations provided herein are for illustrative purposes and are not necessarily drawn to scale.
[0027] It should be noted in advance that, in one alternative embodiment, except for independent descriptions, the same symbols or letters appearing in all formulas have the same meaning and value.
[0028] In one optional embodiment, please refer to Figure 1 This invention provides a water quality monitoring and prediction method based on multi-factor correlation analysis, the method comprising the following steps: S1. Determine water quality evaluation indicators and collect water quality evaluation indicator data of water bodies in the monitoring basin, thereby evaluating the water quality level and obtaining the water quality monitoring dataset of water bodies in the monitoring basin.
[0029] Step S1 specifically includes the following steps: S11. Determine water quality evaluation indicators, including pH value, dissolved oxygen concentration, chemical oxygen demand, five-day biochemical oxygen demand, ammonia nitrogen, total nitrogen, total phosphorus, permanganate index, lead concentration, mercury concentration, and fluoride.
[0030] Specifically, in this embodiment, the water quality evaluation indicators for the monitored water bodies include, but are not limited to, pH value, dissolved oxygen concentration, chemical oxygen demand (COD), five-day biochemical oxygen demand (BOD5), ammonia nitrogen, total nitrogen, total phosphorus, permanganate index, lead concentration, mercury concentration, and fluoride. These indicators were chosen because they are representative of common water quality monitoring practices and can reflect the water body's pH, redox state, organic pollution level, nutrient content, and heavy metal pollution, which are crucial for assessing the water quality of the monitored watershed. More water quality evaluation indicators can be found in the current "Surface Water Environmental Quality Standard (GB3838—2002)".
[0031] S12. Collect water quality evaluation index data of the water body and preprocess the water quality evaluation index data to obtain time series data of the water quality evaluation index data.
[0032] Specifically, in this embodiment, water quality evaluation index data of the monitored water body are collected in real time using different types of water quality monitoring sensors. The collected data undergoes preprocessing operations such as outlier removal, missing value filling, discretization, and normalization to obtain the time series data of each water quality evaluation index. Preprocessing the collected data improves data quality, thereby enhancing the accuracy of water quality monitoring and prediction.
[0033] S13. After data preprocessing, the water quality level is evaluated using the current water quality evaluation index data, and a water quality monitoring dataset is constructed using all the time series data.
[0034] Specifically, in this embodiment, the current "Surface Water Environmental Quality Standard (GB3838-2002)" classifies surface water into five categories based on the function and protection objectives of the water area. Therefore, the water quality classification method of the current "Surface Water Environmental Quality Standard (GB3838-2002)" can be used to evaluate the water quality level using the single-factor evaluation method based on the current water quality evaluation index data, that is, to determine which category the water in the monitoring basin belongs to.
[0035] Furthermore, a sliding window method is used to extract the time series data of each water quality evaluation indicator into multiple subsequences of the same length to construct a water quality monitoring dataset, providing a reliable data foundation for water quality prediction. The window length is set to 24, and the sliding step size to 4.
[0036] S2. Based on the water quality monitoring dataset, the correlation between the various water quality evaluation indicators is determined through correlation analysis, and then data support sets for each of the water quality evaluation indicators are constructed respectively.
[0037] This step uses correlation analysis to obtain the interrelationships between various indicators, thereby classifying the water quality evaluation indicators and constructing data support sets for each. This facilitates the application of different analysis methods to different types of indicators and provides a data foundation for indicator prediction, improving the accuracy of water quality prediction. Step S2 specifically includes the following steps: S21. Based on the water quality monitoring dataset, calculate the comprehensive correlation between each water quality evaluation indicator, and then classify the water quality evaluation indicators into single-factor indicators, multi-factor indicators, and isolated indicators.
[0038] Specifically, step S21 includes the following steps: S211. Calculate the average of the maximum information coefficient and time-delay correlation between each of the water quality evaluation indicators, and use it as the comprehensive correlation between each of the water quality evaluation indicators.
[0039] Specifically, in this embodiment, when calculating the time-lag correlation between the various water quality evaluation indicators, a time lag range of 5 days is selected, and the time-lag correlation is represented by the Pearson correlation coefficient. The calculation method of the maximum information coefficient and the time-lag correlation is a prior art method and will not be described in detail here.
[0040] More specifically, the overall correlation satisfies the following relationship: Where R represents the comprehensive correlation, MIC represents the maximum information coefficient, and r represents the time-lag correlation. This method calculates the comprehensive correlation between various water quality evaluation indicators, taking into account both linear and nonlinear correlations, as well as time-lag relationships, making the correlation analysis more comprehensive and accurate, and thus enabling rapid and accurate classification of the indicators.
[0041] S212. For any water quality evaluation index A, if there are other water quality evaluation indexes B whose comprehensive correlation with it is greater than the correlation threshold, then B is an associated index of A.
[0042] Specifically, in this embodiment, the correlation threshold needs to be set for different water bodies. Generally, if the correlation threshold is set too high, it will reduce the accuracy of water quality prediction; however, if it is set too low, it will increase the amount of data processing, which is not conducive to improving the efficiency of water quality prediction. Therefore, the correlation threshold should not be too high or too low. Based on this consideration, this embodiment sets the correlation threshold to 0.5.
[0043] S213. For any water quality evaluation indicator A, if it has one of the associated indicators, then A is the single-cause indicator; if it has multiple associated indicators, then A is the multi-cause indicator; if it does not have any associated indicators, then A is the isolated indicator.
[0044] S22. Based on the water quality monitoring dataset, construct corresponding data support sets for each of the single-factor indicators, multi-factor indicators, and isolated indicators.
[0045] This step constructs separate data support sets for different water quality assessment indicators, facilitating the application of different prediction methods for different types of indicators and providing a data foundation for indicator prediction, thereby improving the accuracy of water quality prediction. Step S22 specifically includes the following steps: S221. For any of the single-factor indicators, extract the time series data of the single-factor indicator and the time series data of its related indicators from the water quality monitoring dataset to construct the data support set of the single-factor indicator.
[0046] S222. For any of the multi-factor indicators, select multiple indicators from its associated indicators as its main associated indicators according to the magnitude of the comprehensive correlation.
[0047] Specifically, in this embodiment, for any multi-factor indicator, the two most correlated indicators with which it has the strongest overall correlation are selected as its primary correlated indicators. For multi-factor indicators, selecting two primary correlated indicators not only enables joint prediction of the indicators but also prevents excessive data processing, thereby improving prediction efficiency to some extent. In other alternative embodiments, other numbers of primary correlated indicators may be selected, or no primary correlated indicators may be selected.
[0048] S223. Extract the time series data of the multi-factor indicators and their main related indicators from the water quality monitoring dataset, and then construct the data support set of the multi-factor indicators.
[0049] Specifically, in this embodiment, for any multifactor indicator, the time series data of the multifactor indicator and its main related indicators are extracted from the water quality monitoring dataset, and then a data support set for the multifactor indicator is constructed.
[0050] S224. For any of the isolated indicators, extract its time series data from the water quality monitoring dataset to construct its data support set.
[0051] S3. Based on the data support set of the water quality evaluation indicators, use the indicator prediction model to obtain the predicted values of the water quality evaluation indicators.
[0052] The indicator prediction models include three types: Type I, Type II, and Type III. This step involves classifying and modeling different types of water quality assessment indicators to fully leverage the advantages of each model, improve the accuracy of water quality indicator prediction, and thus enhance the overall accuracy of water quality forecasting. Step S3 specifically includes the following steps: S31. For the single-cause index, a first-class index prediction model is constructed using its data support set and conditional generative adversarial network, thereby realizing the prediction of the single-cause index.
[0053] Specifically, in this embodiment, for single-cause indicators, the data support set is strictly divided into training set, validation set and test set according to time order and a ratio of 7:1:2 to ensure that the training set data is earlier than the validation set data, and the validation set data is earlier than the test set data, thus preventing future information leakage.
[0054] Furthermore, a conditional generative adversarial network (GAN) based on Wasserstein distance and gradient penalty is constructed. The time-series data of the associated indicators of the single-cause indicator are used as conditional inputs to the generator and discriminator. The generator produces predicted values for the future single-cause indicator under given conditions. The core network architecture of the generator and discriminator in the GAN employs models capable of effectively capturing temporal dependencies, such as Long Short-Term Memory (LSTM) networks. The GAN uses the Wasserstein distance loss function and introduces a gradient penalty term to stabilize the training process and prevent mode collapse. The discriminator aims to maximize the Wasserstein distance estimate between real and generated sample pairs, while the generator aims to minimize this distance estimate.
[0055] Furthermore, the conditional generative adversarial network (GAN) is trained, validated, and tested using training, validation, and test sets, respectively. During training, mean squared error, mean absolute error, accuracy, and recall are used to evaluate the prediction results, and the final generator is used as the first-class indicator prediction model. A corresponding first-class indicator prediction model needs to be constructed for each single-cause indicator. After obtaining the first-class indicator prediction model, the time series data of the associated indicators of the single-cause indicator are used as conditional inputs to generate N samples. The mean of the N samples is taken as the final predicted value of the single-cause indicator.
[0056] S32. For the multifactor index, perform multivariate empirical mode decomposition on each time series data in its data support set, and further obtain the mode decomposition set of the multifactor index.
[0057] Specifically, step S32 includes the following steps: S321. For the multifactor index, multivariate empirical mode decomposition is used to decompose the time series data of each water quality evaluation index in the same period of its data support set to obtain the IMF components and residual terms of the multifactor index and its main related indicators.
[0058] Specifically, in this embodiment, for multifactor indicators, multivariate empirical mode decomposition (EMD) is used to decompose each set of time series data in its data support set. In the data support set of multifactor indicators, a set of time series data contains time series data of multifactor indicators and their main related indicators for the same time period. Through EMD, the various IMF components and residual terms of multifactor indicators and their main related indicators can be obtained. This can effectively extract different frequency components of the data, providing richer information for subsequent analysis and prediction.
[0059] S322. The IMF components of the multifactor index and its main related index are divided into high-frequency components and low-frequency components by sample entropy.
[0060] Specifically, in this embodiment, the sample entropy value is calculated for the IMF component of each indicator obtained in step S321, and a sample entropy threshold is set to divide the IMF component into high-frequency and low-frequency components. Specifically, a sample entropy threshold is set for different water quality evaluation indicators. For a specific time series data of a water quality evaluation indicator, if the sample entropy of its IMF component is not less than its sample entropy threshold, the IMF component is determined to be a high-frequency sub-component; otherwise, it is a low-frequency sub-component. This step, by dividing the IMF component into high-frequency and low-frequency components through sample entropy, helps to distinguish different features in the data and provides a basis for subsequent conditional vector construction.
[0061] More specifically, when determining the sample entropy threshold for different water quality assessment indicators, the sample entropy can be calculated for all IMF components of the indicator first, and then a histogram of the sample entropy can be plotted to manually select the sample entropy threshold. In other alternative embodiments, a fixed sample entropy threshold can be used for all water quality assessment indicators, thereby reducing the amount of data processing, but this will also negatively impact the accuracy of water quality prediction.
[0062] Furthermore, for any water quality evaluation index, all its high-frequency sub-components are added together to obtain the high-frequency component of that water quality evaluation index. Similarly, its low-frequency component can be obtained.
[0063] S323. Using a weighted summation method, the high-frequency components, low-frequency components, and residual terms of each major related indicator of the multi-factor index are fused into a high-frequency condition vector, a low-frequency condition vector, and a residual condition vector.
[0064] Specifically, in this embodiment, the comprehensive correlation between the multi-factor index and its various main related indicators is normalized and used as a reference weight for fusing the high-frequency components, low-frequency components, and residual terms of the main related indicators. Then, a weighted summation method is used to fuse the high-frequency components, low-frequency components, and residual terms of each main related indicator of the multi-factor index into a high-frequency condition vector, a low-frequency condition vector, and a residual condition vector, respectively. The high-frequency condition vector satisfies the following relationship: in, This is a high-frequency condition vector, where N is the number of the main related indicators of the multi-factor index. The reference weight for the i-th primary related indicator. Let be the high-frequency component of the i-th primary correlation index. The calculation methods for the low-frequency condition vector and residual condition vector can refer to the calculation method for the high-frequency condition vector, and will not be elaborated here.
[0065] This step uses a weighted summation method to fuse the high-frequency components, low-frequency components, and residual terms of each major related indicator of the multi-factor index into a high-frequency condition vector, a low-frequency condition vector, and a residual condition vector, respectively. This facilitates a comprehensive consideration of the impact of different components of each major related indicator on the major related indicators, thereby improving the performance of the prediction model.
[0066] S324. Construct the mode decomposition set using the time series data, high-frequency components, low-frequency components, and residual terms of the multifactor index, as well as the high-frequency condition vector, the low-frequency condition vector, and the residual condition vector.
[0067] Specifically, in this embodiment, for any multifactor index, a mode decomposition set is constructed using the high-frequency components, low-frequency components, residual terms, high-frequency conditional vectors, low-frequency conditional vectors, and residual conditional vectors corresponding to its different time series data, providing comprehensive and rich data support for the subsequent training and validation of conditional hierarchical generative adversarial networks.
[0068] S33. Use the modality decomposition set of the multi-factor index to complete the training and validation of the conditional hierarchical generative adversarial network, thereby obtaining the second type of index prediction model, which is used to predict the multi-factor index.
[0069] Specifically, in this embodiment, standard CGAN typically uses a simple concatenation method to input conditional information into the generator and discriminator. This approach may be effective when dealing with low-dimensional, homogeneous, and semantically simple conditions. However, under complex conditional inputs with multiple factors, multiple scales, and clear physical meanings, the simple concatenation of conditional information of different types and scales leads to information mixing, making it difficult for the generator and discriminator to effectively extract and utilize this information, resulting in feature entanglement and affecting the quality of the generated results. The single discriminator of standard CGAN is unable to impose fine-grained, multi-scale conditional constraints on complex conditional inputs. In water quality prediction scenarios, different conditions may have different degrees and ways of affecting the prediction results, and a single discriminator cannot effectively judge and respond to these differences, leading to large errors in the prediction results.
[0070] Please see Figure 2 To improve the accuracy of water quality assessment indicator prediction and thus the accuracy of water quality assessment, this embodiment uses a conditional hierarchical generative adversarial network. The conditional hierarchical generative adversarial network includes a condition combination module, a generator module, and a four-level discriminant module. The four-level discriminant module includes three condition discriminators and a master discriminator. The condition discriminators include a first condition discriminator, a second condition discriminator, and a third condition discriminator, each corresponding to one condition.
[0071] More specifically, when predicting multi-factor indicators, the corresponding high-frequency condition vector, low-frequency condition vector, and residual condition vector are used as three conditions. The input conditions for the first, second, and third condition discriminators are the high-frequency condition vector, low-frequency condition vector, and residual condition vector, respectively. The operation flow of the condition-level generative adversarial network is as follows: First, the condition combination module extracts the features of the three input conditions through three lightweight feature extraction networks (1D CNNs) and concatenates them to obtain a condition fusion vector. Then, the condition fusion vector and random noise are input into the generator module to generate simulated samples of multi-factor indicators. Next, the conditions, simulated samples of multi-factor indicators, and real samples are input into the condition discriminators corresponding to the conditions to determine the authenticity of the simulated samples and whether they meet the corresponding conditions. At the same time, the condition fusion vector, simulated samples of multi-factor indicators, and real samples are input into the overall discriminator to determine the authenticity of the simulated samples and whether they meet the condition fusion vector.
[0072] Through adversarial training between the generator module and the fourth-level discriminator module, the generator module continuously optimizes its ability to generate simulated samples, making them more consistent with the distribution and constraints of real samples; the fourth-level discriminator module continuously improves its ability to distinguish simulated samples. Furthermore, to improve the stability of model training, a gradient penalty technique is introduced into the conditional hierarchical generative adversarial network. After multiple rounds of training and validation, a stable generator is obtained and used as the second-class indicator prediction model to predict the future values of multi-factor indicators. The conditional hierarchical generative adversarial network in this embodiment can fully consider the influence of different conditions, improve the quality and accuracy of generated samples, and thus improve the accuracy of water quality prediction.
[0073] In a conditional hierarchical generative adversarial network, the loss functions of the generator and discriminator satisfy the following relationships: in, Let z be the generator loss; E be the expected value, representing the average over all possible z and c; z be random noise from the prior distribution. Mid-sampling is used to introduce randomness; c is a condition, derived from the conditional distribution. Medium sampling; M is the number of conditions; For the j-th conditional discriminator to generate samples and the j-th condition The judgment result; For the total discriminator to generate samples Conditional fusion vector The judgment result; The loss of the j-th conditional discriminator is calculated using existing techniques. This is the weight coefficient for the gradient penalty, usually taken as 10; The gradient penalty for the j-th conditional discriminator; To obtain real samples, samples are taken from the data distribution p; For the j-th conditional discriminator on the real sample and the j-th condition The judgment result, The loss of the total discriminator; The gradient penalty for the overall discriminator is calculated in accordance with existing techniques. For the overall discriminator to the real samples Conditional fusion vector The judgment result; For the total discriminator to generate samples Conditional fusion vector The judgment result.
[0074] S34. For the isolated indicator, a third type of indicator prediction model is constructed using its data support set and generative adversarial network, thereby realizing the prediction of the isolated indicator.
[0075] Specifically, in this embodiment, since isolated indicators lack corresponding conditions, their data support set is directly used to complete the training, validation, and testing of the generative adversarial network (GAN). The resulting generator is then used as a third-class indicator prediction model to predict the future values of isolated indicators. The core network architecture of the GAN's generator and discriminator employs a model capable of effectively capturing temporal dependencies, such as a Long Short-Term Memory (LSTM) network.
[0076] S4. Based on the predicted values of each of the water quality evaluation indicators, predict the future water quality level of the detected water body, and trigger an early warning when the early warning conditions are met.
[0077] Step S4 specifically includes the following steps: S41. Based on the predicted values of the water quality evaluation indicators, use the single-factor evaluation method to predict the water quality level.
[0078] Specifically, in this embodiment, based on the predicted values of water quality evaluation indicators, the future water quality level can be quickly obtained using the single-factor evaluation method, referring to the water quality classification method of the current "Surface Water Environmental Quality Standard (GB3838-2002)".
[0079] S42. Determine the water quality requirements for the water bodies in the monitored watershed, and trigger an early warning when the water quality evaluation indicators and the water quality level do not meet the water quality requirements.
[0080] Specifically, in this embodiment, the water quality grade requirements of the monitored water body and the value range of various water quality evaluation indicators are determined. Then, when the water quality evaluation indicators and water quality grade do not meet the water quality requirements, an early warning is triggered, which is conducive to taking timely measures to deal with water quality problems and ensure water environment safety.
[0081] It should be noted that in some cases, the actions described in the specification can be performed in different orders and still achieve the desired results. In this embodiment, the order of steps is given only to make the embodiment clearer and easier to explain, and not to limit it.
[0082] In an optional embodiment, please refer to Figure 3. To improve the practicality of this method and facilitate its promotion, the present invention also provides a water quality monitoring and prediction system based on multi-factor correlation analysis. The water quality monitoring and prediction system based on multi-factor correlation analysis includes: a data acquisition device 1, a data output device 2, a processor 3, and a storage device 4. The storage device 4 includes a computer-readable storage medium storing a computer program. The computer program includes program instructions, which, when executed by the processor 3, cause the processor 3 to perform the contents described in steps S1 to S4.
[0083] In summary, the present invention has at least the following beneficial effects: First, this method comprehensively considers the linear and nonlinear correlations between indicators, as well as the time lag relationship, and classifies water quality evaluation indicators into single-factor indicators, multi-factor indicators, and isolated indicators, and constructs data support sets for each, providing a data foundation for indicator prediction and helping to improve the accuracy of water quality prediction.
[0084] Secondly, this method employs different analytical approaches and categorizes and models different types of water quality assessment indicators, fully leveraging the advantages of various models to improve the accuracy of water quality indicator predictions, thereby enhancing the overall accuracy of water quality forecasts. Specifically, the conditional hierarchical generative adversarial network constructed for multi-factor indicators fully considers the influence of different conditions, improving the quality and accuracy of generated samples and enabling joint prediction among indicators, thus further enhancing the accuracy of water quality forecasts.
[0085] Furthermore, this method uses a single-factor evaluation method to predict water quality levels, which can quickly obtain the predicted water quality levels. Based on the water quality prediction results and the water quality requirements of the water bodies in the monitored basin, early warnings can be triggered in a timely manner, which is conducive to taking timely measures to deal with water quality problems and ensure water environment safety.
[0086] Finally, the present invention provides a system adapted to the method, which can improve the practicality of the method and facilitate its promotion.
[0087] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be covered within the scope of the claims and specification of the present invention.
Claims
1. A water quality monitoring and prediction method based on multi-factor correlation analysis, characterized in that, Includes the following steps: Determine water quality evaluation indicators, collect water quality evaluation indicator data of water bodies in the monitored watershed, and then evaluate the water quality level and obtain the water quality monitoring dataset of the water bodies in the monitored watershed; Based on the water quality monitoring dataset, correlation analysis is used to determine the correlation between the various water quality evaluation indicators, and then data support sets for each of the water quality evaluation indicators are constructed respectively. Based on the data support set of the water quality evaluation indicators, the predicted values of the water quality evaluation indicators are obtained using the indicator prediction model; The future water quality level of the tested water body is predicted based on the predicted values of each of the aforementioned water quality evaluation indicators, and an early warning is triggered when the early warning conditions are met.
2. The water quality monitoring and prediction method based on multi-factor correlation analysis according to claim 1, characterized in that, The process of determining water quality evaluation indicators, collecting water quality evaluation indicator data of water bodies in the monitored watershed, and then evaluating the water quality level and obtaining the water quality monitoring dataset of the water bodies in the monitored watershed includes the following steps: Determine water quality evaluation indicators, including pH value, dissolved oxygen concentration, chemical oxygen demand, five-day biochemical oxygen demand, ammonia nitrogen, total nitrogen, total phosphorus, permanganate index, lead concentration, mercury concentration, and fluoride; Collect water quality evaluation index data of the water body and preprocess the water quality evaluation index data to obtain time series data of the water quality evaluation index data; After data preprocessing, the water quality level is evaluated using the current water quality evaluation index data, and a water quality monitoring dataset is constructed using all the time series data.
3. The water quality monitoring and prediction method based on multi-factor correlation analysis according to claim 1, characterized in that, The step of determining the correlation between various water quality evaluation indicators based on the water quality monitoring dataset, and then constructing data support sets for each water quality evaluation indicator, includes the following steps: Based on the water quality monitoring dataset, the comprehensive correlation between each water quality evaluation index is calculated, and then the water quality evaluation index is divided into single-factor indexes, multi-factor indexes, and isolated indexes. Based on the water quality monitoring dataset, corresponding data support sets are constructed for each of the single-factor indicators, multi-factor indicators, and isolated indicators.
4. The water quality monitoring and prediction method based on multi-factor correlation analysis according to claim 3, characterized in that, The step of calculating the comprehensive correlation between various water quality evaluation indicators based on the water quality monitoring dataset, and then classifying the water quality evaluation indicators into single-factor indicators, multi-factor indicators, and isolated indicators, includes the following steps: The mean of the maximum information coefficient and time-lag correlation between each of the water quality evaluation indicators is calculated and used as the comprehensive correlation between each of the water quality evaluation indicators. For any water quality evaluation index A, if there are other water quality evaluation index B whose overall correlation with it is greater than the correlation threshold, then B is an associated index of A. For any water quality evaluation indicator A, if it has one of the aforementioned related indicators, then A is the single-cause indicator; if it has multiple of the aforementioned related indicators, then A is the multi-cause indicator; if it does not have any of the aforementioned related indicators, then A is the isolated indicator.
5. The water quality monitoring and prediction method based on multi-factor correlation analysis according to claim 4, characterized in that, The step of constructing corresponding data support sets for each of the single-factor indicators, multi-factor indicators, and isolated indicators based on the water quality monitoring dataset includes the following steps: For any of the single-factor indicators, the time series data of the single-factor indicator and the time series data of its related indicators are extracted from the water quality monitoring dataset to construct the data support set of the single-factor indicator. For any of the multi-factor indicators, select multiple indicators from its associated indicators as its main associated indicators according to the magnitude of the comprehensive correlation. The time series data of the multifactor indicators and their main related indicators are extracted from the water quality monitoring dataset, and then a data support set for the multifactor indicators is constructed. For any of the isolated indicators, time series data are extracted from the water quality monitoring dataset to construct its data support set.
6. The water quality monitoring and prediction method based on multi-factor correlation analysis according to claim 4, characterized in that: The indicator prediction model includes a first type of indicator prediction model, a second type of indicator prediction model, and a third type of indicator prediction model; The step of obtaining the predicted values of the water quality evaluation indicators using an indicator prediction model based on the data support set of the water quality evaluation indicators includes the following steps: For the single-cause index, a first-class index prediction model is constructed using its data support set and conditional generative adversarial network, thereby realizing the prediction of the single-cause index; For the multifactor index, multivariate empirical mode decomposition is performed on each time series data in its data support set, and the mode decomposition set of the multifactor index is further obtained. The modality decomposition set of the multi-factor indicators is used to complete the training and validation of the conditional hierarchical generative adversarial network, thereby obtaining the second type of indicator prediction model, which is used to predict the multi-factor indicators. For the isolated indicator, a third type of indicator prediction model is constructed using its data support set and generative adversarial network, thereby realizing the prediction of the isolated indicator.
7. The water quality monitoring and prediction method based on multi-factor correlation analysis according to claim 6, characterized in that, The step of performing multivariate empirical mode decomposition on each time series data in the data support set for the multifactor index, and further obtaining the mode decomposition set of the multifactor index, includes the following steps: For the multifactor indicators, multivariate empirical mode decomposition is used to decompose the time series data of each water quality evaluation indicator in the same period of its data support set to obtain the IMF components and residual terms of the multifactor indicators and their main related indicators. The IMF components of the multifactor index and its main related indexes are divided into high-frequency components and low-frequency components by using sample entropy. The high-frequency components, low-frequency components, and residual terms of each major correlation indicator of the multi-factor index are fused into a high-frequency condition vector, a low-frequency condition vector, and a residual condition vector by using a weighted summation method. The mode decomposition set is constructed using the time series data, high-frequency components, low-frequency components, and residual terms of the multifactor index, as well as the high-frequency condition vector, the low-frequency condition vector, and the residual condition vector.
8. The water quality monitoring and prediction method based on multi-factor correlation analysis according to claim 6, characterized in that: The conditional hierarchical generative adversarial network includes a condition combination module, a generator module, and a four-level discriminant module. The four-level discriminant module includes three condition discriminators and a total discriminator. The condition discriminators include a first condition discriminator, a second condition discriminator, and a third condition discriminator. The first condition discriminator, the second condition discriminator, and the third condition discriminator each correspond to a condition. The conditional hierarchical generative adversarial network performs the following steps during runtime: The condition combination module extracts features from each input condition through a lightweight feature extraction network and concatenates them to obtain a condition fusion vector. The conditional fusion vector and random noise are input into the generator module to generate simulated samples; The conditions, the simulated samples, and the real samples are input into a condition discriminator corresponding to the conditions to determine the authenticity of the simulated samples and whether they meet the conditions. The conditional fusion vector, the simulated sample, and the real sample are input into the overall discriminator to determine the authenticity of the simulated sample and whether it conforms to the conditional fusion vector.
9. A water quality monitoring and prediction method based on multi-factor correlation analysis according to claim 1, characterized in that, The step of predicting the future water quality level of the tested water body based on the predicted values of each of the water quality evaluation indicators, and triggering an early warning when the early warning conditions are met, includes the following steps: Based on the predicted values of the water quality evaluation indicators, the water quality grade is predicted using the single-factor evaluation method. The water quality requirements for the water bodies in the monitored watershed are determined, and an early warning is triggered when the water quality evaluation indicators and the water quality level do not meet the water quality requirements.
10. A water quality monitoring and prediction system based on multi-factor correlation analysis, characterized in that, The water quality monitoring and prediction system based on multi-factor correlation analysis includes: a data acquisition device, a data output device, a processor, and a storage device. The storage device includes a computer-readable storage medium storing a computer program. The computer program includes program instructions, which, when executed by the processor, cause the processor to implement the water quality monitoring and prediction method based on multi-factor correlation analysis as described in any one of claims 1-9.
Citation Information
Patent Citations
Multi-index associated river water quality evaluation method and system, equipment and storage medium
CN114037202A
Water quality prediction and early warning system based on machine learning method
CN115829120A
Remote sensing image analysis and cyanobacterial bloom prediction method based on four-dimensional generative adversarial network
CN116403103A
Sewage quality multi-index prediction method
CN117236510A
Underground water quality factor dynamic migration analysis method and system
CN119207635A
Cited By
Surface water resource quality dynamic prediction and early warning method
CN122048578A