A method for identifying and warning water quality abnormal events based on spatio-temporal data of pipe network water quality
By constructing an adversarial learning network model (GAN) and timing Bayesian principle in the water supply network, identifying and warning of water quality abnormalities, the problems of low detection accuracy and high false alarm rate in the existing technology are solved, and more efficient monitoring and early warning of water quality abnormalities accidents are achieved.
Patent Information
- Application Number
- CN202211104588.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-09
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-09-09
AI Technical Summary
In actual application, the existing water quality abnormality accident detection methods for water supply networks have problems such as low model accuracy, low pollution detection rate, and high false alarms, especially in complex water supply systems where hydraulic water quality models cannot be accurately obtained.
The early warning method for identifying and warning water quality abnormal events based on the temporal and spatial data of the pipeline network water quality, and standardized preprocessing and superimposed image conversion of the water quality data of multiple sensor sites is used to build an adversarial learning network model (GAN), and the abnormal score is calculated using the GAN model, and the probability of water quality abnormal events is calculated in combination with the time-sequence Bayesian principle.
The correlation analysis of the spatio-temporal distribution characteristics of water quality in complex water supply systems is realized, the accuracy of detection of pollution accidents is improved, the number of false alarms is reduced, and it has strong robustness and scope of application.
Smart Images

Figure CN115470850B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of water treatment in water supply networks, and relates to a method for identifying and warning water quality abnormal events based on spatio-temporal data of network water quality. Background Technique
[0002] Water supply networks are important urban infrastructures that safely and reliably deliver water to users. However, especially in developing countries like China, pollution accidents often occur in water supply networks due to network aging, lack of scheduling and maintenance management, and poor construction quality. When a pollution accident occurs in a water supply network, unless the pollution can be quickly detected and removed, the contaminated water will rapidly spread throughout the entire water supply network, forming a global risk. This will not only cause water supply interruption and huge economic losses, but also cause environmental damage and public health and safety problems. Therefore, quickly and accurately detecting pollution events is crucial for ensuring urban water supply safety, which helps water service groups formulate remedial measures, reduce losses caused by pollution events, and improve the water supply level and social acceptance of water service groups.
[0003] Traditionally, detecting specific substances in drinking water is carried out through on-site sampling and laboratory analysis. This method can test various types of water quality parameters or directly identify pollutants. However, for large-scale water supply networks, this method is very time-consuming and laborious, and most importantly, it cannot provide early warnings of pollution events in a timely manner. Currently, some water service groups in Chinese cities have established relatively complete sensor monitoring networks (SCADA systems) in the local water supply systems. By using online water quality monitoring sensors, the changes in conventional water quality parameters at key nodes in the network can be monitored in real time. On this basis, rapid analysis and response to water quality accidents in water supply networks can be carried out to improve the water supply service level of the company.
[0004] Research on the identification of water quality pollution accidents can be mainly divided into methods based on statistical analysis models, hydraulic models, and machine learning models. In the methods based on statistical analysis models, the detection of pollution accidents is usually based on the distribution of water quality parameter data in the water supply network. Due to the non-linear and non-stationary characteristics of water quality data, statistical methods are usually not applicable to detecting minor abnormal changes in the water quality of the water supply network. The method based on the hydraulic model detects pollution accidents by comparing the observed real-time water quality data with the predicted values obtained using the hydraulic water quality model of the pipe network. However, due to the complexity of the pipe network topology and data limitations, it is difficult to obtain an accurate hydraulic water quality model in practical applications. Machine learning algorithms are considered as alternative methods for real-time predicting changes in water quality parameters and identifying pollution accidents. Various machine learning algorithms have been applied to the detection of water quality abnormal accidents in the water supply network, such as artificial neural network (ANN), support vector machine (SVM), long short-term memory (LSTM), etc. These models can capture the characteristics of the water quality time series data at a single sensor site. However, these models do not utilize the spatial distribution relationship of sensor data at multiple sites, and when there are large hydraulic operation changes during the normal operation of water quality monitoring sites, there will be more false alarms. Generally speaking, the existing models have problems such as low model accuracy, low detection rate of pollution detection accidents, and more false alarms in practical applications. Summary of the Invention
[0005] To address the above deficiencies, the problem to be solved by the present invention is to provide a method for identifying and warning water quality abnormal accidents in a water supply network, which can perform correlation analysis on multi-source water quality data in the spatial range in a complex water supply system where it is impossible to accurately obtain a hydraulic water quality model, obtain the characteristics of the spatio-temporal distribution of the water quality of the water supply network, realize the real-time automation of water quality pollution accident monitoring, improve the detection rate of the model for pollution accidents, and reduce the number of false alarms of the model.
[0006] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0007] A method for identifying and warning water quality abnormal events based on spatio-temporal data of pipe network water quality, comprising the following steps:
[0008] (1) Select the monitored water quality parameters and water quality sensor sites, and by analyzing the time interval and change trend of the sensors receiving polluted water, group and analyze adjacent monitoring sites in the water supply network with the same change trend. Divide the N sensor detection sites into N + 1 groups: in the first N groups, each group contains the multi-parameter water quality monitoring data of a single sensor site, and there is another group containing the multi-parameter water quality monitoring data of all N sensor sites.
[0009] (2) The water quality data of group N+1 are standardized and preprocessed, and the water quality time series data in each group after preprocessing are converted into superimposed images. Assuming that each station collects Nr water quality parameters, at each moment, V water quality parameter data (V = Nr × N) can be collected from the N sensor stations analyzed. At time t, the distance image of each group can be expressed as in is the superimposed image m at time t t The element in the i-th row and j-th column is calculated as the sum of the i-th and j-th data of the V water quality parameter data collected at time t. When N = 1, the stacking image conversion is only used for multi-parameter data fusion of a single sensor site; when N>1, its stacking image conversion is used to fuse the spatiotemporal distribution data of multiple sites. In order to reduce the impact of noise, the stacking image at each moment takes the average of the past d time lengths.
[0010] (3) Construct the adversarial learning network model GAN.
[0011] The adversarial learning network model is divided into a generator G and a discriminator D. The generator G uses a convolutional neural network including an encoder for compressing image information and a decoder for restoring image information; the discriminator D is a classification convolutional neural network used to compress and extract image features. First, the multivariate time series data of the water quality monitoring station analyzed under normal operation is converted into a superimposed image using step (2). Then, the historical superimposed images of the past K moments are used as the input of the GAN model, and the superimposed image of the next moment is generated through the generator network G. Then, the image generated by the generator is converted into a superimposed image. The superimposed image m based on the measured data is input to the discriminator network. The discriminator is used to distinguish the superimposed image m (real) of the measured data from the superimposed image reconstructed by the generator G. (Forgery).
[0012] The training of the GAN model uses the improved W-GAN loss as the training loss L of the discriminator D , used to stabilize the training process, specifically expressed as,
[0013]
[0014] Where E represents the mathematical expectation of the calculation; D(·) is the feature vector finally output by the discriminator D; is the superimposed image reconstructed by the generator; m is the superimposed image based on the measured data; It is an interpolated image composed of a pair of real superimposed images and a reconstructed superimposed image; is the interpolated image Find the gradient; λ GP is the coefficient of the gradient penalty term, which is 10 according to experience.
[0015] The generator training uses the reconstruction loss between the real superimposed image and the generated superimposed image as the generator loss L G , to help the generator learn the distribution of normal water quality data. Its calculation is expressed as,
[0016]
[0017] In the GAN model, the two networks G and D are trained and updated simultaneously. The ultimate goal of model training is not to minimize the loss of any single network, but to find a stable state when the losses of G and D both converge to a stable state.
[0018] (4) Construct the anomaly score ψ based on the generator and discriminator of the GAN model trained in step (3). The anomaly score calculation based on GAN includes two parts: anomaly recognition using the generator network and anomaly recognition using the discriminator network. Among them, the anomaly recognition based on the generator network is calculated by comparing the image generated by the generator network with the reconstruction loss between the superimposed image m constructed based on the measured data. The anomaly recognition based on the discriminator network is calculated by comparing the feature vectors D(m) and obtained from the final outputs after respectively inputting the superimposed image based on the measured data and the image generated by the generator into the discriminator network. The anomaly score finally constructed based on the GAN model is specifically expressed as,
[0019]
[0020] where ψ(t) represents the anomaly score calculated using the GAN model at time t. The closer the anomaly score is to 0, the more normal the current water quality state is. On the contrary, the larger the anomaly score, the greater the difference between the current water quality state and the normal operating water quality, and the more likely the water quality is in an abnormal state at that moment. λ s is a weighted parameter that adjusts the relative importance of the reconstruction loss and the feature loss with respect to the anomaly score.
[0021] (5) Select an appropriate anomaly score threshold ψ thre , and when the anomaly score obtained in step (4) exceeds ψ thre , it is used as the initial anomaly point recognition. Since the construction and training of the GAN model in step (3) are both based on the water quality data under normal operating conditions, the selection of the anomaly score threshold is mainly determined by the anomaly score distribution obtained from the dataset used to train the GAN network. After obtaining the anomaly scores for the water quality dataset used in the construction and training of the GAN model in step (3), arrange the anomaly score values from small to large, and according to the actual situation of the data distribution, select the 96%-99% percentile points in its distribution as the anomaly score threshold ψthre , so that most of the points in the water quality dataset used to construct the GAN model are within the normal range.
[0022] (6) Use the principle of temporal Bayesian to calculate the probability of water quality abnormal events. Based on the recognition results of the abnormal points obtained in step (5), calculate the probability P(t) of pollution events occurring at each moment. Specifically, it can be expressed by the following formula:
[0023]
[0024] In the formula, TPR is the true positive rate, which is calculated as the ratio of the number of time steps correctly classified as abnormal when the water supply network is polluted to the total number of time steps of pollution. When there is no prior knowledge of pollution events, the value of TPR is 0.5. FPR is the false positive rate, which is calculated as the ratio of the number of time steps identified as abnormal points by the model when the water supply network is operating normally without pollution to the total number of time steps of normal operation. This is equivalent to the ratio of the number of moments exceeding the abnormal score threshold in the training set to the total number of the training dataset.
[0025] The probability of a pollution event occurring at the initial moment is given as P(0). Since pollution events are rare in real life, a small probability value P(0) ∈ [10 -6 , 10 -4 is taken. When the calculated probability P(t) exceeds a certain threshold P thre , the model identifies it as a pollution event. A low probability threshold can increase the event detection rate but may increase the number of false alarms; a high probability threshold can improve the reliability of event alarms and reduce the number of false alarms, but at the same time, the number of detected pollution events may be smaller. The probability threshold P thre for pollution event recognition is set according to the decision maker's trade-off between the event detection rate and the false alarm rate, and is usually set to a probability value exceeding 70%.
[0026] (7) The normal operation of the hydraulic changes in the water supply network will cause sudden changes in water quality parameters in the short term. To distinguish normal water quality changes from pollution events, a simple exponential smoothing model is used to smooth the calculated probability. Specifically, it can be expressed by the following formula:
[0027] P(t) = αP(t) + (1 - α)P(t - 1) (6)
[0028] In the formula, α is the smoothing parameter, which determines the degree of emphasis on the probability of the most recently updated event. α ∈ [0.3, 0.9]. The smaller α is, the slower the event probability is updated, and the more abnormal points are required to be identified to reach the alarm threshold.
[0029] (8) Calculate the probability of water quality abnormal accidents using the single-site model and the multi-site model respectively. The single-site model means applying the anomaly detection methods proposed in steps (3)-(7) separately to the first N groups of water quality data in steps (1)-(2), constructing a GAN model for each site, and taking the maximum value of the calculated pollution event occurrence probabilities of the N GAN models as the event probability P calculated by the single-site model. single The multi-site model means applying the anomaly detection methods proposed in steps (3)-(7) to the (N + 1)-th group of water quality data in steps (1)-(2). This group of data includes the water quality data monitored by N water quality monitoring stations. Construct a GAN model to obtain the pollution event occurrence probability P calculated by the final multi-site model. multi .
[0030] (9) In order to make full use of the water quality relationships between and within monitoring stations, fuse the event probabilities calculated by the single-site model and the multi-site model, and based on the multivariate water quality parameters of all monitoring stations, obtain the combined event probability P reflecting the possibility of pollution events. all , which can be specifically expressed by the following expression:
[0031] P all (t) = ηP single (t) + (1 - η)P multi (t) (7)
[0032] In the formula, P single (t) and P multi (t) respectively represent the pollution event occurrence probabilities calculated using the single-site model and the multi-site model at time t. η is an important weight that adjusts the influence of the single-site model and the multi-site model on the synchronous decision-making of the final pollution event recognition. Both the single-site and multi-site models in the present invention are unsupervised models and do not know the pollution information in advance. Therefore, set η = 0.5 to reflect the same importance of the single-site model and the multi-site model for the final pollution detection and alarm. When the combined event probability P all exceeds the preset probability threshold P thre in step (6), give the alarm signal of the final model and give the probability P of the occurrence of the water quality abnormal accident. all .
[0033] The beneficial effects of the present invention are:
[0034] (1) The water quality abnormal accident detection method proposed by the present invention considers the water quality data information related to multiple sites, and improves the detection accuracy of pollution events by fusing the results of the single-site and multi-site anomaly detection models.
[0035] (2) The water quality anomaly accident detection method proposed by the present invention has strong robustness and can adapt to a certain degree of noise points and unsteady water quality data. There are no specific requirements for the types of water quality indicators detected and the number of water quality monitoring stations, which improves the scope of application of the present invention.
[0036] (3) Most of the existing methods for identifying water pollution accidents in pipe networks require actual pollution data for model training or setting relevant parameters, and it is difficult to collect a large number of pollution events for training and learning in real life. Compared with these methods, the water quality anomaly accident detection method proposed by the present invention is an unsupervised learning method. The construction and training of the model only require water quality data under the normal operation of the water supply pipe network, and the application scope of the model is wider and the practicability is stronger. Description of the Drawings
[0037] Figure 1 It is a flowchart for model construction;
[0038] Figure 2 It is a layout diagram of a certain city's water supply pipe network and sensor monitoring stations;
[0039] Figure 3 It is a schematic diagram of the structure of the adversarial learning model (GAN);
[0040] Figure 4 It is the final water quality anomaly event identification and early warning situation obtained by using the single-site model, multi-site model, and the combined model of single-site and multi-site for the test set data respectively. Detailed Embodiment
[0041] In order to present the technical solution and advantages of the present invention more clearly, the present invention will be described in detail below with reference to the drawings and embodiments. It should be noted that the embodiments are only specific interpretations of the present invention, but the implementation manners of the invention are not limited thereto.
[0042] Embodiment 1.
[0043] Refer to the attached Figure 1 , the specific implementation steps of the present invention are as follows:
[0044] S1. Preparation and processing of data. The model constructs and trains by using the water quality data of normal sensor monitoring stations, and evaluates the performance of the model by using the data containing water quality pollution events. The water quality data of multiple sensor monitoring stations in the water supply pipe network are simulated by using the pipe network hydraulic water quality model. The selected N sensors are divided into N + 1 groups. The first N groups contain the water quality parameter data of their respective single sites, and the last group contains the water quality parameter data of N sites. The required data includes two types: normal operation water quality data and water quality data with pollution events, which are divided into the following two parts:
[0045] S11, Normal water quality data. Input the time series values of multiple water quality indicators under normal operating conditions at the water source. The water quality indicators include, but are not limited to, residual chlorine, pH, temperature, conductivity, turbidity, TOC (total organic carbon), etc. By running the hydraulic water quality model of the pipeline network, collect the water quality data of the selected multiple monitoring stations. Divide the original normal data into two parts, a 70% training set and a 30% test set.
[0046] S12, Water quality data containing pollution events. To ensure that pollution events can affect the selected water quality monitoring stations, set pollution events near the water source. Since there are few records of water quality abnormal events during the operation of the pipeline network and the occurrence of water quality events depends greatly on the environment of the pipeline network, in this invention, the method of simulating event occurrence in relevant research is referred to. By simulating the change of water quality parameters with a Gaussian shape distribution, set the duration of pollution occurrence to 10 hours, randomly simulate the number of water quality indicators affected when different pollution times occur (3 - 6), the change trend (increase or decrease) of the corresponding monitored water quality indicators, and the change amplitude (1.0 - 2.5). Add pollution to the data in the test set and use the hydraulic water quality model to obtain the change of water quality parameters at each station.
[0047] S2, Superimposed image conversion. Perform data image conversion on the normal and pollution - containing water quality data of sensor stations 1, 2, and 3. To ensure the consistency of the size of the superimposed images of single - site and multi - site, after the operation of converting time - series data into superimposed images, use the image filling method to set the final superimposed images to the same size (32 * 32), and fill the outermost circle of the image with 0. To reduce the influence of noise, take the average value of the past 5 time lengths for the superimposed image at each moment.
[0048] S3, Construct an adversarial learning model and calculate the anomaly score at each moment using the trained adversarial learning model. For the superimposed images of water quality data conversion of N single - sites and one multi - site, construct adversarial learning models (GANs) for training respectively. The structure and parameter settings of each adversarial learning model are the same. The generator adopts an auto - encoding structure and a convolutional neural network, uses the historical 30 images (2.5 hours) as input, and the output is the estimated superimposed image at the current moment. The discriminator D adopts a convolutional neural network structure, takes the historical 30 images as conditional input, and at the same time inputs the superimposed image constructed from the real data at the current moment and the superimposed image generated by the generator, and the output is the feature vector for evaluating the authenticity of the input picture. Calculate the anomaly score of water quality at each moment by integrating the reconstruction loss of the generator and the feature loss formula (4) of the discriminator.
[0049] S4. For the anomaly scores calculated by the adversarial learning models trained in S3, it is necessary to determine the normal value threshold ranges of the anomaly scores calculated by different models. For single-site and multi-site models, the selection of their thresholds is related to the distribution of the anomaly scores calculated in the training set. The general principle is to keep the anomaly scores calculated in most of the training set within the normal range.
[0050] S5. Update the probability of water quality anomaly events and give early warnings. Use the principle of sequential Bayesian to update the probability of water quality anomaly events. When the probability exceeds a certain threshold, the model identifies it as a pollution accident alarm. The probability threshold for pollution event alarm is set to P thre = 80%, and the probability of pollution events at the initial moment is set to a relatively low value P(0) = 10 -5 , and the smoothing parameter takes the value of α = 0.6. Use formulas (5) and (6) to update the probability of water quality anomaly events for N single-site models and one multi-site model respectively.
[0051] S6. Statistically analyze the alarm results of multiple models at the same moment. Calculate the probability of pollution events of the final model based on the probabilities of pollution events obtained from the multi-site model and the single-site model. all When it exceeds the probability threshold P thre = 80% of the pollution event alarm, give an alarm for the final model, and the probability of pollution events P all as well as the sensor sites where pollution is detected and the corresponding water quality parameters can be output.
[0052] Apply the method of the present invention to the water supply network of a certain city in China ( Figure 2 ). A total of 33 sensor detection sites are placed in this water supply network. Select sensor sites 1, 2, and 3 close to the pollution for the construction and performance evaluation of the model method. Collect the water quality data records of the three sites at 5-minute intervals for 2 months. The designed water quality includes residual chlorine, pH, temperature, conductivity, turbidity, and total organic carbon. Divide the data into a training set and a test set. The data in the training set are all the data of the normal operation of the water supply network and are used for the construction and training of the model. The data in the test set are the data with pollution events added and are used for the performance evaluation of the model to detect pollution accidents. Figure 4 Compare the final detection results of the model with the single-site model based on a single site and the multi-site model. It can be seen that the proposed model combines the advantages of the single-site and multi-site models. Among 7 pollution accidents, the model successfully detected 6 pollution events and only had 2 short-term false alarms. Its results are better than those of the detection results using only the single-site or multi-site model. The method proposed by the present invention can improve the detection accuracy and increase the detection rate of pollution events. The application of this example also confirms that the method proposed by the present invention has good practicability, a high effective alarm rate, a low false alarm rate, and has good application value in the actual water supply network.
[0053] The above-described embodiments merely represent the implementation modes of the present invention, but should not be construed as limiting the scope of the patent for the present invention. It should be noted that for those skilled in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all fall within the protection scope of the present invention.
Claims
1. A method for identifying and warning water quality abnormal events based on spatio-temporal data of pipe network water quality, characterized in that, it includes the following steps: (1) Select the monitored water quality parameters and water quality sensor sites. By analyzing the time interval and change trend of the sensors receiving polluted water, group and analyze adjacent monitoring sites in the water supply pipe network with the same change trend; divide N sensor detection sites into N + 1 groups: in the first N groups, each group contains multi-parameter water quality monitoring data of a single sensor site, and there is another group that contains multi-parameter water quality monitoring data of all N sensor sites; (2) Standardize and preprocess the water quality data of N + 1 groups, and convert the water quality time series data within each group after preprocessing into superimposed images respectively; assume that each site collects Nr water quality parameters, and for each moment, V water quality parameter data can be collected from the N sensor sites under analysis (V = Nr × N); at time t, the distance image of each group can be expressed as where is the element in the i-th row and j-th column of the superimposed image m t at time t, calculated as the sum of the i-th and j-th data among the V water quality parameter data collected at time t; when N = 1, the superimposed image conversion is only used for the multi-parameter data fusion of a single sensor site; when N > 1, its superimposed image conversion is used for fusing the spatio-temporal distribution data of multiple sites; in order to reduce the influence of noise, the superimposed image at each moment takes the average value of the past d time lengths. (3) Construct an adversarial learning network model GAN; The described adversarial learning network model includes a generator G and a discriminator D. The generator G uses a convolutional neural network including an encoder for compressing image information and a decoder for restoring image information. The discriminator D is a classification convolutional neural network used to compress and extract image features. First, use step (2) to convert the multivariate time series data of the water quality monitoring sites analyzed under the normal operating state into a superimposed image. Then, use the historical superimposed images of the past K moments as the input of the GAN model, and generate the superimposed image of the next moment through the generator network G. Finally, input the image generated by the generator and the superimposed image m based on the measured data into the discriminator network. The discriminator is used to distinguish the superimposed image m (real) of the measured data and the superimposed image (forged) reconstructed by the generator G; The training of the GAN model uses an improved W-GAN loss as the training loss L of the discriminator D , which is used to stabilize the training process and is specifically expressed as Where E represents the calculated mathematical expectation; D(·) is the feature vector finally output by the discriminator D; is the superimposed image reconstructed by the generator; m is the superimposed image based on the measured data; is the interpolated image composed of a pair of real superimposed images and reconstructed superimposed images; is the interpolated image to calculate the gradient; λ GP is the coefficient of the gradient penalty term, which is taken as 10 according to the empirical value; The generator training uses the reconstruction loss between the real superimposed image and the generated superimposed image as the generator loss L G , to help the generator learn the distribution of normal water quality data; Its calculation representation is, In the GAN model, train and update both the G and D networks simultaneously; the ultimate goal of model training is to find a stable state when the losses of G and D both converge to a stable state; (4) Based on the generator and discriminator of the GAN model trained in step (3), construct an anomaly score ψ; (5) Select an appropriate anomaly score threshold ψ thre , such that most points in the water quality dataset used to construct the GAN model are within the normal range, and when the anomaly score obtained in step (4) exceeds ψ thre , it is used as the initial anomaly point identification; (6) Use the principle of temporal Bayesian to calculate the probability of water quality abnormal events, and calculate the probability P(t) of pollution events occurring at each moment based on the recognition results of the abnormal points obtained in step (5). Specifically, it can be expressed by the following formula, In the formula, TPR is the true positive rate, calculated as the ratio of the number of time steps correctly classified as abnormal when the water supply pipe network is polluted to the total number of time steps of pollution; when there is no prior knowledge of pollution events, the value of TPR is 0.5; FPR is the false positive rate, calculated as the ratio of the number of time steps identified as abnormal points by the model when the water supply pipe network is operating normally without pollution to the total number of time steps of normal operation, which is equivalent to the number of moments exceeding the anomaly score threshold in the training set to the total number of the training data set; The probability of a pollution event occurring at the initial moment is given as P(0). When the calculated probability P(t) exceeds the threshold P thre , the model identifies it as a pollution event; the probability threshold P thre for identifying pollution events is set according to the decision maker's consideration of the trade-off between the detection rate and false alarm rate of pollution events. (7) The normal operation hydraulic changes in the water supply pipe network will cause sudden changes in water quality parameters in the short term; in order to distinguish normal water quality changes and pollution events, use an exponential smoothing model to smooth the calculated probability; (8) Calculate the probability of water quality abnormal accidents using single-site models and multi-site models respectively; The single-site model means that the anomaly detection methods proposed in steps (3)-(7) are separately applied to the first N groups of water quality data in steps (1)-(2), a GAN model is constructed for each site, and the maximum value of the pollution event occurrence probabilities calculated by the N GAN models is used as the event probability P calculated by the single-site model single ; The multi-site model means applying the anomaly detection method proposed in steps (3)-(7) to the (N + 1)-th group of water quality data in steps (1)-(2). This group of data includes water quality data monitored by N water quality monitoring stations, constructing a GAN model, and obtaining the probability P of pollution events calculated by the final multi-site model multi ; (9)In order to make full use of the water quality relationships between and within monitoring stations, the event probabilities calculated by the single-station model and the multi-station model are fused, and based on the multivariate water quality parameters of all monitoring stations, the combined event probability P reflecting the likelihood of pollution events is obtained all , which is expressed by the following expression: P all P(t) = ηP single P(t) + (1 - η)P multi P(t) (7) Wherein, P single (t) and P multi (t) respectively represent the occurrence probabilities of pollution events calculated using the single-site model and the multi-site model at time t; η is an important weight that adjusts the influence of single-site models and multi-site models on the synchronous decision-making of the final pollution event recognition; Both the single-site and multi-site models described above are unsupervised models and do not know the pollution information in advance. Therefore, η = 0.5 is set to reflect the same importance of the single-site model and the multi-site model for the final pollution detection alarm; when the combined event probability P all exceeds the preset probability threshold P thre in step (6), an alarm signal of the final model is given, and the probability P all of the occurrence of the water quality abnormal accident is given.
2. A method for identifying and warning water quality abnormal events based on spatio-temporal data of pipe network water quality according to claim 1, characterized in that, In the step (4) described above, the anomaly score calculation based on GAN includes two parts: anomaly recognition using the generator network and anomaly recognition using the discriminator network; among them, the anomaly recognition based on the generator network is calculated by comparing the image generated by the generator network with the reconstruction loss between the superimposed image m constructed based on the measured data; the anomaly recognition based on the discriminator network is by comparing the feature vectors D(m) and obtained after respectively inputting the superimposed image based on the measured data and the image generated by the generator into the discriminator network; the anomaly score finally constructed based on the GAN model is specifically expressed as Among them, ψ(t) represents the anomaly score calculated by the GAN model at time t; the closer the anomaly score is to 0, the more normal the current water quality state is. On the contrary, the larger the anomaly score is, the greater the difference between the current water quality state and the normal operation state is, and the more likely the water quality is in an abnormal state at this moment; λ s is a weighted parameter that adjusts the relative importance of the reconstruction loss and the feature loss with respect to the anomaly score.
3. A method for identifying and warning water quality abnormal events based on spatio-temporal data of pipe network water quality according to claim 1, characterized in that, In the said step (5), the abnormal score threshold ψ thre is confirmed in the following manner: Since both the construction and training of the GAN model in step (3) are based on the water quality data under normal operating conditions, the selection of the abnormal score threshold is mainly determined by the abnormal score distribution obtained from the dataset for training the GAN network; after obtaining the abnormal scores for the water quality dataset used in the construction and training of the GAN model in step (3) by using step (4), the abnormal score values are arranged from small to large, and according to the actual situation of the data distribution, the 96%-99% quantile point in its distribution is selected as its abnormal score threshold ψ thre , so that most of the points in the water quality dataset used to construct the GAN model are within the normal range.
4. A method for identifying and warning water quality abnormal events based on spatio-temporal data of pipe network water quality according to claim 1, characterized in that, In the said step (6), the probability of the pollution event occurring is P(0) ∈ [10 -6 , 10 -4 ; the probability threshold P thre for the said pollution event recognition is set to a probability value exceeding 70%.
5. A method for identifying and warning water quality abnormal events based on spatio-temporal data of pipe network water quality according to claim 1, characterized in that, In the said step (7), the following formula is used for smoothing processing: P(t) = αP(t) + (1 - α)P(t - 1) (6) In the formula, α is the smoothing parameter, which determines the degree of emphasis on the probability of the most recently updated event, α ∈ [0.3, 0.9]; the smaller α is, the slower the event probability is updated, and the more abnormal points are required to reach the alarm threshold.
Citation Information
Patent Citations
Detection method of abnormal event of multi-variable water quality parameter time sequence data
CN106872657A
Water quality abnormal event identification and early warning method based on pipe network multivariate water quality time series data
CN111191855A