Pipe burst detection and performance evaluation method based on single point and time series anomaly detection
The integration of single-point and time-series anomaly detection using unsupervised stacking integration and statistical theory enhances pipe burst detection accuracy in water supply systems, reducing false alarms and ensuring timely response to pipe ruptures.
Patent Information
- Application Number
- JP2023572727
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-08-21
- Filing Date
- 2023-10-20
- Publication Date
- 2025-09-29
- Estimated Expiration
- 2043-10-20
AI Technical Summary
Existing data-driven methods for pipe burst detection in water supply systems suffer from high false alarm rates due to limited monitoring data and interference from noisy data, making it difficult to accurately identify pipe bursts and distinguish between normal and abnormal conditions.
A method combining single-point and time-series anomaly detection using unsupervised stacking integration and statistical theory to analyze real-time monitoring data, integrating results to accurately detect pipe bursts and distinguish between normal and abnormal conditions.
The method effectively reduces false alarms and improves detection accuracy by identifying pipe bursts and fault conditions in water supply networks, ensuring timely response to pipe ruptures and maintaining water quality.
Smart Images

Figure 2025531958000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to the field of urban water supply pipe burst detection technology, and in particular to a pipe burst detection and performance evaluation method based on single point and time series anomaly detection. [Background technology]
[0002] Pipe bursts are a major form of water loss in water supply systems, resulting in significant water losses despite their short duration. Pipe bursts not only waste water resources, but also cause a drop in pressure in the pipe network, affecting normal water supply. After a pipe bursts, contaminants are likely to enter the system, potentially affecting the quality of drinking water. Pipe burst detection is an important measure for ensuring urban water safety, aiming to respond to pipe burst incidents in a timely manner and reduce the harm caused by pipe burst incidents.
[0003] Data-driven methods are a new approach in the field of pipe burst detection, addressing pipe burst detection as an anomaly detection problem and identifying pipe bursts by detecting abnormal features in real-time monitoring data. This approach does not require the use of hardware equipment or a pipe network model to detect pipe bursts; instead, it simply uses software to perform anomaly detection on the pipe network monitoring data, offering the advantages of low cost, short processing time, and low labor intensity. At the same time, due to unfavorable factors such as limited pipe burst monitoring data and interference from dirty data, data-driven methods at this stage have a high false alarm rate for pipe burst detection.
[0004] Data-driven methods can be further divided into methods based on single-point anomaly detection and methods based on time-series anomaly detection. Single-point anomaly detection methods detect anomalies in a single value of monitored data, and if anomalies are detected over multiple time periods (time windows), a pipe burst warning is issued. These methods can timely detect abnormal values in monitored data and distinguish between normal and abnormal working conditions in a pipe network. They are widely used in the field of pipe burst detection. However, these methods cannot accurately identify the various triggering factors of anomalies. Therefore, appropriate data preprocessing methods must be used to filter out abnormal values from the monitored data before pipe burst detection. Under normal working conditions, the monitored data of a pipe network follows a specific cyclical pattern, and all monitored data are normal. When a pipe burst occurs in a pipe network, dirty data may appear at some or all monitored points, changing the original cyclical pattern of the monitored data and resulting in the occurrence of abnormal values in the monitored data. If the triggering factors of these abnormal values cannot be accurately identified after detecting these abnormal values, the false alarm rate for pipe burst detection will increase, and if no warning is given after detecting these abnormal values, the detection rate for pipe burst will decrease.
[0005] Methods based on time series anomaly detection consider the correlation between monitoring data and find anomalous time series by analyzing the dissimilarity changes of the time series, and methods based on time series anomaly detection can detect these changes. Such methods eliminate the interference of unstable conditions on the results of pipe burst detection. However, such methods are susceptible to noisy data, and a small amount of noise can dominate the similarity measure, increasing the false alarm rate of such methods. Summary of the Invention [Problem to be solved by the invention]
[0006] SUMMARY OF THE INVENTION The present invention aims to provide a method for detecting and evaluating the performance of a pipe burst based on single-point and time-series anomaly detection, so as to make up for the above-mentioned shortcomings and solve the problems raised in the background art. [Means for solving the problem]
[0007] The present invention uses the following technical solutions to solve the above technical problems: A water supply pipe burst detection and performance evaluation method based on single-point and time-series anomaly detection, which includes: (1) performing single-point anomaly detection on real-time monitoring data of a water supply pipe network to distinguish between normal and abnormal working conditions of the water supply pipe network; (2) performing time-series anomaly detection on the real-time monitoring data of the water supply pipe network to distinguish between normal and abnormal working conditions of the water supply pipe network; and (3) integrating the results of the single-point and time-series anomaly detection on the water supply pipe network to accurately detect and distinguish between water supply pipe network bursts and fault conditions in the monitoring system.
[0008] Preferably, step (1) specifically includes: a step (1.1) of preparing historical data and real-time single-point monitoring data for each monitoring point in the water supply network, so as to perform single-point anomaly detection on the real-time monitoring data; a step (1.2) of using an unsupervised stacking integration algorithm to perform single-point anomaly detection on the real-time monitoring data of the water supply network, and identifying single-point anomalies in the real-time monitoring data; and a step (1.3) of using statistical theory to qualitatively detect single values in the monitoring data of the water supply network, and identifying single-point anomalies in the real-time monitoring data.
[0009] Preferably, step (2) specifically includes: a step (2.1) of preparing the history and time series of real-time monitoring data of each monitoring point in the water supply network so as to detect time series abnormalities in the real-time monitoring data; a step (2.2) of analyzing the differences between the time series of monitoring data of different monitoring points to determine abnormal time series between the monitoring points; and a step (2.3) of analyzing the differences between the time series of monitoring data of the same monitoring point to determine abnormal time series of the monitoring point itself.
[0010] Preferably, step (3) specifically includes the steps of: (3.1) using a single-point anomaly detection method to perform single-point anomaly detection for various abnormal situations in the water supply network; (3.2) using a time-series anomaly detection method to perform time-series anomaly detection for various abnormal situations in the water supply network; (3.3) integrating the results of the single-point and time-series anomaly detection to detect and identify various abnormal situations in the water supply network; and (3.4) using a pipe burst detection method to detect and identify various abnormal situations in the water supply network and evaluating the pipe burst detection performance of the provided method.
[0011] Preferably, step (3.1) specifically includes the following: When using the single-point abnormal value detection method to perform single-point abnormal value detection for various abnormal situations in a water supply network, the following four situations are considered: (1) a pipe burst occurs in the pipe network; (2) dirty data appears at a single monitoring point; (3) dirty data appears at some monitoring points; and (4) dirty data appears at all monitoring points. Each situation includes 25 sets of abnormal monitoring data, and single-point abnormal value detection is performed from the time the abnormal value begins, and whether to issue an alert is determined based on the abnormal value detection results at five consecutive times.
[0012] Preferably, step (3.2) specifically includes the following: When using the time series anomaly detection method to perform time series anomaly detection for various abnormal situations in a water supply pipe network, the following four situations are considered: (1) a pipe burst occurs in the pipe network; (2) dirty data appears at a single monitoring point; (3) dirty data appears at some monitoring points; and (4) dirty data appears at all monitoring points. Each situation includes 25 sets of abnormal monitoring data, and time series anomaly detection is performed between monitoring points starting from the time when the abnormal value begins, and whether to issue an alert is determined based on the time series anomaly detection results at five consecutive times.
[0013] Preferably, step (3.3) specifically includes the steps of: detecting whether there is a single point or time series anomaly in the monitoring data; identifying a situation in which anomalies appear in the monitoring data of some monitoring points; identifying a situation in which pipe bursts occur and dirty data appears at all monitoring points; and evaluating and analyzing the performance of the method for detecting pipe bursts (3.3.4).
[0014] Preferably, step (3.3.1) is specifically as follows: First, distinguish between the normal working conditions of the piping network and situations in which abnormalities appear in the monitoring data. If a single-point abnormal value is detected in the single-point monitoring data based on an unsupervised stacking integration algorithm or statistical theory, mark the time when the single-point abnormal value appears, and continue single-point abnormal value detection for the monitoring data at the next time. If an abnormal time series is found in the time-series monitoring data based on time-series anomaly detection between monitoring points or time-series anomaly detection within the monitoring point, mark the time when the abnormal time series appears, and continue time-series anomaly detection for the monitoring data at the next time. Step (3.3.2) is specifically as follows: If abnormalities appear in the single points and time series of the monitoring data of some monitoring points, but no abnormal situation appears in the single points and time series of the monitoring data of the remaining monitoring points, it is determined that dirty data appears at some monitoring points. Step (3.3.3) is specifically as follows: If anomalies are detected in both the single-point and time-series monitoring data of all monitoring points, the increase or decrease in the anomaly values will distinguish between a situation in which a pipe burst will occur in the pipe network and a situation in which dirty data appears at all monitoring points. If the single-point qualitative detection results of the monitoring data of all monitoring points are all "decreasing," it will be determined that a pipe burst will occur in the pipe network. Conversely, if the single-point qualitative detection results of the monitoring data are "increasing," it will be determined that dirty data has appeared at all monitoring points in the pipe network. In step (3.3.4), the performance of the method for pipe burst detection is evaluated using the following indicators: (1) detection accuracy rate (σ1); (2) anomaly identification rate (σ2); and (3) anomaly detection rate (σ3).
[0015] Preferably, the detection accuracy rate is given as follows: σ1=N dn / N n In the formula, N n denotes the total number of abnormal scenes, and N dn indicates the number of abnormal scenes detected.
[0016] The anomaly identification rate is shown as follows: σ2=N in / N n In the formula, N in denotes the number of correctly identified anomalous scenes, and N n indicates the total number of abnormal scenes.
[0017] The anomaly detection rate is shown as follows: σ3=t dn / t t In the formula, t dn and t t and denote the duration due to abnormal event detection and the actual duration, respectively.
[0018] Preferably, step (3.4) specifically includes: a step (3.4.1) of removing single-point abnormal value detection results when detecting and identifying various abnormal situations in the water supply network, and obtaining the abnormality detection rate, detection accuracy rate, and abnormality identification rate for various abnormal situations; a step (3.4.2) of removing single-point value qualitative detection results when detecting and identifying various abnormal situations in the water supply network, and obtaining the abnormality detection rate, detection accuracy rate, and abnormality identification rate for various abnormal situations; a step (3.4.3) of removing time-series abnormality detection results between monitoring points when detecting and identifying various abnormal situations in the water supply network, and obtaining the abnormality detection rate, detection accuracy rate, and abnormality identification rate for various abnormal situations; and a step (3.4.4) of removing time-series abnormality detection results of the monitoring point itself when detecting and identifying various abnormal situations in the water supply network, and obtaining the abnormality detection rate, detection accuracy rate, and abnormality identification rate for various abnormal situations. [Effects of the Invention]
[0019] The beneficial effects of the present invention are as follows: The method of the present invention utilizes the history and real-time monitoring data of a water supply network to effectively detect and distinguish pipe ruptures and fault conditions at monitoring points in the water supply network, and performs single-point and time-series anomaly detection on the real-time monitoring data of the water supply network to timely discover abnormal working conditions in the water supply network, and can accurately distinguish between various abnormal situations according to the various warning conditions of the single-point and time-series anomaly detection results. That is, the method of the present invention can accurately identify pipe ruptures and distinguish between abnormal conditions in various monitoring point data. [Brief explanation of the drawings]
[0020] [Figure 1] 1 is a flowchart of a water network pipe burst detection and performance assessment method based on single point and time series anomaly detection. [Figure 2] FIG. 1 is a schematic diagram of single point value anomaly detection provided by the present invention. [Figure 3] 1 is a schematic diagram of single-point outlier detection based on unsupervised stacking integration provided by the present invention: (ac) monitoring data time series; (de) comparison of real-time data with historical data; (fg) single-point outlier detection. [Figure 4] 1 is a schematic diagram of single-point outlier detection based on statistical theory provided by the present invention: (ac) monitoring data time series; (de) comparison of real-time data with historical data; (fg) single-point outlier detection. [Figure 5] 1 is a schematic diagram of time series anomaly detection between monitoring points provided by the present invention: (a) monitoring data time series; (bc) comparison of real-time monitoring data time series with historical data time series; (d) monitoring data time series anomaly detection. [Figure 6] 1 is a flowchart of time series anomaly detection between monitoring points provided by the present invention; [Figure 7]1 is a schematic diagram of the time series anomaly detection of the monitoring point itself provided by the present invention: (a) monitoring data time series; (bc) comparison of real-time monitoring data time series with historical monitoring data time series; (c) monitoring data time series anomaly detection. [Figure 8] 1 is a flowchart of a time series anomaly detection process for a monitoring point itself provided by the present invention; [Figure 9] 1 is a schematic diagram of single-point abnormal value detection after a pipe burst occurs in a pipeline network in an embodiment of the present invention: (a) monitoring data time series curve; (b) monitoring data single-point value; [Figure 10] 1 shows the single-point anomaly detection results after dirty data appears at some monitoring points in an embodiment of the present invention: (a) monitoring data time series curve; (b) monitoring data single-point value. [Figure 11] 1 shows the single-point anomaly detection results after dirty data appears at all monitoring points in an embodiment of the present invention: (a) monitoring data time series curve; (b) monitoring data single-point value. [Figure 12] 1 shows time series anomaly detection results after a pipe rupture occurs in a piping network in an embodiment of the present invention: (a) monitoring data time series curve; (b) time series between monitoring points; (cd) time series of the monitoring point itself. [Figure 13] 10 shows time series anomaly detection results after dirty data appears at some monitoring points in an embodiment of the present invention: (a) monitoring data time series curve; (b) time series between monitoring points; (cd) time series of the monitoring points themselves. [Figure 14] 1 shows time series anomaly detection results after dirty data appears at all monitoring points in an embodiment of the present invention: (a) monitoring data time series curve; (b) time series between monitoring points; (cd) time series of the monitoring point itself. [Figure 15] FIG. 1 is a schematic diagram of pressure monitoring point distribution in a Net 3 piping network model according to an embodiment of the present invention. [Figure 16] 1 shows the experimental conditions and the process of data preparation techniques in an embodiment of the present invention. [Figure 17] FIG. 10 is a schematic diagram of a monitoring data .csv file in an embodiment of the present invention. [Figure 18]The results of single-point outlier detection based on unsupervised stacking integration in an embodiment of the present invention are: (a) a pipe burst occurs in the pipe network; (b) dirty data appears at a single monitoring point; (c) dirty data appears at some monitoring points; and (d) dirty data appears at all monitoring points. [Figure 19] The single-point abnormal value detection results based on statistical theory in an embodiment of the present invention are: (a) a pipe burst occurs in a pipe network; (b) dirty data appears at a single monitoring point; (c) dirty data appears at some monitoring points; (d) dirty data appears at all monitoring points. [Figure 20] The results of time-series anomaly detection between monitoring points in an embodiment of the present invention are: (a) a pipe burst occurs in the piping network; (b) dirty data appears at a single monitoring point; (c) dirty data appears at some monitoring points; and (d) dirty data appears at all monitoring points. [Figure 21] The results of time-series anomaly detection at the monitoring point itself in an embodiment of the present invention are: (a) a pipe burst occurs in the piping network; (b) dirty data appears at a single monitoring point; (c) dirty data appears at some monitoring points; and (d) dirty data appears at all monitoring points. [Figure 22] 1 shows the anomaly detection rate, detection accuracy rate, and anomaly identification rate for various abnormal scenes in an embodiment of the present invention. [Figure 23] FIG. 10 illustrates the removal of single-point anomaly detection results in an embodiment of the present invention. [Figure 24] FIG. 10 illustrates the removal of single-point value qualitative detection results in an embodiment of the present invention. [Figure 25] FIG. 10 is a diagram illustrating the removal of time-series abnormality detection results between monitoring points in an embodiment of the present invention. [Figure 26] FIG. 10 is a diagram illustrating the removal of time-series abnormality detection results of a monitoring point itself in an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0021] The present invention will now be described in more detail with reference to the accompanying drawings and specific examples.
[0022] Example 1 As shown in Figure 1, the method for detecting and evaluating the performance of a water supply pipe burst based on single-point and time-series anomaly detection includes the steps of: (1) performing single-point anomaly detection on real-time monitoring data of a water supply pipe network to distinguish between normal and abnormal working conditions of the water supply pipe network; (2) performing time-series anomaly detection on real-time monitoring data of the water supply pipe network to distinguish between normal and abnormal working conditions of the water supply pipe network; and (3) integrating the results of the single-point and time-series anomaly detection of the water supply pipe network to accurately detect and distinguish between water supply pipe network bursts and fault conditions in the monitoring system.
[0023] Furthermore, the step (1) specifically includes the steps of: (1.1) preparing history and real-time single-point monitoring data for each monitoring point in the water supply network so as to perform single-point abnormal value detection on the real-time monitoring data; (1.2) using an unsupervised stacking integration algorithm to perform single-point abnormal value detection on the real-time monitoring data of the water supply network and identify single-point abnormal values in the real-time monitoring data; and (1.3) using statistical theory to qualitatively detect single values in the monitoring data of the water supply network and identify single-point abnormalities in the real-time monitoring data.
[0024] Furthermore, the step (1.1) specifically includes the following: x denotes the monitoring data of the monitoring point, the monitoring data of the monitoring point 1 is denoted by x1, and the real-time monitoring data is denoted by x1. 0 and historical monitoring data is indicated by x1 i (i=1,2,...,n d ) and the monitoring data of monitoring point 1 at time m is denoted by x1(m). Single point anomaly detection is performed by comparing the real-time monitoring data of the monitoring point with the past n d For example, if it is necessary to determine whether the real-time monitoring data at time m of monitoring point 1 is an abnormal value, the real-time monitoring data x1 0 (m) and past n d1x historical monitoring data for m days i (m)(i=1,2,...,n d ) and various anomaly detection methods are used to compare x1 0 Determine whether (m) is an outlier.
[0025] Before performing single-point outlier detection on the monitoring data x of each monitoring point, the real-time monitoring data and historical monitoring data of each monitoring point need to be prepared, as shown below:
[0026]
number
[0027] In the formula, x 0 (i) (i=1,2,···,m) denotes the real-time monitoring data of the monitoring point, and x 0 (m) denotes the mth real-time monitoring data, and m is the total number of daily monitoring data at each monitoring point. x j (i) (i=1,2,···,m; j=1,2,···,d) indicates the historical monitoring data of the monitoring point, and x j (m) is the mth historical monitoring data for the past j days at the monitoring point, and d is the total number of days of historical monitoring data. Each row of the matrix X represents the real-time and historical monitoring data at a certain time at the monitoring point, for example, X(i)=[x 0 (i),x 1 (i),···,x j (i),···,x d (i)] indicates the real-time and historical monitoring data of the i-th monitoring point. Each row of the matrix X indicates the monitoring data of the monitoring point at each time of the day, for example, X(j)=[x j (1), x j (2),···,x j (i),···,x j (m)] T indicates the historical monitoring data for each time point for the past j days at the monitoring point.
[0028] After collecting real-time monitoring data from each monitoring point, it is compared with the historical monitoring data for the corresponding time in the past d days to check whether the real-time monitoring data is an abnormal value. When performing abnormal value detection for the m-th real-time monitoring data, prepare the m-th real-time monitoring data for all monitoring points and the m-th historical monitoring data for the past d days.
number
[0029] In the formula, x k 0 (m) (k=1, 2, . . . , n) denotes the m-th real-time monitoring data of monitoring point k, and n is the number of monitoring points distributed in the piping network. k j (m)(k=1,2,···,n;j=1,2,···,d) represents the mth historical monitoring data for the past j days at monitoring point k. Each row of the matrix X(m) represents the mth historical and real-time monitoring data for a certain monitoring point. For example, x k (m)=[x k 0 (m),x k 1 (m),···,x k j (m),···,x k d (m)](j=1,2,...,d) indicates the m-th historical and real-time monitoring data of monitoring point k. Each row of the matrix X(m) indicates the m-th monitoring data of a day at all monitoring points in the pipeline network. For example, x j (m)=[x1 j (m), x2 j (m),···,x k j (m),···,x n j (m)] T (k=1,2,···,n) indicates the mth historical monitoring data for the past j days at the monitoring point.
[0030] Furthermore, the above step (1.2) specifically includes the following: For single-point anomaly detection in monitoring data, the present invention employs a machine learning algorithm based on unsupervised stacking integration. This algorithm is primarily used in power system fault diagnosis, and can effectively detect and accurately identify power system fault events and abnormalities in monitoring data by utilizing real-time monitoring data from SCADA systems. The algorithm is divided into three layers: (1) Layer 1 is an isolation forest algorithm; (2) Layer 2 is a k-means clustering algorithm and a local outlier factor algorithm; and (3) Layer 3 is an integration of the k-means clustering algorithm and the local outlier factor algorithm.
[0031] When performing single-point anomaly detection on monitored data, each row of X(m) is input to the unsupervised stacking integration algorithm as the data to be detected. For the monitored data of n monitoring points, single-point anomaly detection must be performed n times at each time. For example, taking the monitored data of monitoring point 1 as an example, x1(m)=[x1 0 (m), x1 1 (m),···,x1 d A total of d+1 data items (m) to be detected are input to the unsupervised stacking integration algorithm to obtain the probability that each monitored data item is an anomaly. 0If (m) has the highest probability of being an anomaly, mark the real-time monitoring data as an anomaly. Otherwise, it is determined that the real-time monitoring data is a normal value and should not be marked. Step (1.2) mainly includes the following steps: inputting the data to be detected into the separation forest algorithm to obtain an anomaly score for each data (1.2.1); using each anomaly score obtained by the separation forest algorithm in layer 1 as input to the k-means clustering algorithm in layer 2, and clustering each anomaly score in the k-means clustering algorithm to obtain binary data 0 (normal) and 1 (abnormal); using each anomaly score obtained by the separation forest algorithm in layer 1 as input to the local outlier factor algorithm in layer 2, and obtaining an outlier factor for each anomaly score using the local outlier factor algorithm (1.2.3); and obtaining the probability that each anomaly score will be an anomaly based on the output results of the k-means clustering algorithm and the local outlier factor algorithm, i.e., the probability that each data to be detected will be an anomaly (1.2.4).
[0032] Furthermore, the above step (1.3) specifically includes the following: The present invention provides a single-point anomaly detection method based on statistical theory, so as to avoid identifying the occurrence of dirty data at a monitoring point as a pipe network burst. Because a pipe burst incident usually leads to a rapid decrease in monitoring data, the abnormal monitoring value after a pipe burst occurs should be smaller than the normal monitoring value. If the abnormal monitoring value is larger than the normal monitoring value, it can be determined that no pipe burst accident has occurred in the pipe network and dirty data has appeared at the monitoring point. In addition, to avoid identifying normal monitoring values due to fluctuations in water demand as abnormal monitoring values, real-time monitoring data should only be marked as an abnormal value after exceeding a certain range.
[0033] For monitoring data x of a single monitoring point, historical monitoring data x i (i=1,2,···,d) determines the qualitative threshold [ζ - (x),ζ +(x)]. ζ - (x)=μ(x i )-m k σ(x i ) ζ + (x)=μ(x i )+m k σ(x i )
[0034] x 0 ζ - (x) and ζ + Compared to (x), there are three situations: (1) x 0 <ζ - In the case of (x), the real-time monitoring data is an abnormal value and shows a "decreasing" trend compared to the historical monitoring data, and the abnormality detection result is indicated by -1; (2) ζ - (x)≦x 0 ≦ζ + In the case of (x), the real-time monitoring data indicates a normal value, and the abnormality detection result is shown as 0; (3) ζ + (x) <x 0 In this case, the real-time monitoring data is an abnormal value and has an "increasing" tendency compared to the historical monitoring data, and the abnormality detection result is indicated as 1.
[0035] Furthermore, step (2) specifically includes step (2.1) of preparing the history and time series of real-time monitoring data of each monitoring point in the water supply pipe network so as to detect time series abnormalities in the real-time monitoring data, step (2.2) of analyzing the differences between the time series of monitoring data of different monitoring points to determine abnormal time series between the monitoring points, and step (2.3) of analyzing the differences between the time series of monitoring data of the same monitoring point to determine abnormal time series of the monitoring point itself.
[0036] Furthermore, the step (2.1) specifically includes the following: The monitoring data time series refers to a vector including multiple monitoring values, for example, the monitoring data time series of monitoring point 1 is S1=[x1 m-l+1 ,x1 m-l+2 ,···,x1 m ] can be shown as x1m indicates the monitoring data at time m at monitoring point 1, l is the length of the monitoring data time series, i.e., the monitoring data time series S1 includes l monitoring data, and x1 m-l+1 indicates the monitoring data at monitoring point 1 at time m-l+1. If all the monitoring data in S1 are normal, i.e., if the shape of the monitoring data time series S1 does not change significantly, S1 is considered to be a normal monitoring data time series. Otherwise, if an abnormality appears in the monitoring data in S1, i.e., if the shape of the monitoring data time series S1 changes (compared to the normal monitoring data time series), S1 is considered to be an abnormal monitoring data time series. To detect abnormalities in monitoring data time series, the present invention uses a method based on detecting dissimilarity between monitoring data time series. This method detects abnormal monitoring data time series using the distance between two monitoring data time series, and for two monitoring data time series S1 and S2, the distance between the monitoring data time series is expressed as follows:
number
[0037] If each monitoring data in monitoring data time series S1 and S2 is constant, d(S1,S2) is also constant, i.e., d(S1,S2) remains constant. A change in the monitoring data in monitoring data time series S1 and S2 will cause a change in d(S1,S2). Conversely, a change in d(S1,S2) indicates that the monitoring data in monitoring data time series S1 or S2 has changed, i.e., an anomaly exists in the monitoring data. Therefore, the present invention detects abnormal monitoring data time series by detecting distance changes in the monitoring data time series. Depending on the source of the monitoring data in the data time series to be detected, it can be divided into (1) time series anomaly detection between monitoring points and (2) time series anomaly detection within the monitoring point itself. Before detecting anomalies in the monitoring data time series, the monitoring data time series for each monitoring point must be prepared. For a single monitoring point 1, the mth monitoring data time series for the past j days is shown as follows: S1 j (m)=[x1 j (m-l+1),x1 j (m-l+2),···x1 j (m-1), x1 j (m)] In the formula, l indicates the length of the monitoring data time series, i.e., the monitoring data time series contains l monitoring data, the starting monitoring data of the mth monitoring data time series for monitoring point 1 is x1(m-l+1), and the last monitoring data is x1(m).
[0038] The monitoring data time series is shown below:
number
[0039] In the formula, S 0 (i) (i=1,2,···,m) indicates the real-time monitoring data time series of the monitoring point, and S 0 (m) denotes the m-th real-time monitoring data time series of a monitoring point, and m is the total number of daily monitoring data at each monitoring point. j (i) (i=1,2,···,m;i=1,2,···,d) indicates the time series of historical monitoring data of the monitoring point, and S j(m) indicates the m-th historical monitoring data time series for the past j days at the monitoring point, and d is the total number of days of historical monitoring data. Each row of the matrix S indicates the real-time and historical monitoring data time series at a certain time at the monitoring point, for example, S(i) = [S 0 (i),S 1 (i),···,S j (i),···,S d (i)] indicates the i-th real-time and historical monitoring data time series of the monitoring point. Each column of the matrix S indicates the monitoring data time series at each time of a day at the monitoring point. For example, S(j)=[S j (1), S j (2),···,S j (i),···,S j (m)] T indicates the time series of historical monitoring data at each time for the past j days at the monitoring point.
[0040] When performing anomaly detection on the mth real-time monitoring data time series, the mth real-time monitoring data time series at all monitoring points and the mth historical monitoring data time series for the past d days are prepared, as shown below.
number
[0041] In the formula, S k 0 (m) (k=1, 2, . . . , n) denotes the m-th real-time monitoring data time series of monitoring point k, and n is the number of monitoring points distributed in the pipeline network. k j (m) (k=1,2,···,n; j=1,2,···,d) represents the m-th historical monitoring data for the past j days at monitoring point k. Each row of the matrix S(m) represents the m-th historical and real-time monitoring data for a certain monitoring point. For example, S k (m)=[S k 0 (m),S k 1 (m),···,S k j (m),···,S kd (m)](j=1,2,...,d) indicates the m-th historical and real-time monitoring data of monitoring point k. Each column of the matrix S(m) indicates the m-th monitoring data time series of a day at a monitoring point. For example, S j (m)=[S1 j (m), S2 j (m),···,S k j (m),···,S n j (m)] T (k=1,2,···,n) indicates the mth monitoring data time series for the past j days at the monitoring point.
[0042] Furthermore, the step (2.2) specifically includes:
[0043] When the monitoring data time series to be detected are from different monitoring points, it is called inter-monitoring point time series anomaly detection. Under normal working conditions, the monitoring data from different monitoring points in a pipeline network follow the same law, that is, the distance between the monitoring data time series from different monitoring points does not change much. If an abnormal value appears in the monitoring data of a certain monitoring point, it will cause a change in the shape of the monitoring data time series from that monitoring point, and therefore cause a change in the distance between the monitoring data time series from the monitoring points. Through inter-monitoring point time series anomaly detection, such changes can be found, and then the abnormal value in the monitoring data can be detected.
[0044] To detect whether the time series distance between two monitoring points has changed, we first extract the row vectors corresponding to the monitoring data time series of the two monitoring points, and then calculate and obtain the distance between the real-time and historical monitoring data time series at the two monitoring points. Time series anomaly detection between monitoring points can be divided into three steps: (1) The distance d between the historical monitoring data time series m i (S1, S2) (i = 1, 2, , d) is calculated and obtained; (2) the distance d of the real-time monitoring data time series m 0 Calculate (S1,S2) and obtain (3)d m 0 (S1,S2) and dm i Compare with (S1,S2), d m 0 Check whether (S1,S2) is an outlier.
[0045] Distance d of historical monitoring data time series m i (S1, S2) is the distance between the monitoring data time series at time m between monitoring points 1 and 2 for the past d days, and the distance d between the real-time monitoring data time series before this m 0 (S1,S2) can be directly exported. Therefore, when detecting anomalies in the time series between monitoring points, the distance d of the real-time monitoring data time series m 0 It is only necessary to calculate and obtain (S1, S2). The distance between monitoring points 1 and 2 in the monitoring data time series at time m is as follows: d m (S1,S2)=[d m 0 (S1,S2),d m 2 (S1,S2),···,d m d (S1,S2)] In the formula, d m 0 (S1,S2) indicates the distance between the real-time monitoring data time series, and d m i (S1, S2) (i = 1, 2, , d) is the distance of the historical monitoring data time series.
[0046] The decision threshold ξ1(S1,S2) is calculated and obtained according to the distance of the historical monitoring data time series, and ξ1(S1,S2) is d m i It is obtained by calculating the mean value and variance of (S1, S2). ξ1(S1,S2)=μ(d m i (S1,S2))+m S1,S2 σ(d m i (S1,S2) In the formula, μ(d m i(S1,S2)) and σ(d m i (S1,S2)) are the distance d m i The mean and variance of (S1, S2) (i=1, 2, , d) are shown.
[0047] After obtaining the decision threshold ξ1(S1,S2), we use it as the distance d m 0 Compared with (S1, S2), the distance d m 0 Determine whether (S1, S2) is an abnormal value. d m 0 If (S1, S2) ≦ ξ1(S1, S2), it indicates that there is no abnormality in the monitoring data at the current time of monitoring points 1 and 2, and the detection result is output as "normal" and displayed as 0. In all other cases, the detection result is output as "abnormal" and displayed as 1.
[0048] Furthermore, the above step (2.3) specifically includes the following: When the monitoring data time series to be detected are from the same monitoring point, it is called monitoring point itself time series anomaly detection. Under normal working conditions, the monitoring data of each monitoring point in the pipeline network follows the same rule, that is, the distance between the monitoring data time series of the same monitoring point does not change much. When an abnormal value appears in the monitoring data of a certain monitoring point, it will cause a change in the shape of the monitoring data time series of that monitoring point, and therefore the distance between the monitoring data time series of the monitoring point itself will also change.
[0049] When detecting anomalies in the time series of the monitoring point itself, it is necessary to calculate and obtain the distance between the real-time and historical monitoring data time series, and the distance between the historical and historical monitoring data time series, and then compare them. The time series anomaly detection of the monitoring point itself can be divided into three main steps: (1) The distance d between the historical and historical monitoring data time series m i,j (i≠j; i,j=0,1,···,d) and obtain the distance d between the real-time and historical monitoring data time series. m 0,j (j=1,2,···,d) and obtain (3)d m0,j and d m i,j Compared with d m 0,j Check whether there are any outliers in
[0050] In fact, the distance d between the history and the historical monitoring data time series m i,j is already calculated in the previous monitoring point itself time series anomaly detection. m 0,j Only the minimum value min{d m 0,j}:min{d m 0,j}=min{d m 0,1 ,d m 0,2 ,···,d m 0,d}. The distance d of the mth monitoring data time series for the past d days m i,j The threshold ξ2(d m i,j ):ξ2(d m i,j )=μ(d m i,j )+mσ(d m i,j ) is obtained, where μ(d m i,j ) and σ(d m i,j ) are the distances d m i,j (i≠j;i,j=0,1,···,d) are the mean and variance.
[0051] Decision threshold ξ2(d m i,j ) is obtained, and then it is calculated as the minimum distance min{d m 0,j}, and min{d m 0,j Determine whether min{d m 0,j}≦ξ2(d mi,j ), it indicates that there is no abnormality in the monitoring data time series at the current time of the monitoring point, and the detection result is output as 0; otherwise, the detection result is output as 1.
[0052] Furthermore, step (3) specifically includes the following: Step (3.1): Use the provided single-point anomaly detection method to perform single-point anomaly detection for various abnormal scenarios in a water supply network. To determine the detection effectiveness of the unsupervised classification-based single-point anomaly detection method for various abnormal scenarios, various abnormal scenarios are used to test the provided method, mainly considering the following four scenarios: (1) a pipe burst occurs in the pipeline network; (2) dirty data appears at a single monitoring point; (3) dirty data appears at some monitoring points; and (4) dirty data appears at all monitoring points. Each scenario includes 25 sets of abnormal monitoring data, and single-point anomaly detection is performed starting from the time the abnormal value begins, and whether to issue an alert is determined based on the anomaly detection results at five consecutive times.
[0053] Step (3.2) uses the provided time series anomaly detection method to perform time series anomaly detection on various abnormal scenarios in a water supply network. To determine the detection effectiveness of the time series anomaly detection-based method on various abnormal scenarios, the provided method is tested using various abnormal scenarios, mainly considering the following four scenarios: (1) a pipe burst occurs in the pipeline network; (2) dirty data appears at a single monitoring point; (3) dirty data appears at some monitoring points; and (4) dirty data appears at all monitoring points. Each scenario contains 25 sets of abnormal monitoring data, and single-point anomaly detection is performed starting from the time the abnormal value begins, and whether to issue an alert is determined based on the anomaly detection results at five consecutive times.
[0054] In step (3.3), the single-point and time-series anomaly detection results are integrated to detect and identify various abnormal situations in the water supply network. The purpose of integrating the single-point and time-series anomaly detection results is to detect abnormal values in real-time monitoring data in a timely manner and accurately identify situations in which pipe ruptures occur in the pipeline network or dirty data appears at monitoring points. Anomaly detection is performed on the monitoring data of each monitoring point using a sliding time window, completing single-point and time-series anomaly detection for one time at a time. When a single-point or time-series anomaly is detected at a monitoring point, the time corresponding to the single-point or time-series anomaly is marked, and a data cleansing method is used to replace the anomaly. When the marked time window reaches a certain length, an anomaly warning is issued and various abnormal situations are identified. The process is mainly divided into the following steps:
[0055] Step (3.3.1) detects whether there are single-point or time-series anomalies in the monitoring data. First, distinguish between the normal working conditions of the piping network and the circumstances under which anomalies appear in the monitoring data. For single-point monitoring data, if a single-point anomaly is detected based on an unsupervised stacking integration algorithm or statistical theory, mark the time at which the single-point anomaly appears, and then perform single-point anomaly detection on the monitoring data at the next time. For time-series monitoring data, if an abnormal time series is found based on time-series anomaly detection between monitoring points or time-series anomaly detection at the monitoring point itself, mark the time at which the abnormal time series appears, and then perform time-series anomaly detection on the monitoring data at the next time.
[0056] Step (3.3.2) identifies a situation in which anomalies appear in the monitoring data of some monitoring points. If anomalies appear in the single points and time series of the monitoring data of some monitoring points, but no anomalies appear in the single points and time series of the monitoring data of the remaining monitoring points, it is determined that dirty data appears at some monitoring points. For example, if an anomaly appears in the monitoring data of monitoring point 1, but no anomalies appear in the monitoring data of monitoring points 2 and 3, a single-point anomaly can be detected in the monitoring data of monitoring point 1, and the time series of monitoring point 1 itself will also have a continuous abnormal warning status, while neither monitoring points 2 nor 3 will have a warning status. The time series distance d(S1,S2) of the monitoring data of monitoring points 1 and 2 and the time series distance d(S1,S3) of the monitoring data of monitoring points 1 and 3 are both abnormal values, and the time series distance d(S2,S3) of the monitoring data of monitoring points 2 and 3 is a normal value.
[0057] Step (3.3.3) identifies the situation where a pipe burst occurs and the situation where dirty data appears at all monitoring points. If abnormalities are detected in both the single point and time series of the monitoring data at all monitoring points, the situation where a pipe burst occurs in the pipe network and the situation where dirty data appears at all monitoring points are identified through the increase and decrease of the abnormal values. If the single point qualitative detection results of the monitoring data at all monitoring points are all "decreasing", it can be determined that a pipe burst will occur in the pipe network. Conversely, if the single point qualitative detection result of a certain monitoring data is "increasing", it can be determined that dirty data has appeared at all monitoring points in the pipe network.
[0058] In step (3.3.4), the pipe burst detection performance of the method is evaluated and analyzed. The provided method aims to detect single-point and time-series anomalies in real-time monitoring data, and combines the single-point and time-series anomaly detection results to identify situations in which a pipe burst occurs in the pipe network and situations in which dirty data appears at the monitoring point. Therefore, when evaluating the method, the following indicators are mainly considered: (1) detection accuracy rate (σ1); (2) anomaly identification rate (σ2); and (3) anomaly detection rate (σ3). The detection accuracy rate refers to the probability of detecting various abnormal situations, i.e., identifying and alerting on single-point or time-series anomalies in abnormal situations. The detection accuracy rate is expressed as follows: σ1=N dn / N n In the formula, N n denotes the total number of abnormal scenes, and N dn indicates the number of abnormal scenes detected.
[0059] The anomaly identification rate refers to the probability that various abnormal situations are correctly identified. For example, in 10 types of abnormal situations, the number of pipe bursts occurring in the pipeline network is 8, and the number of detected pipe burst incidents is 4, so the anomaly identification rate of pipe bursts occurring in the pipeline network is 4 / 8=50%. The anomaly identification rate can be expressed as follows: σ2=N in / N n In the formula, N in denotes the number of correctly identified anomalous scenes, and N n indicates the total number of abnormal scenes.
[0060] The anomaly detection rate is the ratio between the duration of various detected abnormal scenes and the actual duration of the abnormal scenes. For example, if the duration of dirty data appearing at a certain monitoring point is 1 hour, and the duration of dirty data appearing at the monitoring point detected is 0.5 hours, the anomaly detection rate is 0.5 / 1=50%. The anomaly detection rate can be expressed as follows: σ3=t dn / t t In the formula, t dn and t t and denote the detected duration and the actual duration of the abnormal event, respectively. Obviously, the higher σ1, σ2 and σ3 are, the better.
[0061] Step (3.4): The pipe burst detection method is used to detect and identify various abnormal situations in a water supply network, and the pipe burst detection performance of the provided method is evaluated. Step (3.4.1): When detecting and identifying various abnormal situations in a water supply network, single-point abnormal value detection results are removed, and the abnormality detection rate, detection accuracy rate, and abnormality identification rate for various abnormal situations are obtained. Step (3.4.2): When detecting and identifying various abnormal situations in a water supply network, single-point qualitative value detection results are removed, and the abnormality detection rate, detection accuracy rate, and abnormality identification rate for various abnormal situations are obtained. Step (3.4.3): When detecting and identifying various abnormal situations in a water supply network, time series abnormality detection results between monitoring points are removed, and the abnormality detection rate, detection accuracy rate, and abnormality identification rate for various abnormal situations are obtained. Step (34.4): When detecting and identifying various abnormal situations in a water supply network, time series abnormality detection results from the monitoring point itself are removed, and the abnormality detection rate, detection accuracy rate, and abnormality identification rate for various abnormal situations are obtained.
[0062] In the above process, various abnormal situations are detected using single-point anomaly detection methods and time-series anomaly detection methods, respectively. The single-point and time-series anomaly detection results are integrated to obtain the final anomaly detection results for various situations, and four integration policies are used: i. Remove single-point anomaly value detection results; ii. Remove single-point value qualitative detection results; iii. Remove time-series anomaly detection results between monitoring points; iv. Remove time-series anomaly detection results for the monitoring point itself. The final anomaly detection results under various situations are obtained and saved in anomaly detection results.csv. Then, the performance indicators of the provided method for various abnormal situations, namely, the detection accuracy rate, anomaly identification rate, and anomaly detection rate, are calculated and obtained, and saved in anomaly evaluation results.csv.
[0063] A specific example is as follows.
[0064] An embodiment of the present invention provides a method for detecting and evaluating the performance of a water supply network pipe rupture based on single-point and time-series anomaly detection. The method utilizes the history and real-time monitoring data of the water supply network to effectively detect and identify pipe ruptures and monitoring point failures in the water supply network, and by performing single-point and time-series anomaly detection on the real-time monitoring data of the water supply network, it is possible to timely discover abnormal working conditions in the water supply network, and accurately distinguish between various abnormal situations according to the various warning situations of the single-point and time-series anomaly detection results. The method specifically includes the following steps:
[0065] Step (1): Perform single-point anomaly detection on the real-time monitoring data of the water supply network to distinguish between normal and abnormal working conditions of the water supply network.
[0066] Step (1.1): A hydraulic simulation is performed on a water supply network to obtain pressure monitoring data under various abnormal conditions in the water supply network, so as to perform single-point anomaly detection on the abnormal monitoring data. Figure 2 shows a schematic diagram of single-point anomaly detection. The present invention mainly considers three types of abnormal conditions: (a) a pipe burst occurs in the water supply network; (b) a failure occurs at some monitoring points in the water supply network; and (c) a failure occurs at all monitoring points in the water supply network. Step (1.2): Based on the provided unsupervised stacking integration algorithm, single-point anomaly detection is performed on the pressure monitoring data under various abnormal working conditions in the water supply network.
[0067] Figure 3 shows a schematic diagram of single-point outlier detection based on the unsupervised stacking integration algorithm. The diagram includes real-time and historical monitoring data from two monitoring points, designated Monitoring Point 1 and Monitoring Point 2, respectively. Figures 3(a) and 3(b) show the historical monitoring data from Monitoring Point 1 and Monitoring Point 2, while Figure 3(c) shows the real-time monitoring data from Monitoring Point 1 and Monitoring Point 2. Both monitoring points contain 15 days of historical monitoring data, i.e., d = 15. The monitoring data collection frequency is 15 minutes per session. The first monitoring data is collected at 0:15 and the second monitoring data is collected at 0:30. The total number of daily monitoring data at each monitoring point is 96, i.e., m = 96.
[0068] As shown in Figure 3(a), the top 95 real-time monitoring data of monitoring points 1 and 2 are all normal values, and there is no significant difference compared to the historical monitoring data in Figures 3(b) and 3(c). As shown in Figure 3(a), the top 96 real-time monitoring data (x1 0 (96) and x2 0 (96)) is the historical monitoring data (x1 i (96) and x2 i (96) (i=1,2,...,15)). As shown in Figure 3(d), monitoring point 1 is the 96th historical monitoring data x1 for the past 15 days. i (96) is distributed within the yellow area (i.e., the normal range), and is the 96th real-time monitoring data x1 of monitoring point 1. 0 (96) is located above the yellow area and clearly deviates from the normal value. As shown in Figure 3(e), monitoring point 2 is the 96th historical monitoring data x2 from the past 15 days. i (96) is distributed entirely within the yellow area (i.e., the normal range), and is the 96th real-time monitoring data of monitoring point 2 x2 0 (96) is distributed above the yellow area and clearly deviates from the normal value.
[0069] As shown in Fig. 3(f), the m-th real-time monitoring data x n 0 If you need to detect whether (m) is an outlier, then xn 0 (m) and the mth historical monitoring data for the past d days x n i (m) (i = 1, 2, , d) A total of d + 1 values are input to the unsupervised stacking integration algorithm, and each x n Obtain the probability that (m) is an outlier. x n 0 If (m) has the highest probability of being an outlier, mark it as an outlier; otherwise, x n 0 (m) indicates normal value and is not marked.
[0070] Step (1.3) utilizes statistical theory to perform qualitative detection on single values in the water supply network monitoring data and identify single-point anomalies in the real-time monitoring data. Figure 4 shows a schematic diagram of single-point anomaly detection based on statistical theory. The diagram includes real-time and historical monitoring data from two monitoring points, designated monitoring points 1 and 2, respectively. Figures 4(a) and 4(b) show the historical monitoring data from monitoring points 1 and 2, while Figure 4(c) shows the real-time monitoring data from monitoring points 1 and 2. Both monitoring points contain 15 days of historical monitoring data, i.e., d = 15. The monitoring data collection frequency is 15 minutes per session. The first monitoring data is collected at 0:15 and the second monitoring data is collected at 0:30. The total number of daily monitoring data at each monitoring point is 96, i.e., m = 96.
[0071] As shown in Figure 4(a), the top 95 real-time monitoring data of monitoring points 1 and 2 are all normal, and the historical monitoring data (x1 i (96) and x2 i There is no significant difference compared to (96) (i=1,2,...,15). The 96th monitoring data x1 of monitoring point 1 0 (96) is historical monitoring data x1 i (96) (i=1,2,···,15) is clearly lower than the 96th monitoring data x2 0 (96) is historical monitoring data x2 i This is a clear increase compared to (96) (i = 1, 2, · · · , 15).
[0072] As shown in Fig. 4(d) and Fig. 4(e), the light gray area indicates the normal range of the monitored data, and the dark gray line indicates the average value of the 15 monitored data. k σ and μ-m k σ indicates the upper and lower limits of the range, respectively. As shown in Figure 4(d), the 96th historical monitoring data for monitoring point 1 are all distributed within the light gray region, and the 96th real-time monitoring data is located below the light gray region, i.e., the real-time monitoring data is smaller than the historical monitoring data. As shown in Figure 4(e), the 96th historical monitoring data for monitoring point 2 are all distributed within the light gray region, and the 96th real-time monitoring data is located above the light gray region, i.e., the monitoring data is increasing. Therefore, based on statistical theory, single values of the monitoring data can be detected and divided into three types: (1) abnormal values (increase); (2) normal values; and (3) abnormal values (decrease).
[0073] Figure 5 shows a schematic diagram of single-point anomaly detection after a pipe rupture occurs in a pipeline network. Figure 5(a) shows the time series curves of pressure monitoring data from monitoring points 1 and 2. The monitoring data collection frequency is 15 minutes, and each monitoring point contains 96 monitoring data per day. Normal values are distributed within the light gray area, while abnormal values are distributed within the green area. From November 12 to November 15, the pipeline network was under normal operating conditions. November 16 is real-time monitoring data. A pipe rupture occurred in the pipeline network at 10:30, and monitoring points 1 and 2 showed a continuous decline from the 42nd monitoring data.
[0074] Figure 5(b) provides a schematic diagram comparing five sets of real-time and historical monitoring data after a pipe burst occurs in a pipeline network at monitoring points 1 and 2. As shown in the figure, the five sets of real-time monitoring data (42, 43, 44, 45, and 46) at monitoring points 1 and 2 all show a clear decrease compared to the corresponding historical monitoring data for the past four days. In the figure, all five sets of historical monitoring data are located within the light gray area (normal), while all five sets of real-time monitoring data are located within the dark gray area (abnormal). When single-point anomaly detection is performed on the five sets of real-time and historical monitoring data using an unsupervised stacking integration algorithm, the real-time monitoring data is marked as "abnormal." At the same time, the real-time monitoring data is all decreased compared to the historical data. That is, when single-point qualitative detection is performed on the real-time and historical monitoring data using statistical theory, the detection result is "decreased." Therefore, after a pipe burst occurs in a pipeline network, single-point anomaly detection typically yields the following results: (1) Anomalies are detected in the monitoring data at each monitoring point. (2) Abnormal values tend to be lower than normal values.
[0075] Figure 6 shows a schematic diagram of single-point anomaly detection after a pipe rupture occurs in a pipeline network. Figure 6(a) shows the time series curves of pressure monitoring data from monitoring points 1 and 2. The monitoring data collection frequency is 15 minutes, and each monitoring point contains 96 monitoring data per day. Normal values are distributed within the yellow region, and abnormal values are distributed within the green region. From November 12 to November 15, the pipeline network was under normal operating conditions. November 16 is real-time monitoring data. A pipe rupture occurred in the pipeline network at 10:30, and monitoring points 1 and 2 showed a continuous decline from the 42nd monitoring data.
[0076] Figure 6(b) provides a schematic diagram comparing five sets of real-time and historical monitoring data after a pipe rupture occurred in the pipeline network at monitoring points 1 and 2. As shown in the figure, five sets of real-time monitoring data (42, 43, 44, 45, and 46) from monitoring points 1 and 2 all show a significant decline compared to the corresponding historical monitoring data for the past four days. In the figure, all five sets of historical monitoring data are located within the yellow region (normal), while all five sets of real-time monitoring data are located within the green region (abnormal). When single-point anomaly detection is performed on the five sets of real-time and historical monitoring data using an unsupervised stacking integration algorithm, the real-time monitoring data is marked as "abnormal." At the same time, the real-time monitoring data is all decreased compared to the historical data. That is, when single-point qualitative detection is performed on real-time and historical monitoring data using statistical theory, the detection result is "decreased." Therefore, after a pipe burst occurs in a pipeline network, single-point anomaly detection usually produces the following results: (1) Anomalies are detected in the monitoring data of each monitoring point; (2) Anomalies tend to decrease compared to normal values.
[0077] Figure 7 shows a schematic diagram of single-point anomaly detection after dirty data appears at some monitoring points. Figure 7(a) shows the time series curves of pressure monitoring data from monitoring points 1 and 2. The monitoring data collection frequency is 15 minutes, and each monitoring point contains 96 monitoring data per day. The data from November 12th to November 15th is historical monitoring data, and the data from November 16th is real-time monitoring data. Monitoring point 1 has continuously displayed multiple dirty data points (42, 43, 44, 45, 46, and 47) since 10:30, while the real-time monitoring data from monitoring point 2 is all normal.
[0078] As shown in Figure 7(b), five sets of real-time monitoring data (42, 43, 44, 45, and 46) at monitoring point 1 are significantly different from the historical monitoring data corresponding to the past four days. Here, real-time monitoring data 42, 43, 44, 45, and 46 show increases, while real-time monitoring data 44 and 45 show decreases. In contrast, the five sets of real-time monitoring data at monitoring point 2 are not significantly different from the historical monitoring data and are all distributed within the yellow region. When using an unsupervised stacking integration algorithm to perform single-point anomaly detection on the five real-time monitoring data at monitoring points 1 and 2, the anomaly detection result for monitoring point 1 is "abnormal," and the anomaly detection result for monitoring point 2 is "normal." At the same time, when using statistical theory to perform qualitative detection on the real-time monitoring data at monitoring points 1 and 2, the anomaly detection result for monitoring point 1 is "increase" or "decrease," and the anomaly detection result for monitoring point 2 is "normal." Therefore, after dirty data appears at some monitoring points, single-point anomaly detection usually produces the following results: (1) anomalies are detected at some monitoring points; (2) the abnormal values tend to decrease or increase compared to the normal values.
[0079] Figure 8 shows a schematic diagram of single-point anomaly detection after dirty data appears at all monitoring points. Figure 8(a) shows the time series curves of pressure monitoring data for monitoring points 1 and 2. The monitoring data collection frequency is 15 minutes, and each monitoring point contains 96 monitoring data points per day. The historical monitoring data from November 12 to November 15 is normal. The real-time monitoring data from November 16 shows that multiple dirty data points (42, 43, 44, 45, 46, and 47) appeared consecutively at monitoring points 1 and 2 from 10:30.
[0080] As shown in Figure 8(b), five sets of real-time monitoring data (42, 43, 44, 45, and 46) at monitoring points 1 and 2 are significantly different from the historical monitoring data corresponding to the past four days. Here, the real-time monitoring data of 42, 43, and 46 increase, while the real-time monitoring data of 44 and 45 decrease. When using the unsupervised stacking integration algorithm to perform single-point outlier detection on the five sets of real-time monitoring data at monitoring points 1 and 2, the following results are obtained: x1 0 (42), x1 0 (43), x1 0 (44), x1 0 (45), x1 0 (46), x2 0 (42), x2 0 (43), x2 0 (44), x2 0 (45), and x2 0 (46) are both abnormal values, that is, the single-point abnormal value detection results of monitoring points 1 and 2 are "abnormal". When qualitative detection is performed on the real-time monitoring data of monitoring points 1 and 2 using statistical theory, the following results are obtained: x1 0 (42), x1 0 (43), x1 0 (46), x2 0 (42), x2 0 (43), and x2 0 (46) increases by x1 0 (44), x1 0 (45), x2 0 (44), and x2 0 (45) decreases, that is, the single-point anomaly qualitative detection results of monitoring points 1 and 2 are "increase" or "decrease". Therefore, after dirty data appears at all monitoring points, the single-point anomaly detection results are as follows: (1) Anomalies are detected at all monitoring points; (2) The anomaly values tend to increase or decrease compared to the accurate values.
[0081] In summary, single-point abnormal values are detected in both the situation where a pipe burst occurs in a pipeline network and the situation where dirty data appears at a monitoring point. The single-point abnormal value detection results are different for different abnormal events, as shown below: (1) After a pipe burst occurs in a pipeline network, single-point abnormal values are detected at all monitoring points, and all abnormal values are smaller than normal values; (2) After dirty data appears at some monitoring points, single-point abnormal values are detected at the monitoring points where dirty data appears, and the abnormal values tend to increase or decrease; (3) After dirty data appears at all monitoring points, single-point abnormal values are detected at all monitoring points, and the abnormal values tend to increase or decrease.
[0082] Step (2): Perform time series anomaly detection on the real-time monitoring data of the water supply network to distinguish between normal and abnormal working conditions of the water supply network. Step (2) specifically includes the following: Step (2.1): Perform hydraulic simulation on the water supply network to obtain pressure monitoring data under various abnormal scenarios of the water supply network, so as to perform time series anomaly detection on the abnormal monitoring data. The present invention mainly considers three types of abnormal scenarios: (a) a pipe burst occurs in the water supply network; (b) a fault occurs at some monitoring points in the water supply network; or (c) a fault occurs at all monitoring points in the water supply network. Step (2.2): Perform anomaly detection on the time series of monitoring data from different monitoring points, i.e., identify time series anomalies between monitoring points, as shown in Figure 10.
[0083] Figure 9 shows a schematic diagram of time series anomaly detection between monitoring points. Figure 9(a) shows the time series curves of real-time and historical monitoring data at monitoring points 1 and 2. December 1st to December 15th are historical monitoring data, and December 16th is real-time monitoring data. The historical monitoring data are normal values, the 95th real-time monitoring data at monitoring point 1 is an abnormal value, and all other real-time monitoring data are normal values. d m j If (S1, S2) denotes the distance between the mth monitoring data time series on the jth day at monitoring points 1 and 2, the distance between the mth monitoring data time series at monitoring points 1 and 2 is d m (S1,S2)=[dm 0 (S1,S2),d m 1 (S1,S2),···,d m j (S1,S2),···,d m 15 (S1,S2)], and d m 0 (S1, S2) is the distance between the real-time monitoring data time series of monitoring points 1 and 2, and d m j (S1, S2) (j=1, 2, , 15) is the distance between the historical monitoring data time series of monitoring points 1 and 2. Time series anomaly detection between monitoring points is performed by d m 0 (S1,S2) and d m j Compare with (S1,S2), d m 0 The goal is to check whether (S1,S2) is an outlier. m 0 If (S1, S2) is an abnormal value, it is found that the real-time monitoring data time series S1(m) or S2(m) of monitoring point 1 or 2 changes, and the inter-monitoring point time series is marked as "abnormal." m 0 If (S1, S2) is a normal value, mark the inter-point time series as "normal".
[0084] FIG. 9(b) shows the 94th monitoring data time series at monitoring points 1 and 2. The monitoring data at monitoring points 1 and 2 are both normal values, and the 94th monitoring data time series at monitoring point 1 is S1(94)=[S1 0 (94), S1 1 (94),···,S1 15 (94))] is similar in shape to the 94th monitoring data time series S2(94) = [S2 0 (94), S2 1 (94),···,S2 15 (94))] is small. Therefore, the distance d 94 0 (S1, S2) is the distance d 94j Compared to (S1,S2)(j=1,2,···,15), the change is not large, i.e., d 94 0 (S1, S2) is the normal value, and the time series between monitoring points 1 and 2 is "normal."
[0085] FIG. 9(c) shows the 95th monitoring data time series at monitoring points 1 and 2. The 95th real-time monitoring data at monitoring point 1 is an abnormal value, while the other monitoring data are all normal values. As shown in the figure, the 95th monitoring data time series S1 at monitoring point 1 0 (95) is the historical monitoring data time series S1 i (95) (j=1,2,···,15), the shape has changed significantly, and the 95th monitoring data time series S2(95)=[S2 0 (95), S2 1 (95),···,S2 15 (95))] is small. Therefore, the distance d 95 0 (S1, S2) is the historical monitoring data time series d 95 j (S1,S2)(j=1,2,···,15) 95 0 (S1, S2) is an anomaly, and the time series between monitoring points 1 and 2 is "anomalous."
[0086] Step (2.3), anomaly detection is performed on the monitoring data time series at the same monitoring point, i.e., the time series anomalies of the monitoring point itself are identified, as shown in Fig. 11 .
[0087] Figure 10 shows a schematic diagram of time series anomaly detection for the monitoring point itself. Figure 10(a) shows the time series curves of real-time and historical monitoring data for monitoring points 1 and 2. December 1st to December 15th are historical monitoring data, and December 16th is real-time monitoring data. All historical monitoring data are normal values, the 95th real-time monitoring data at monitoring point 1 is an abnormal value, and all other real-time monitoring data are normal values.
[0088] d m i,j If (S1) (i≠j; i,j=0,1,···,15) denotes the distance between the m-th monitoring data time series on the i-th day and the j-th day at monitoring point 1, the distance between the m-th real-time monitoring data time series at monitoring point 1 and the m-th historical monitoring data time series for the past 15 days is d m 0,j (S1)=[d m 0,1 (S1),d m 0,2 (S1),···,d m 0,15 (S1)], and the distance between the mth historical monitoring data time series for the past 15 days is d m i,j (S1)=[d m 1,2 (S1),d m 1,3 (S1),···,d m 14,15 (S1)]. The time series abnormality detection of the monitoring point itself is d m 0,j (S1)(j=1,2,···,15) and d m i,j (S1) (i≠j; i, j=0,1,···,15) and d m 0,j (S1) is to check whether the distance of the monitoring data time series is an anomaly. m 0,j If (S1) is an abnormal value, it is found that the m-th real-time monitoring data time series at monitoring point 1 has changed, and the monitoring point 1's own time series is marked as "abnormal", and all d m 0,j If (S1) is a normal value, the time series anomaly detection result of monitoring point 1 itself is marked as "normal."
[0089] Figure 10(b) shows the 94th monitoring data time series at monitoring point 1. All of the monitoring data at monitoring point 1 are normal values, and the shapes of the 94th monitoring data time series at monitoring point 1 are similar. Therefore, the distance d m i,j(S1) (i≠j; i,j=0,1,···,15) changes are small, that is, all d m 0,j (S1) is a normal value, and the time series anomaly detection result of monitoring point 1 itself is marked as "normal."
[0090] FIG. 10(c) shows the 95th monitoring data time series at monitoring point 1. The 95th real-time monitoring data at monitoring point 1 is an abnormal value, while the other monitoring data are all normal values. As shown in the figure, the 95th monitoring data time series S1 at monitoring point 1 0 (95) is the historical monitoring data time series S1 i (95) (i = 0, 1, . . . , 15), the shape changes significantly. Therefore, the distance d m 0,j (S1)(j=1,2,...,15) is the distance d m i,j Compared to (S1) (i≠j; i,j=0,1,···,15), there is a similarly large change, i.e., d m 0,j An abnormal value exists in (S1), and the time series abnormality detection result of monitoring point 1 itself is marked as "abnormal."
[0091] Figure 12 shows a schematic diagram of time-series anomaly detection after a pipe rupture occurs in a piping network. Figure 10(a) shows the time-series curves of pressure monitoring data from monitoring points 1 and 2, each of which contains 96 daily monitoring data. November 14th and 15th are historical monitoring data, and both are normal values. November 16th is real-time monitoring data, and a pipe rupture occurred in the piping network at 10:30. Monitoring points 1 and 2 show a continuous decline from the 42nd monitoring data.
[0092] Figure 12(b) shows a schematic diagram of time series anomaly detection between monitoring points. For monitoring points 1 and 2, the shape of the monitoring data time series of four sets (42nd, 43rd, 44th, and 45th) on the 16th changes significantly compared to the monitoring data time series on the 15th. However, since the monitoring data time series of monitoring points 1 and 2 change simultaneously and the rules of change are similar, the change in the time series distance between monitoring points 1 and 2 is not significant. For example, for monitoring points 1 and 2, the distance between the 42nd time series on the 16th is d 42 0 (S1, S2), and the distance d 42 1 The change is smaller than (S1, S2), i.e., the anomaly detection result of the 42nd time series on the 16th is "normal" for monitoring points 1 and 2. As shown in the figure, when anomaly detection is performed on the distances of the 43rd, 44th, and 45th monitoring data time series on the 16th for monitoring points 1 and 2, there is a possibility that all the detection results will be "normal."
[0093] Figure 12(c) and Figure 12(d) show schematic diagrams of time series anomaly detection for monitoring points 1 and 2, respectively. As shown in Figure 12(c), for monitoring point 1, the shape of the monitoring data time series for three sets (42nd, 43rd, and 44th) on the 16th changes significantly compared to the monitoring data time series for the 14th and 15th, so the change in distance between the real-time and historical monitoring data time series is large. For example, the distance between the 42nd real-time (16th) and historical (15th) monitoring data time series at monitoring point 1 is d 45 16,15 (S1), and the distance between the 42nd history (15th day) and the history (14th day) is d 45 15,14 (S1) and d 45 16,15 (S1) is clearly d 45 15,14 (S1), that is, the time series anomaly detection result of monitoring point 1 itself is "abnormal." Similarly, the 42nd, 43rd, and 44th time series anomaly detection results of monitoring point 2 itself are all "abnormal."
[0094] In summary, after a pipe burst occurs in a pipeline network, the time series anomaly detection usually achieves the following results: (1) The time series detection results between monitoring points may be abnormal or may be normal; (2) The time series of the monitoring points themselves are all detected as abnormal.
[0095] Figure 13 shows a schematic diagram of time series anomaly detection after dirty data appears at some monitoring points. Figure 13(a) shows the time series curves of pressure monitoring data for monitoring points 1 and 2, each of which contains 96 daily monitoring data. November 14th to 15th are historical monitoring data, all of which are normal values. November 16th is real-time monitoring data, and monitoring point 1 has multiple dirty data items (44th, 45th, 46th, and 45th) appearing consecutively from 10:30, while monitoring point 2 has all normal values.
[0096] Figure 13(b) shows a schematic diagram of time series anomaly detection between monitoring points. For monitoring point 1, the four sets of monitoring data time series (42nd, 43rd, 44th, and 45th) on the 16th change significantly compared to the monitoring data time series on the 15th. At the same time, the four sets of monitoring data time series for monitoring point 2 change less significantly. Therefore, the real-time monitoring data time series for monitoring points 1 and 2 change significantly compared to the historical monitoring data time series. For example, the distance between monitoring points 1 and 2 for the 42nd monitoring data time series on the 16th is d 42 0 (S1, S2), and the distance d 42 1 The change is larger than (S1, S2), that is, the anomaly detection result for the 42nd monitoring data time series on the 16th at monitoring points 1 and 2 is "Abnormal." As shown in the figure, when anomaly detection is performed on the distances of the 43rd, 44th, and 45th monitoring data time series on the 16th at monitoring points 1 and 2, the detection result is "Abnormal."
[0097] Figure 13(c) and Figure 13(d) show schematic diagrams of time series anomaly detection for monitoring points 1 and 2, respectively. As shown in Figure 11(c), the monitoring data time series of three sets (42nd, 43rd, and 44th) on the 16th day at monitoring point 1 show a large change in shape compared to the monitoring data time series on the 14th and 15th days, so the change in distance between the real-time and historical monitoring data time series is large. For example, the distance between the 42nd real-time (16th) and historical (15th) monitoring data time series at monitoring point 1 is d 42 16,15 (S1), and the distance between the 42nd history (15th day) and the history (14th day) is d 42 15,14 (S1) and d 42 16,15 (S1) is clearly d 42 15,14 (S1), that is, the time series anomaly detection result of monitoring point 1 itself is "abnormal."
[0098] As shown in Figure 13(d), at monitoring point 2, the monitoring data time series of the three sets (42nd, 43rd, and 44th) on the 16th show smaller changes than the monitoring data time series on the 14th and 15th, and the change in the distance between the real-time and historical monitoring data time series is not large. For example, at monitoring point 1, the distance between the 42nd real-time (16th) and historical (15th) monitoring data time series is d 42 16,15 (S2), and the distance between the 42nd history (15th day) and the history (14th day) is d 42 15,14 (S2) and d 42 16,15 (S2) is d 42 15,14 There is no significant difference compared to (S2), that is, the time series S anomaly detection result for monitoring point 2 itself is "normal."
[0099] In summary, after dirty data occurs at some monitoring points, time series anomaly detection usually produces the following results: (1) anomalies are detected in the time series between monitoring points; (2) anomalies are detected in all of the time series of some monitoring points themselves.
[0100] Figure 14 shows a schematic diagram of time series anomaly detection after dirty data appears at all monitoring points. Figure 4.4.6(a) shows the pressure monitoring data time series curves for monitoring points 1 and 2, each containing 96 daily monitoring data. November 14th and 15th are historical monitoring data, and all historical monitoring data are normal values. November 16th is real-time monitoring data, and monitoring points 1 and 2 have seen multiple abnormal values (42nd, 43rd, 44th, 45th, and 46th) appear consecutively from 10:30.
[0101] Figure 14(b) shows a schematic diagram of time series anomaly detection between monitoring points. At monitoring points 1 and 2, the four sets of monitoring data time series (43rd, 44th, 45th, and 46th) on the 16th change significantly compared to the monitoring data time series on the 15th. If the change trends of the real-time monitoring data time series at monitoring points 1 and 2 match, the time series anomaly detection result between monitoring points is "normal." If the real-time monitoring data time series at monitoring points 1 and 2 change significantly, the time series anomaly detection result between monitoring points is "abnormal." For example, the distance between monitoring points 1 and 2 for the 43rd monitoring time series on the 16th is d 43 0 (S1, S2), and the distance d 43 1 If the change is smaller than (S1, S2), the abnormality detection result of the 43rd time series on the 16th is "normal" for monitoring points 1 and 2. The distance between the 46th monitoring data time series on the 16th is d 46 0 (S1, S2), and the distance d 46 1 There is a large change compared to (S1, S2), that is, for monitoring points 1 and 2, the 46th time-series anomaly detection result on the 16th is "abnormal."
[0102] Figure 14(c) and Figure 14(d) show schematic diagrams of time series anomaly detection for monitoring points 1 and 2, respectively. As shown in Figure 14(c), for monitoring point 1, the shape of the three sets of monitoring data time series (42nd, 43rd, and 44th) on the 16th changes significantly compared to the monitoring data time series on the 14th and 15th, so the change in distance between the real-time and historical monitoring data time series is large. For example, the distance between the 42nd real-time (16th) and historical (15th) monitoring data time series at monitoring point 1 is d 42 16,15 (S1), and the distance between the 42nd history (15th day) and the history (14th day) is d 42 15,14 (S1) and d 42 16,15 (S1) is clearly d 42 15,14 (S1), that is, the time series anomaly detection result of monitoring point 1 itself is "abnormal." Similarly, as shown in FIG. 14(d), the time series anomaly detection results of monitoring point 2 itself are all "abnormal."
[0103] In summary, after dirty data appears at all monitoring points, time series anomaly detection usually achieves the following results: (1) the time series detection results between monitoring points may be abnormal or may be normal; (2) all the monitoring points' own time series will have anomalies detected.
[0104] The situation where a pipe burst occurs in a pipe network and dirty data appears at some (all) of the monitoring points causes anomalies in the monitoring data time series, but various time series anomaly detection results appear, as shown below: (1) A pipe burst occurs in a pipe network. The time series anomaly detection results between the monitoring points may be normal or may be abnormal. Anomalies are detected in the time series of all the monitoring points themselves; (2) Anomalies are detected in the time series of some of the monitoring points themselves; (3) Dirty data appears at all detection points. The time series anomaly detection results between the monitoring points may be normal or may be abnormal. Anomalies are detected in the time series of all the monitoring points themselves.
[0105] Step (3): Experiments are conducted using the reference test pipe network Net 3 pipe network model to verify and evaluate the pipe burst detection performance of the proposed method.
[0106] As shown in FIG. 15, it is assumed that three pressure monitoring points are distributed in the piping network, and the distributed nodes are 169, 204 and 275 in order, which are designated as pressure monitoring points 1, 2 and 3, respectively.
[0107] Using the monitoring data from three pressure monitoring points, single-point and time-series anomaly detection was performed. The experimental conditions and data preparation technique process are shown in Figure 16 and include the following steps: (a) through simulation, obtain monitoring data regarding the occurrence of pipe ruptures in the pipeline network or the appearance of dirty data at the monitoring points; (b) based on a single-point unsupervised classification method, perform single-point anomaly detection on the monitoring data; (c) based on a time-series unsupervised classification method, perform time-series anomaly detection on the monitoring data; (d) integrate the single-point and time-series anomaly detection results to obtain detection results for various abnormal scenes.
[0108] Figure 16 shows a schematic diagram of the monitoring data.csv file, which includes the collection date, time, time step, status and tab information of the monitoring data at each monitoring point. The status of the monitoring data is divided into two types: normal value (0) and abnormal value (1). The monitoring data tab is divided into monitoring data under normal working conditions of the piping network (0) and monitoring data under abnormal working conditions of the piping network (1-9). The situations in which abnormalities appear in various monitoring data are shown in Table 1.
[0109] Table 1. Various situations where abnormalities appear in monitoring data [Table 1]
[0110] Figure 17 shows the results of single-point anomaly detection based on the unsupervised stacking integration algorithm. Four scenarios are considered: (1) As shown in Figure 17(a), a pipe burst occurs in the pipeline network, and the monitoring data from monitoring points 1, 2, and 3 all show abnormal values; (2) As shown in Figure 17(b), the monitoring data from monitoring point 1 shows all abnormal values, and the monitoring data from monitoring points 2 and 3 all show normal values; (3) As shown in Figure 17(c), the monitoring data from monitoring points 1 and 2 all show abnormal values, and the monitoring data from monitoring point 3 all show normal values; (4) As shown in Figure 17(d), the monitoring data from monitoring points 1, 2, and 3 all show abnormal values. If the monitoring data at five consecutive times are all detected as abnormal, an abnormality alert is issued and the anomaly detection result is output as "Abnormal," indicated by a dot in the yellow area in the figure. Conversely, if no abnormal event occurs, the anomaly detection result is output as "Normal," indicated by a dot in the green area in the figure. As shown in Figure 17, the unsupervised stacking integration algorithm can detect most of the abnormal values in the anomaly detection data, while not issuing an abnormality warning for some of the abnormal monitoring data. Some of the abnormal values in the monitoring data may have small changes, making it difficult to meet the requirement of detecting an abnormality at all five consecutive times. Therefore, some of the abnormal values do not issue an abnormality warning. In addition, the unsupervised stacking integration algorithm can detect situations where a pipe burst occurs in a pipe network and situations where dirty data appears at all monitoring points, but it cannot distinguish between these two types of abnormal situations.
[0111] Figure 18 shows the results of single-point anomaly detection based on statistical theory. Four scenarios are considered: (1) As shown in Figure 18(a), a pipe burst occurs in the pipeline network, and all the monitoring data at monitoring points 12 and 3 are abnormal values; (2) As shown in Figure 4.5.5(b), all the monitoring data at monitoring point 1 is abnormal, but all the monitoring data at monitoring points 2 and 3 are normal; (3) As shown in Figure 18(c), all the monitoring data at monitoring points 1 and 2 are abnormal, and all the monitoring data at monitoring point 3 is normal; (4) As shown in Figure 18(d), all the monitoring data at monitoring points 1, 2, and 3 are abnormal. If all the monitoring data at five consecutive times are detected as abnormal, an abnormality warning is issued: (1) if the abnormal value is greater than the normal value, the abnormality detection result is "increase," indicated by a dot in the yellow area; (2) if all the abnormal values are less than the normal value, the abnormality detection result is "decrease," indicated by a dot in the orange area. On the other hand, it indicates that there is no abnormal event, and outputs the abnormality detection result "normal", which is indicated by a dot in the green area. As shown in Figure 18, the method based on statistical theory can detect some abnormal monitoring data, but most of the abnormal monitoring data has not been detected yet, but the method can distinguish between a situation where a pipe burst occurs in the pipe network and a situation where dirty data appears at all monitoring points.
[0112] Figure 19 shows the time series anomaly detection results between monitoring points, and considers four situations: (1) As shown in Figure 19(a), a pipe burst occurs in the piping network, and the monitoring data of monitoring points 1, 2, and 3 all show abnormal values; (2) As shown in Figure 19(b), the monitoring data of monitoring point 1 all shows abnormal values, and the monitoring data of monitoring points 2 and 3 all show normal values; (3) As shown in Figure 19(c), the monitoring data of monitoring points 1 and 2 all show abnormal values, and the monitoring data of monitoring point 3 all show normal values; (4) As shown in Figure 19(d), the monitoring data of monitoring points 1, 2, and 3 all show abnormal values. If the time series detection results of five consecutive times are all abnormal, an abnormality warning is issued: (1) The distance (e.g., d) between the monitoring data time series of two monitoring points is too large. m (S1, S2) and d mIf all of the (S1, S3) are abnormal, the result "Two abnormalities" is output and shown as a point in the orange area. (2) The distance (e.g., d m (S1, S2) or d m If (S1, S3) is abnormal and the distance between the monitoring data time series of one other monitoring point is normal, the result "Single Abnormal" is output and indicated by a dot in the yellow area; (3) If the distances between the monitoring data time series of two monitoring points are all normal, it is determined that there is no abnormal event, and the result "Normal" is output and indicated by a dot in the green area. As shown in Figure 19, the detection rate for pipe ruptures in the pipe network based on the time series between monitoring points is low, but the detection rate for dirty data at the monitoring points is high. This may be because a pipe rupture in the pipe network causes a drop in the monitoring data at each monitoring point, and the change patterns of the monitoring data at each monitoring point are similar, resulting in small changes in the time series distance between monitoring points.
[0113] Figure 20 shows the time-series anomaly detection results for the monitoring point itself, considering four situations: (1) As shown in Figure 20(a), a pipe burst occurs in the piping network, and the monitoring data for monitoring points 1, 2, and 3 all show abnormal values; (2) As shown in Figure 20(b), the monitoring data for monitoring point 1 all shows abnormal values, and the monitoring data for monitoring points 2 and 3 all show normal values; (3) As shown in Figure 20(c), the monitoring data for monitoring points 1 and 2 all show abnormal values, and the monitoring data for monitoring point 3 all show normal values; and (4) As shown in Figure 20(d), the monitoring data for monitoring points 1, 2, and 3 all show abnormal values. If all five consecutive time-series anomaly detection results are abnormal, an abnormality warning is issued, indicated by a dot in the yellow area. Conversely, a "normal" result is output, indicated by a dot in the green area. As shown in Figure 20, based on the time series of the monitoring points themselves, various abnormal scenes can be effectively detected, and at the same time, the scenes in which dirty data appears at a single monitoring point and at some monitoring points can be distinguished, but the scenes in which a pipe burst occurs in the pipe network and the scenes in which dirty data appears at all monitoring points cannot be accurately distinguished.
[0114] Detecting situations where dirty data appears at monitoring points. Except for scenes 4 and 8, σ2 is approximately 98% for all scenes. As shown in Figure 21, the provided method can accurately detect abnormal values in monitoring data and effectively identify various abnormal scenes. Here, scene 4, where dirty data appears at three monitoring points, is identified as a scene where a pipe burst occurs in the pipe network. When identifying abnormal scenes where a pipe burst occurs in a pipe network or dirty data appears at monitoring points, it is usually easy to identify abnormal scenes where dirty data appears at one or some monitoring points. However, if abnormalities appear in the monitoring data of all monitoring points (all monitoring data is clearly degraded), this may be identified as a scene where a pipe burst occurs in the pipe network. In consideration of water supply safety, a pipe burst warning should be issued and the pipe network suspected of a pipe burst should be investigated.
[0115] Figure 22 shows the anomaly detection results for various scenes after removing single-point anomaly detection results. As shown in Figure 22, removing single-point anomaly detection results does not significantly affect the detection results for scenes 2-9, i.e., the impact on abnormal scenes where dirty data appears at monitoring points is small. However, the detection results for pipe bursts in a pipe network are significantly affected, mainly reflected in the anomaly detection rate and anomaly identification rate. After removing single-point anomaly detection results, some abnormal monitoring data caused by a pipe burst in the pipe network is not detected, resulting in a decrease in the anomaly detection rate. At the same time, when identifying abnormal scenes where a pipe burst occurs in a pipe network or dirty data appears at a monitoring point, the pipe burst in the pipe network at a certain time is identified as a situation where dirty data appears at the monitoring point, resulting in a decrease in the anomaly identification rate. This is because, after a certain period of time after a pipe burst, all monitoring data at each monitoring point drops, resulting in no abnormalities being detected in the time series of the monitoring point itself and the time series between monitoring points. Single-point anomaly detection can precisely detect the drop in pressure monitoring data after the pipe burst.
[0116] Figure 23 shows the detection results of various abnormal scenes after removing the single-point value qualitative detection results. As shown in Figure 23, the anomaly detection rates for various scenes all reach 100%, while the detection accuracy rate and anomaly identification rate all decrease significantly. This indicates that although there are false alarms and false detections of various abnormal scenes, all abnormal scenes are detected. Obviously, the single-point value qualitative detection results are mainly used to remove false alarms. When the single-point value qualitative detection results are integrated, the detection accuracy rate and anomaly identification rate for various scenes are both high, while the anomaly detection rate is less than 100%. Under normal operating conditions, fluctuations in water demand due to weather or holidays cause changes in the monitoring data at each monitoring point in the pipe network. These changes are detected simply by detecting time series anomalies at the monitoring point itself and time series anomalies between monitoring points. Through the single-point value qualitative detection results, abnormal alarm situations in the monitoring data caused by fluctuations in water demand can be removed. When the single-point value qualitative detection results are removed, the false alarm rate for pipe burst detection increases.
[0117] As shown in Figure 24, after removing the time series anomaly detection results between monitoring points, the detection results for a pipe burst in the pipeline network (Scenario 1) are not significantly affected, but the abnormal scenarios where dirty data appears at the monitoring points are significantly affected, resulting in a significant decrease in the anomaly detection rate for all eight scenarios. This indicates that multiple anomalies are not detected in the abnormal scenarios where dirty data appears at the monitoring points. After dirty data appears at a monitoring point, the shape of the time series of the monitoring data for each monitoring point or some monitoring points changes. If the shape of the time series of the monitoring data for a monitoring point matches the shape of the time series of its historical monitoring data, anomalies will not be detected in the time series of the monitoring point itself. If the changes in these monitoring values are small, it is difficult to detect anomalies through single-point anomaly detection. Clearly, these anomalies can be detected through time series anomaly detection between monitoring points alone.
[0118] Figure 25 shows the detection results for various abnormal scenes after the time series anomaly detection results of the monitoring point itself are removed. As shown in Figure 25, the anomaly detection rate and anomaly identification rate for various abnormal scenes all decrease. Obviously, after dirty data appears at a monitoring point, the shape of its own time series changes. When an abnormal value appears in some monitoring data, the magnitude of the monitoring value does not change significantly, so it cannot be detected by single-point anomalies or time series anomaly detection methods between monitoring points, resulting in a decrease in the anomaly detection rate. In addition, the lack of time series warning situations for each monitoring point itself also results in a decrease in the identification rate for various abnormal scenes.
[0119] (Addendum) (Appendix 1) (1) performing single-point anomaly detection on real-time monitoring data of a water supply network to distinguish between normal and abnormal working conditions of the water supply network; (2) performing time series anomaly detection on real-time monitoring data of the water supply pipe network to distinguish between normal and abnormal working conditions of the water supply pipe network; and (3) integrating the results of single-point and time-series anomaly detection in the water supply pipe network to accurately detect and identify pipe bursts in the water supply pipe network and failures in the monitoring system.
[0120] (Appendix 2) Specifically, the step (1) includes: (1.1) preparing historical and real-time single-point monitoring data at each monitoring point in a water distribution network so as to perform single-point outlier detection on the real-time monitoring data; (1.2) performing single-point anomaly detection on real-time monitoring data of a water supply network using an unsupervised stacking integration algorithm to identify single-point anomalies in the real-time monitoring data; and (1.3) utilizing statistical theory to perform qualitative detection on single values of the water supply network monitoring data and identify single-point anomalies in the real-time monitoring data.
[0121] (Appendix 3) Specifically, the step (2) includes: A step (2.1) of preparing historical and real-time monitoring data time series at each monitoring point in the water supply network so as to perform time series anomaly detection on the real-time monitoring data; (2.2) performing a dissimilarity analysis on the monitoring data time series at different monitoring points to determine abnormal time series between the monitoring points; and (2.3) performing a dissimilarity analysis on the monitoring data time series at the same monitoring point to determine the abnormal time series of the monitoring point itself.
[0122] (Appendix 4) Specifically, the step (3) includes: (3.1) using a single-point outlier detection method to perform single-point outlier detection on various abnormal scenes in the water supply network; (3.2) performing time series anomaly detection for various abnormal situations in the water supply pipe network using a time series anomaly detection method; (3.3) Integrating the single point and time series anomaly detection results to detect and identify various abnormal situations in the water supply pipe network; and (3.4) using the pipe burst detection method to detect and identify various abnormal situations in the water supply pipe network and evaluate the pipe burst detection performance of the provided method.
[0123] (Appendix 5) Specifically, the step (3.1) is The method for detecting and evaluating pipe ruptures in a water supply network based on single-point and time-series anomaly detection, as described in Appendix 4, is characterized in that when using the single-point anomaly detection method to detect single-point anomalies for various abnormal situations in a water supply network, the method considers the following four situations: (1) a pipe rupture occurs in the network; (2) dirty data appears at a single monitoring point; (3) dirty data appears at some monitoring points; and (4) dirty data appears at all monitoring points. Each situation contains 25 sets of abnormal monitoring data. The method performs single-point anomaly detection from the time the abnormal value begins, and determines whether to issue an alert based on the anomaly detection results at five consecutive times.
[0124] (Appendix 6) Specifically, the step (3.2) is The method for detecting and evaluating performance of a water supply network pipe rupture based on single-point and time-series anomaly detection, as described in Appendix 4, is characterized in that when using a time-series anomaly detection method to detect time-series anomalies in various abnormal situations in a water supply network, the method considers the following four situations: (1) a pipe rupture occurs in the network; (2) dirty data appears at a single monitoring point; (3) dirty data appears at some monitoring points; and (4) dirty data appears at all monitoring points. Each situation contains 25 sets of abnormal monitoring data. The method performs time-series anomaly detection between monitoring points from the time when the abnormal value begins, and determines whether to issue an alert based on the time-series anomaly detection results at five consecutive times.
[0125] (Appendix 7) Specifically, the step (3.3) is Detecting whether a single point or time series anomaly exists in the monitoring data (3.3.1); a step (3.3.2) of identifying a situation in which an anomaly appears in the monitoring data at some monitoring points; A step (3.3.3) of identifying burst pipes and situations where dirty data appears at all monitoring points; and (3.3.4) evaluating and analyzing the performance of the method for detecting pipe bursts in a water supply network based on single-point and time-series anomaly detection.
[0126] (Appendix 8) Specifically, first, distinguish between the normal working conditions of the piping network and the circumstances under which abnormalities appear in the monitoring data; if the existence of a single-point abnormal value is detected for the single-point monitoring data based on an unsupervised stacking integration algorithm or statistical theory, mark the time when the single-point abnormal value appears, and then perform single-point abnormal value detection for the monitoring data at the next time; if an abnormal time series is found for the time-series monitoring data based on time-series abnormality detection between monitoring points or time-series abnormality detection for the monitoring point itself, mark the time when the abnormal time series appears, and then perform time-series abnormality detection for the monitoring data at the next time (step 3.3.1); Specifically, when an anomaly appears in the single point and time series of the monitoring data at some monitoring points, and no anomaly appears in the single point and time series of the monitoring data at the remaining monitoring points, it is determined that dirty data appears at some monitoring points (3.3.2); Specifically, when abnormalities are detected in both the single point and time series of the monitoring data at all monitoring points, the situation of a pipe burst occurring in the piping network and the situation of dirty data appearing at all monitoring points are identified through the increase and decrease of the abnormal values. When the single point qualitative detection results of the monitoring data at all monitoring points are all "decreasing", it is determined that a pipe burst has occurred in the piping network. Conversely, when the single point qualitative detection result of a certain monitoring data is "increasing", it is determined that dirty data has appeared at all monitoring points in the piping network. Step (3.3.3) The method for detecting and evaluating pipe bursts in a water supply network based on single-point and time-series anomaly detection and performance evaluation described in Appendix 7, characterized in that evaluating the pipe burst detection performance in the method includes a step (3.3.4) of considering several indicators: (1) detection accuracy rate (σ1), (2) anomaly identification rate (σ2), and (3) anomaly detection rate (σ3).
[0127] (Appendix 9) The detection accuracy rate is shown as follows: σ1=N dn / N n In the formula, N n indicates the total number of abnormal scenes, N dn denotes the number of detected abnormal scenes, The anomaly identification rate is shown as follows: σ2=N in / N n In the formula, N in denotes the number of correctly identified anomalous scenes, and N n indicates the total number of abnormal scenes, The anomaly detection rate is shown as follows: σ3=t dn / t t In the formula, t dn and t t The method for detecting and evaluating the performance of a water supply network pipe burst based on single-point and time-series anomaly detection described in Appendix 8, characterized in that the above terms respectively indicate the duration due to the abnormal event detection and the actual duration.
[0128] (Appendix 10) Specifically, step (3.4) includes: When detecting and identifying various abnormal situations in the water supply network, removing single-point abnormal value detection results to obtain an abnormality detection rate, detection accuracy rate and abnormality identification rate for various abnormal situations (3.4.1); When detecting and identifying various abnormal situations in the water supply pipe network, removing the single-point qualitative detection results to obtain the abnormality detection rate, detection accuracy rate and abnormality identification rate of various abnormal situations (3.4.2); When detecting and identifying various abnormal situations in the water supply pipe network, removing the time series abnormality detection results between monitoring points to obtain the abnormality detection rate, detection accuracy rate and abnormality identification rate of various abnormal situations (3.4.3); The method for detecting and evaluating pipe bursts in a water supply network based on single-point and time-series anomaly detection and performance evaluation described in Appendix 4, characterized in that it includes a step (3.4.4) of removing the time-series anomaly detection results of the monitoring point itself when detecting and identifying various abnormal situations in the water supply network, thereby obtaining the anomaly detection rate, detection accuracy rate, and anomaly identification rate for various abnormal situations.
Claims
1. (1) performing single-point outlier detection on real-time monitoring data of a water supply network to distinguish between normal and abnormal working conditions of the water supply network; (2) performing time series anomaly detection on real-time monitoring data of the water supply network to distinguish between normal and abnormal working conditions of the water supply network; and (3) integrating the results of single-point and time-series anomaly detection in the water supply network to accurately detect and identify pipe bursts in the water supply network and failures in the monitoring system.
2. Specifically, the step (1) includes: (1.1) preparing historical and real-time single-point monitoring data at each monitoring point in a water distribution network so as to perform single-point outlier detection on the real-time monitoring data; (1.2) performing single-point anomaly detection on real-time monitoring data of a water supply network using an unsupervised stacking integration algorithm to identify single-point anomalies in the real-time monitoring data; and (1.3) utilizing statistical theory to perform qualitative detection on single values of the water supply network monitoring data and identify single-point anomalies in the real-time monitoring data.
3. Specifically, step (2) includes: A step (2.1) of preparing historical and real-time monitoring data time series at each monitoring point in the water supply network so as to perform time series anomaly detection on the real-time monitoring data; (2.2) performing a dissimilarity analysis on the monitoring data time series at different monitoring points to determine abnormal time series between the monitoring points; and (2.3) performing a difference analysis on the monitoring data time series at the same monitoring point to determine the abnormal time series of the monitoring point itself.
4. Specifically, step (3) includes: (3.1) using a single-point outlier detection method to perform single-point outlier detection on various abnormal scenes in the water supply network; (3.2) performing time series anomaly detection for various abnormal situations in the water supply pipe network using a time series anomaly detection method; (3.3) Integrating the single point and time series anomaly detection results to detect and identify various abnormal situations in the water supply network; The method for detecting and evaluating the performance of a water supply network pipe burst based on single point and time series anomaly detection, as described in claim 1, further comprising the step (3.4) of using the pipe burst detection method to detect and identify various abnormal situations in the water supply network and evaluating the pipe burst detection performance of the provided method.
5. Specifically, the step (3.1) is The method for detecting and evaluating the performance of a water supply network pipe rupture based on single-point and time-series anomaly detection, as claimed in claim 4, characterized in that when using the single-point anomaly detection method to perform single-point anomaly detection for various abnormal situations in a water supply network, the following four situations are considered: (1) a pipe rupture occurs in the network; (2) dirty data appears at a single monitoring point; (3) dirty data appears at some monitoring points; and (4) dirty data appears at all monitoring points. Each situation contains 25 sets of abnormal monitoring data. The method includes performing single-point anomaly detection from the time when the abnormal value begins, and determining whether to issue an alert based on the abnormal value detection results at five consecutive times.
6. Specifically, step (3.2) includes: The method for detecting and evaluating the performance of a water supply network pipe rupture based on single-point and time-series anomaly detection, as claimed in claim 4, further comprising: when using the time series anomaly detection method to perform time series anomaly detection for various abnormal situations in a water supply network, the method considers the following four situations: (1) a pipe rupture occurs in the network; (2) dirty data appears at a single monitoring point; (3) dirty data appears at some monitoring points; and (4) dirty data appears at all monitoring points, each of which contains 25 sets of abnormal monitoring data; performing time series anomaly detection between monitoring points from the time when the abnormal value begins; and determining whether to issue an alert based on the time series anomaly detection results at five consecutive times.
7. Specifically, the step (3.3) is Detecting whether a single point or time series anomaly exists in the monitored data (3.3.1); Identifying a situation where abnormalities appear in the monitoring data at some monitoring points (3.3.2); Identifying burst pipes and situations where dirty data appears at all monitoring points (3.3.3); The method for detecting and evaluating the performance of a water supply network pipe burst based on single point and time series anomaly detection, as described in claim 4, further comprising a step (3.3.4) of evaluating and analyzing the performance of the method for detecting pipe burst.
8. Specifically, first, distinguish between the normal working conditions of the piping network and the circumstances in which abnormalities appear in the monitoring data; if the existence of a single-point abnormal value is detected for the single-point monitoring data based on an unsupervised stacking integration algorithm or statistical theory, mark the time when the single-point abnormal value appears, and then perform single-point abnormal value detection for the monitoring data at the next time; if an abnormal time series is found for the time-series monitoring data based on time-series abnormality detection between monitoring points or time-series abnormality detection for the monitoring point itself, mark the time when the abnormal time series appears, and then perform time-series abnormality detection for the monitoring data at the next time (3.3.1); Specifically, when an anomaly appears in the single point and time series of the monitoring data at some monitoring points, and no anomaly appears in the single point and time series of the monitoring data at the remaining monitoring points, it is determined that dirty data appears at some monitoring points (3.3.2); Specifically, when abnormalities are detected in both the single point and time series of the monitoring data at all monitoring points, the situation of a pipe burst occurring in the piping network and the situation of dirty data appearing at all monitoring points are identified through the increase and decrease of the abnormal values. When the single point qualitative detection results of the monitoring data at all monitoring points are all "decreasing", it is determined that a pipe burst has occurred in the piping network. Conversely, when the single point qualitative detection result of a certain monitoring data is "increasing", it is determined that dirty data has appeared at all monitoring points in the piping network. (3.3.3) The pipe burst detection performance of the method is evaluated based on the following criteria: (1) detection accuracy rate (σ 1 ), (2) Anomaly identification rate (σ 2 ), (3) Anomaly detection rate (σ 3 and (3.3.4) considering several indicators, such as:
9. The detection accuracy rate is shown as follows: s 1 =N dn / N n In the formula, N n indicates the total number of abnormal scenes, N dn denotes the number of detected abnormal scenes, The anomaly identification rate is shown as follows: s 2 =N in / N n In the formula, N in denotes the number of correctly identified anomalous scenes, and N n indicates the total number of abnormal scenes, The anomaly detection rate is shown as follows: σ 3 =t dn / t t In the formula, t dn and t The method for detecting and evaluating the performance of a water supply network pipe burst based on single-point and time-series anomaly detection as claimed in claim 8, wherein the time periods indicate the duration due to the detected anomaly and the actual duration, respectively.
10. Specifically, step (3.4) includes: When detecting and identifying various abnormal situations in the water supply network, removing single-point abnormal value detection results to obtain anomaly detection rates, detection accuracy rates and anomaly identification rates for various abnormal situations (3.4.1); (3.4.2) When detecting and identifying various abnormal situations in the water supply network, removing the single-point qualitative detection results to obtain the abnormality detection rate, detection accuracy rate and abnormality identification rate of various abnormal situations; When detecting and identifying various abnormal situations in the water supply network, removing the time series abnormal detection results between monitoring points to obtain the abnormality detection rate, detection accuracy rate and abnormality identification rate of various abnormal situations (3.4.3); The method for detecting and evaluating pipe burst performance in a water supply network based on single-point and time-series anomaly detection, as described in claim 4, further comprising the step (3.4.4) of removing the time-series anomaly detection results of the monitoring point itself when detecting and identifying various abnormal situations in the water supply network, to obtain the anomaly detection rate, detection accuracy rate and anomaly identification rate for various abnormal situations.
Citation Information
Patent Citations
City water supply network burst detection method based on dynamic neural network prediction
CN108167653A
Fire-fighting water pressure anomaly monitoring system and unsupervised anomaly detection method
CN113413568A
System and method for identifying geographical locations potentially affected by anomalies in a water supply network.
JP2014510261A
Abnormality detection device and abnormality detection method
JP2023106472A
Competition-based tool for anomaly detection of business process time series in it environments
US20190228353A1