A detection method for abnormal data in a water supply network
The integration of k-means clustering and box-plot analysis for anomaly detection in water supply networks addresses the inefficiencies of existing methods by utilizing spatial correlations, improving detection accuracy and reducing computational complexity.
Patent Information
- Application Number
- CN202211312033.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-25
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-10-25
AI Technical Summary
The prior art fails to fully utilize the spatial and temporal correlation between monitoring points in the detection of abnormal data of water supply pipeline networks, resulting in complex model construction and low efficiency, and high false positive rate, especially when the data fluctuations increase during the season change period, the fitting performance becomes poor.
The k-means clustering model and box-type graph discrimination method are used to reasonably group monitoring points, and the spatiotemporal correlation between nodes is used to establish a traffic pressure data clustering model for local neighboring nodes, and various critical thresholds are determined to detect abnormal data.
It realizes efficient and accurate detection of abnormal data on flow pressure of water supply pipeline networks, reduces false positive rates, simplifies the model construction process, and ensures the correct analysis of the operating status of the pipeline network.
Smart Images

Figure CN115628776B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of online monitoring of water supply networks, and particularly to a method for detecting abnormal data in a water supply network. Background Art
[0002] With the development of Internet of Things technology, the Supervisory Control and Data Acquisition (SCADA) system for water supply network data collection and monitoring has become increasingly popular. It is very difficult to manually process the massive measured data generated by online monitoring. In recent years, various data-driven methods have emerged, automatically detecting potential outliers therein, and performing a preliminary screening on the massive data, which can greatly reduce the amount of manual identification and processing. In Reference [1], the online monitoring data of a water supply network monitoring station in Beijing for 56 months was segmented in time and season, an autoregressive moving average (ARMA) model was constructed, and the outliers in the artificially simulated sequence were identified through the confidence interval established by this model, so as to realize the self-identification of the data of an independent node itself. In Reference [2], the measured pressure data of a single pressure monitoring point in a certain community was used as sample data, the sample data was denoised through wavelet analysis, 10% reduction of the sample data was used as abnormal data, and the abnormal data was identified based on the local outlier factor algorithm and the K-means algorithm. When detecting outliers in the sample, the sample was processed by taking the difference between adjacent data of the sample to be measured, the abnormal detection results of the sample data before and after processing were compared, and the detection results of abnormal data under different algorithms were compared. The results show that the abnormal detection effect of the processed sample data is better, and the result of the K-means algorithm is relatively better. Reference [3] provides a method for detecting abnormal data in a water supply network based on big data. After preprocessing, normal distribution processing and three times standard deviation (3σ) criterion detection on the big data of water supply network monitoring, the abnormal data of each sensor is determined, and the detection accuracy is high. References [1], [3] and [4] only utilize the time-series data of a single monitoring point, and do not utilize the implicit law of spatio-temporal correlation of many monitoring points in the water supply network, and the false positive rate is relatively high.
[0003] Reference [4] segmented the time series of the online monitoring data of 10 monitoring points in a domestic water supply network over 8 months. Through the analysis of the spatial topological relationship between the monitoring points, the optimal number of monitoring points was determined, and the data of the selected monitoring points were used to construct a support vector regression (SVR) model. After optimizing the model, the confidence interval established by it was used for artificial simulation of outlier identification. The results showed that the SVR model had good fitting performance, and the outlier detection rate was as high as 90%, realizing the interactive identification between the data of the selected monitoring points. Although Reference [4] utilized the SVR model to mine the implicit logical relationship between the monitoring points to predict the change trend of the data of a specific monitoring point, and then gave a prediction confidence interval to identify outliers. However, to ensure the accuracy of the SVR model, the monitoring data was segmented into 96 segments at 15-minute intervals to reduce data fluctuations. In this way, it was necessary to construct 96 SVR models, which was actually cumbersome to process and inefficient. In case of seasonal changes, the data fluctuations in the same period increased, and the fitting performance of the SVR model deteriorated, directly affecting outlier identification.
[0004] Therefore, there is an urgent need for a method for detecting abnormal data in water supply networks that can not only make full use of the spatio-temporal correlation between monitoring points, but also be simple in modeling and calculation.
[0005] References:
[0006] [1] Liu Shuming, Wu Yipeng, Che Han. Quality control of water supply network monitoring data using self-identification [J]. Journal of Tsinghua University (Science and Technology), Vol. 57 No. 9 2017.
[0007] [2] Yang Qihang. Research on the identification of abnormal detection data in water supply networks [D]. Tianjin University of Technology, 2022. DOI: 10.27360 / d.cnki.gtlgy.2022.000124.
[0008] [3] Liu Hang, Liu Chi, Lang Daizhi, Tian Wuping, Han Xing. Method for detecting abnormal data in water supply networks based on big data [P]. Chongqing: CN112612824A, 2021-04-06.
[0009] [4] Liu Shuming, Wu Yipeng, Che Han. Detection of outliers in water supply network data based on interactive identification [J]. Water & Wastewater Engineering, 2015, Vol. 41 No. 11. Summary of the Invention
[0010] Aiming at the above problems, the present invention provides a method for detecting abnormal data in water supply networks, which combines the k-means clustering model and the box-plot outlier detection method. Utilizing the spatio-temporal correlation between the nodes of the water supply network, the flow and pressure data of local neighboring nodes are selected to establish a clustering model, and specific critical thresholds for each category are determined to realize the detection of abnormal flow and pressure data in the water supply network.
[0011] The present invention includes the following steps:
[0012] Step 1, reasonably grouping the monitoring points based on clustering of normal operating conditions.
[0013] The water supply network is composed of numerous nodes and connecting pipes. Water supply companies generally arrange flow and pressure monitoring points at the water demand nodes at the end and the branch nodes of the pipe network to perceive the hydraulic operation state of the pipe network. In view of the close correlation of the hydraulic states of adjacent nodes, reasonably dividing the monitoring points and trying to form a group of the measuring points with strong correlation is conducive to exploring the operating conditions of the involved area.
[0014] For the monitoring points covering the entire pipe network, they are divided into 1 group, 2 groups,..., N / 2 groups (N is the number of measuring points, and if N is odd, round down N / 2) according to the proximity principle. Suppose there are L grouping methods, and the k-means clustering is respectively used to analyze the operating conditions of the measuring point groups under the L grouping methods. Calculate the separation index ε(l)' and the silhouette coefficient S(l)' of the clustering for each grouping method, where l = 1, 2,..., L. The larger the values of the two indexes, the better the clustering effect of the grouping method, that is, the more reasonable the grouping of the monitoring points. Since the trends of the two indexes with the increase of the number of measuring points are different, draw a double-axis display line graph with the separation index ε(l)' and the silhouette coefficient S(l)' as the Y-axis and the number of monitoring points as the X-axis, and select a reasonable grouping method for adjacent monitoring points according to the intersection point of the two index line graphs.
[0015] The specific process of reasonably grouping the monitoring points is as follows:
[0016] First, for the monitoring points covering the entire water supply network, they are divided into 1 group, 2 groups,..., N / 2 groups (N is the number of measuring points, and if N is odd, round N / 2) according to the proximity principle, and there are a total of L grouping methods. For each grouping, for the normal flow and pressure data included in the measuring point group, perform k-means clustering to obtain the clustering result of the measuring point group. When performing k-means clustering, the value of the number of clustering categories k is selected according to the elbow method. The core index of the elbow method is SSE (sum of squared errors), and Equation (1) is the calculation method of SSE:
[0017]
[0018] where k represents the number of clustering categories, E d represents the d-th category, a represents the sample point in E d and c d represents the center of E d . Calculate the SSE (sum of squared errors) under multiple k values. Taking the X-axis as the number of clustering categories k value and the Y-axis as the SSE value, draw a curve graph to obtain an elbow graph, and select the k value corresponding to the elbow as the number of clustering categories.
[0019] The described k-means clustering has the following specific process:
[0020] Suppose a dataset X is given, which contains n sample points, and the dimension of each sample point is m-dimensional, that is: X = {X1, X2, X3, …, X n}.
[0021] (1) First, randomly select k sample points D = {D1, D2, D3, …, D k} in the dataset X as the initial class centers. Calculate the Euclidean distance from each sample point X a to all the clustering centers, and the calculation formula is Equation (2);
[0022]
[0023] X a represents the a-th sample point, a ∈ [1, n], D b represents the b-th clustering center, b ∈ [1, k], X at represents the t-th attribute of the a-th sample point, t ∈ [1, m], D bt represents the t-th attribute of the b-th clustering center.
[0024] (2) Compare the distances from each sample to all the clustering centers, find the class center with the minimum distance, and divide the sample points into the category where the nearest clustering center is located, obtaining k sets of data {F1, F2, F3, …, F k}.
[0025] (3) According to the categories divided in (2), calculate the center point of each category as the new clustering center, and the calculation formula of the clustering center is Equation (3):
[0026]
[0027] C g represents the g-th clustering center, g ∈ [1, k], F g represents the g-th category, N g represents F g contains the number of sample points, X h represents the h-th sample point in F g , h ∈ [1, N g .
[0028] (4) Iterate steps (2) and (3) until the clustering centers no longer change.
[0029] Secondly, calculate the separation index ε of the grouping, and its calculation is as shown in Equation (4)
[0030]
[0031] Among them, D pq represents the Euclidean distance between the clustering centers of class p and class q, which is calculated according to Equation (5). In Equation (5), m represents the dimension of the clustering center point, and C pt and C qt respectively represent the t-th attribute of the p-th class center and the q-th class center.
[0032]
[0033] Among them, d p and d q represent the within-class distance between class p and class q. The within-class distance σ is the standard deviation of the distances from the sample points in each class to the clustering center, which is calculated according to Equation (6). Among them, N represents the total number of distances from the sample points in the same class to the clustering center, x e represents the distance from the sample point e to the class center, and μ represents the average value of all distance values, which is calculated according to Equation (7):
[0034]
[0035]
[0036] Since there are usually multiple classes in each clustering, there is a separation index between two classes. Therefore, the mean value of all separation indexes is taken here as the separation index ε(l)′ of the clustering under the grouping method, and its calculation method is as shown in Equation (8)
[0037]
[0038] Among them, ε o represents the o-th separation index of this grouping, o ∈ [1, U], and U represents that there are U separation indexes in this grouping method.
[0039] Then, calculate the silhouette coefficient S(l)′ of the grouping method. The specific calculation process of the silhouette coefficient is as follows:
[0040] (1) For the sample point w in a certain class, calculate the distances from this point to all other elements in the same class, and then take the average value of the distances, denoted as a(w), which represents the cohesion within the class.
[0041] (2) Take another class outside the class to which the sample point w belongs, calculate the distances from the sample point w to all the sample points in this class, and then find the average value of the distances. Traverse all other classes, find the class with the closest distance, and denote the average value of the distances from the point w to all the sample points in this class as b(w), which represents the separation between classes.
[0042] (3) For the sample point w, its silhouette coefficient calculation formula is as follows
[0043]
[0044] (4) Calculate the silhouette coefficients of all sample points, and find the average value of all silhouette coefficients, which is the silhouette coefficient S(l)′ of this grouping method.
[0045] Finally, obtain the separation index and silhouette coefficient under each grouping method, draw a dual-axis display diagram, and obtain the grouping method with better clustering effect, that is, determine the reasonable grouping of measurement points.
[0046] Step 2: On the basis of the reasonable grouping of monitoring points clustered based on normal working conditions in Step 1, calculate the Euclidean distance from all samples within the group to its class center.
[0047] From Step 1, i measurement point groups under reasonable grouping and j categories of clustering for each measurement point group can be obtained. Then, calculate the Euclidean distance from all samples within each measurement point group after reasonable grouping to its class center according to Equation (2). After obtaining all Euclidean distance values, mark the distance value Dis according to the measurement point group number i and the clustering category j ij , denoted as Disiance.
[0048]
[0049] Among them, Dis ij represents the set of Euclidean distances from all sample points of the j-th category in the i-th measurement point group to its clustering center. Since the clustering categories within each measurement point group may be different, the value range of j is related to the measurement point group number, j ∈ {j1, j2, j3, …, j i}), where j i represents the number of clustering categories in the i-th measurement point group.
[0050] Step 3: Use a box plot to determine the outlier threshold for each category within each measurement point group and test all sample data.
[0051] Analyze the distance data Distance after reasonable grouping using a box plot to obtain the box plot parameters for all categories in each measurement point group. The normal data judgment threshold interval is from the upper limit of the box plot to the lower limit of the box plot. Sample data outside this interval is abnormal data. Mark the upper outlier threshold max ij and the lower outlier threshold min ij , denoted as [min, max].
[0052]
[0053] Among them, max ij and min ijrespectively represent the upper and lower limits of the outlier judgment threshold for the j-th category of the i-th measurement point group. The value of j is related to the group number of the measurement point group, and j ∈ {j1, j2, j3, …, j i}, where j i represents the number of clustering categories of the i-th measurement point group.
[0054] The calculation process of the box plot parameters is as follows:
[0055] (1) Calculate the upper quartile Q3 and lower quartile Q1 of the distance data. For example, if there are v distance data, sort these v data from smallest to largest. Q3 is the number at the T1-th position, and Q1 is the number at the T2-th position. The calculation formulas for T1 and T2 are as shown in Equations (11) and (12)
[0056]
[0057]
[0058] (2) Calculate the interquartile range IQR, as shown in Equation (13)
[0059] IQR = Q3 - Q1 (13)
[0060] (3) Calculate the upper limit Max and lower limit Min. The calculation methods are as shown in Equations (14) and (15) respectively
[0061] Max = Q1 + W * IQR (14)
[0062] Min = Q3 - W * IQR (15)
[0063] where W is the weight coefficient in front of the interquartile range IQR.
[0064] According to the grouping situation of the measurement point group, all sample data are divided into the corresponding groups according to the measurement points, and then the distances from all sample data to the class center are calculated according to Equation (2). According to the distance, the sample data are divided into the category with the closest distance, and the distance value dis ij is marked according to the measurement point group number i and the clustering category j, and is denoted as distance.
[0065]
[0066] where dis ij represents the set of distances from the sample data of the j-th category of the i-th measurement point group to the class center. Here, the value of j corresponds to the number of clustering categories after reasonable grouping, and j ∈ {j1, j2, j3, …, j i}, where j i represents the number of clustering categories of the i-th measurement point group.
[0067] Let disij The distance value in the corresponding threshold interval [max ij ,min ij ]Compared with the above, anything outside the range is abnormal.
[0068] Then, the confusion matrix of each measurement point group is obtained according to the test results, and the detection accuracy is calculated. The confusion matrix of a single measurement point group is shown in Table 1.
[0069] Table 1 Confusion matrix
[0070] Classification Actually abnormal Actually normal Judged as abnormal TP FP Judged as normal FN TN
[0071] Data detection accuracy of the measurement point group i Calculate according to formula (16):
[0072]
[0073] Among them, Accurary i represents the accuracy of the i-th group of measuring points, where TP represents the number of abnormal data detected correctly, FP represents the number of normal data detected as abnormal data, FN represents the number of abnormal data detected as normal data, and TN represents the number of normal data detected correctly.
[0074] If the accuracy is Accurary i If it does not reach 95%, the threshold interval of each category of the measurement point group needs to be adjusted. The weight coefficient W before IQR in formula (14) and formula (15) can be slightly adjusted. The default value of this coefficient is 1.5. After making a small adjustment, all sample data are tested. If the accuracy is i If the standard is still not met, continue to make adjustments until it is met.
[0075] Step 4 Actual abnormal data detection
[0076] The node flow pressure data currently sampled at the monitoring point is:
[0077] (1) Group by measuring points and construct the current samples for each measuring point group.
[0078] (2) For each measurement point group, the current sample is calculated to the center distance of each class in the group (referred to as group r), and is classified into a class s with the closest distance in the group, and the distance within the class dis′ is recorded. rs .
[0079] (3) The intra-class distance dis′ of the current sample of each measurement point group rs The distance value in the class and the threshold interval [max rs ,min rs ] comparison, if it is within the interval, it is normal, otherwise it is abnormal.
[0080] All the current samples in each group have been completely detected, and the detection of differences in the current sampling data is completed.
[0081] Advantages of the present invention: The method of the present invention adopts the k-means clustering and box plot outlier detection method, makes full use of the spatio-temporal correlation between nodes, establishes a clustering model for local nodes, and can accurately identify abnormal data in the detection of water supply networks, providing a guarantee for the correct analysis of the operation status of the pipe network. Description of the Drawings
[0082] Figure 1 Flow chart of the method of the present invention;
[0083] Figure 2 Elbow diagrams under different grouping methods;
[0084] Figure 3 Line charts of separation indices and silhouette coefficients under different grouping methods. Detailed Embodiment
[0085] Select 1955 flow and pressure data of the JS water conveyance project in July 2021 to illustrate the detection of abnormal data in the embodiment of the present invention. Among the 1955 data, there are 257 abnormal data and 1698 normal data, and part of the data on August 6, 2021 is used to verify the actual detection of abnormal data. The specific process of the present invention is as Figure 1 shown, and the specific steps are as follows:
[0086] Step 1: Reasonable grouping of monitoring points based on clustering under normal operating conditions.
[0087] According to the distribution of measuring points in the water supply network, all monitoring points are grouped according to the proximity principle. There are 21 monitoring points in this water diversion project, and the monitoring points are divided into 1 group, 4 groups, 5 groups, 7 groups and 10 groups, as shown in Table 2:
[0088] Table 2 Grouping situation of monitoring points
[0089]
[0090]
[0091] Due to the influence of the number of monitoring points, the grouping cannot exactly make the number of monitoring points in each group the same. Therefore, preliminary clustering is performed on the data corresponding to the number of monitoring points that account for the majority in each group during grouping, and the elbow diagram is used as the clustering basis. The number of monitoring points that account for the majority in each of the 5 grouping methods (number of groups 1, 4, 5, 7, 10) is 21, 6, 4, 3, 2, and the corresponding elbow diagrams are as Figure 2 shown. Figure 2The horizontal axis is the K variable, and the vertical axis is the SSE variable. As K increases, the SSE initially decreases rapidly and then more slowly. The SSE curve resembles an elbow plot. When the decrease slows down, it indicates that the effect of decreasing SSE by increasing the K value is no longer significant. Generally, it is appropriate to select the K value at the moment when the SSE decrease significantly slows down as the number of clusters. Observe in this way Figure 2 the elbow plots in, and the selected K values for grouping methods 1 to 5 are 12, 4, 4, 4, and 4 in sequence. In the elbow plot of grouping method 1, it flattens out after K = 12, but there is a significant decrease in SSE before K = 12. Therefore, 12 is selected as the clustering K value for grouping method 1.
[0092] Calculate the separation index ε(l)′ and the silhouette coefficient S(l) for the 5 groupings, as shown in Table 3.
[0093] Table 3 Separation index ε(l)′ and silhouette coefficient S(l)′
[0094] Grouping method 1 2 3 4 5 Number of groups 1 4 5 7 10 Separation index ε′ 5.08 4.61 4.55 4.44 4.23 Silhouette coefficient S(i) 0.287 0.391 0.418 0.461 0.466
[0095] Draw a double-axis display line chart of the corresponding separation index and silhouette coefficient, as shown in Figure 3 shown. Based on its intersection point, a more reasonable grouping is obtained as grouping method 4, that is, the monitoring points are divided into 7 groups, with each group containing 3 monitoring points.
[0096] Step 2 On the basis of the reasonable grouping of the monitoring points clustered based on the normal working conditions in step 1, calculate the Euclidean distance from all samples within the group to its class center.
[0097] Taking the 5th and 7th measuring point groups in grouping method 4 as an example, calculate the Euclidean distance from all samples within the group to its class center, calculated according to formula (2). Mark the distance value set Dis ij , denoted as Distance.
[0098] Distance = {{Dis 51 , Dis 52 , Dis 53 , Dis 54 , Dis 55}}, {Dis 71 , Dis 72 , Dis 73 , Dis 74}}
[0099] Step 3 Use a box plot to determine the outlier threshold for each class within each measuring point group and test all sample data.
[0100] Use the box plot rule to analyze each class of distance data set Dis in the distance value set Distance of the 5th and 7th groups ij, the discrimination threshold interval of the distance values in each category is obtained as [max ij , min ij , as shown in Table 4. Since there is no negative value for the distance, the lower limit of the negative value is uniformly set to 0.
[0101] Table 4 Discrimination Thresholds for Each Category in Group 5 and Group 7
[0102]
[0103] All sample data are assigned to each measuring point group according to grouping method 4. Calculate the distance from each sample data to the center of all categories in the measuring point group to which it belongs. According to the distance, the sample data is attributed to the category with the closest distance, and the distance value set dis is marked according to group i and category j ij , and distance is obtained.
[0104] distance = {{dis 51 , dis 52 , dis 53 , dis 54 , dis 55}, {dis 71 , dis 72 , dis 73 , dis 74}}
[0105] Then, compare the distance values of all sample data to the center of the category to which they belong with the discrimination threshold interval of the corresponding category. If it is outside the interval, it is determined as abnormal data; if it is within the interval, it is determined as normal data. Finally, the confusion matrix of all sample data detection is obtained, and the results are shown in Table 5.
[0106] Table 5 Confusion Matrix for Sample Data Inspection in Group 5 and Group 7
[0107]
[0108]
[0109] Calculate the accuracies Accurary5 and Accurary7 of data detection, and the results are shown in Table 6.
[0110] Table 6 Detection Accuracy of Abnormal Data
[0111] Group Group 5 Group 7 Accuracy rate 98% 99.2%
[0112] The data detection accuracies of the 5th and 7th measuring point groups both meet the standards, and there is no need to adjust the parameter W in equations (14) and (15).
[0113] Step 4 Actual abnormal data detection.
[0114] Use the outlier threshold intervals obtained in the first 3 steps (as shown in Table 4) to detect some data on August 6. Here, take the time period from 15:55 on August 6 to 16:50 on August 6 (12 sampling data) and the 5th and 7th groups of measurement point groups as examples for illustration:
[0115] (1) Group by measurement points to construct the current samples of the 5th and 7th groups.
[0116] (2) For the current samples of the 5th and 7th groups, calculate the distances to the centers of each category within the group, and classify them into the category j with the closest distance within the group. The 12 data of the current sample of the 5th group belong to the categories {4, 4, 4, 4, 1, 4, 3, 2, 2, 2, 2, 2} respectively, and the 12 data of the current sample of the 7th group belong to the categories {4, 2, 2, 2, 2, 2, 2, 2, 4, 2, 4, 2} respectively.
[0117] Record the within-class distance of the sample points in the first category of the 5th group as dis′ 51 and the within-class distance of the sample points in the second category as dis′ 52 and the within-class distance of the sample points in the third category as dis′ 53 and the within-class distance of the sample points in the fourth category as dis′ 54 ; where dis′ 51 ={51.69}, dis′ 52 ={62.16, 70.18, 65.95, 65.41, 65.07}, dis′ 53 ={97.23}, dis′ 54 ={45.27, 45.56, 45.58, 45.15}; Record the within-class distance of the sample points in the second category of the 7th group as dis′ 72 and the within-class distance of the sample points in the fourth category as dis′ 74 . Where dis′ 72 ={208.06, 207.93, 208.51, 208.28, 208.30, 208.04, 207.86, 208.16, 173.87}, dis′ 74 ={53.25, 53.48, 53.73}.
[0118] (3) For the current samples of the 5th and 7th groups, the distance values in their within-class distance dis′ ij and the outlier threshold interval [min ij , max ijCompare. If it is within the interval, it is normal; otherwise, it is abnormal. After comparison, it is found that the sample points of the 5th group classified into the 1st, 2nd, 3rd, and 4th categories do not exceed the upper limit threshold or are lower than the lower limit threshold, and are judged as normal; the distances from the sample points of the 7th group classified into the 2nd category to the class center all exceed the upper limit threshold and are judged as abnormal, and the distances from the sample points classified into the 4th category to the class center do not exceed the upper limit threshold or are lower than the lower limit threshold, and are judged as normal.
[0119] After investigation, the detected abnormal data conforms to the actual situation. As shown in Table 7, good results of detecting abnormal data are obtained.
[0120] Table 7 Detection Results of Actual Abnormal Data
[0121] Group 5 Actually abnormal Actually normal Predicted as abnormal 0 0 Predicted as normal 0 12 Group 7 Actually abnormal Actually normal Predicted as abnormal 8 0 Predicted as normal 0 4
Claims
1. A method for detecting abnormal data in a water supply network, characterized in that It includes the following steps: Step 1, reasonably grouping the monitoring points based on clustering under normal working conditions; S1.1, for the monitoring points covering the entire water supply network, divide them into Group 1, Group 2, …, Group N / 2 according to the proximity principle, where N is the number of measurement points. If N is odd, round down N / 2. There are L grouping methods in total; Under each grouping, perform k-means clustering on the normal flow and pressure data included in the measurement point group to obtain the clustering result of the measurement point group; When performing k-means clustering, the value of the number of clustering categories k is selected according to the elbow method. The core index of the elbow method is the sum of squared errors SSE; Calculate the SSE under multiple k values. Take the X-axis as the value of the number of clustering categories k and the Y-axis as the SSE value. Plot a curve graph to obtain an elbow graph, and select the k value corresponding to the elbow as the number of clustering categories; S1.2, calculate the separation index ε of the grouping, and its calculation is as shown in Equation (4) Among them, D pq represents the Euclidean distance between the clustering centers of class p and class q, which is calculated according to Equation (5), d p and d p represent the within-class distance of class p and class q; where m represents the dimension of the clustering center point, C pt and C qt represent the t-th attribute of the p-th class center and the q-th class center respectively; The within-class distance σ is the standard deviation of the distances from the sample points in each class to the clustering center, and is calculated according to Equation (6): Among them, N represents the total number of distances from sample points to the cluster center in the same category, and x e represents the distance from the sample point e to the center of this category, and μ represents the average value of all distance values, which is calculated according to Equation (7): Take the mean value of all separation indices as the separation index ε(l)′ of the clustering under the grouping method, and its calculation method is as shown in Equation (8) where ε o represents the o-th separability index of the group, o ∈ [1, U], and U represents that there are U separability indices in this grouping method; S1.3, calculate the silhouette coefficient S(l)′ of the grouping method. The specific calculation process of the silhouette coefficient is as follows: S1.3.1, for a sample point w in a certain category, calculate the distances from this point to all other elements in the same category, and then take the average value of the distances, denoted as a(w), representing the cohesion within the class; S1.3.2, select another class outside the class to which the sample point w belongs, calculate the distances from the sample point w to all sample points in this class, and then find the average value of the distances. Traverse all other classes, find the closest class, and denote the average value of the distances from the sample point w to all sample points in this class as b(w), representing the separation between classes; S1.3.3, for the sample point w, its silhouette coefficient calculation formula is as follows S1.3.4, Calculate the silhouette coefficients of all sample points and find the average of all silhouette coefficients, which is the silhouette coefficient S(l) of this grouping method ′ ; S1.4, obtain the separation indices and silhouette coefficients under each grouping method, draw a dual-axis display line graph with the separation index and silhouette coefficient as the Y-axis and the number of monitoring points as the X-axis, and select a reasonable grouping method for adjacent monitoring points according to the intersection of the line graphs of the two indices to determine the reasonable grouping of the measurement points; Step 2, based on the reasonable grouping of the monitoring points clustered under normal working conditions in Step 1, calculate the Euclidean distance values from all samples within the group to their class centers; After obtaining all the Euclidean distance values, label the Euclidean distance value Dis according to the measurement point group number i and the clustering category j ij , denoted as Distance; Step 3, use a box plot to determine the outlier threshold for each class within each measurement point group and test all sample data; Analyze the Euclidean distance data Distance after reasonable grouping using a box plot to obtain the box plot parameters for all classes in each measurement point group. The normal data judgment threshold interval is from the upper limit to the lower limit of the box plot, denoted as [min, max]; According to the grouping of measurement point groups, all sample data are divided into corresponding groups according to the measurement points, and then the Euclidean distances from all sample data to the center of their respective classes are calculated; according to the proximity of the Euclidean distances, the sample data are divided into the class with the closest Euclidean distance, and the Euclidean distance value dis is marked according to the measurement point group number i and the clustering class j ij , denoted as distance; Compare the Euclidean distance value in dis ij with the corresponding outlier threshold interval [max ij , min ij . If it is outside the interval, it is an outlier; Then obtain a confusion matrix according to the test results and calculate the detection accuracy; Step 4, actual abnormal data detection; For the node flow and pressure data currently sampled at the monitoring points; S4.1, group according to the measurement points and construct the current samples for each measurement point group; S4.
2. For the current samples of each measurement point group, calculate the distances to the centers of each category within this group, and classify them into a certain category s with the closest distance within this group, and record the within-class distance dis′ rs ; S4.3, for the current samples of each measurement point group, the within-class distance dis' rs in the distance value is compared with the outlier threshold interval [max ij , min ij of the belonging class. If it is within the interval, it is normal; otherwise, it is abnormal. After all the current samples of each group have been detected, the detection of the current sampled data is completed.
2. The method for detecting abnormal data of a water supply pipe network according to claim 1, characterized in that: The specific process of the k-means clustering described in Step 1 is as follows: Given a dataset X that contains n sample points, and the dimension of each sample point is m-dimensional, i.e., X = {X1, X2, X3, …, X n}; S1.1.
1. First, randomly select k sample points D = {D1, D2, D3, …, D k} in the data set X as the initial class centers, and calculate the Euclidean distance from each sample point X i to all the clustering centers. The calculation formula is shown in Equation (2); X a represents the a-th sample point, where a ∈ [1, n], D b represents the b-th cluster center, where b ∈ [1, k], X at represents the t-th attribute of the a-th sample point, where t ∈ [1, m], D bt represents the t-th attribute of the b-th cluster center; S1.1.
2. Compare the distances of each sample to each cluster center, find the cluster center with the minimum distance, and assign the sample point to the category where the nearest cluster center is located, obtaining data sets for k categories {F1, F2, F3, …, F k}; S1.1.
3. Calculate the center point of each category according to the categories divided in S1.1.2 as the new clustering center. The calculation formula for the clustering center is formula (3): Among them, C g represents the g-th clustering center, where g ∈ [1, k], F g represents the g-th category, N g represents the number of sample points included in F g , X h represents the h-th sample point in F g , where h ∈ [1, N g ; S1.1.
4. Iterate steps S1.1.2 and S1.1.3 until the clustering center no longer changes.
3. The abnormal data detection method for a water supply pipe network according to claim 1, wherein: The calculation process of the box plot parameters described in step 3 is as follows: S3.
1. Calculate the upper quartile Q3 and the lower quartile Q1 of the distance data. There are v distance data. Sort these v data from smallest to largest. Q3 is the number at the T1-th position, and Q1 is the number at the T2-th position. The calculation formulas for T1 and T2 are as shown in formulas (11) and (12). S3.
2. Calculate the interquartile range IQR, as shown in formula (13). IQR = Q3 - Q1 (13) S3.
3. Calculate the upper limit Max and the lower limit Min. The calculation methods are as shown in formulas (14) and (15) respectively. Max = Q1 + W * IQR (14) Min = Q3 - W * IQR (15) Where W is the weight coefficient of the interquartile range IQR.
4. A method for detecting abnormal data in a water supply pipe network according to claim 3, characterized in that: If the accuracy rate described in step 3 does not reach 95%, adjust the discrimination threshold interval for each category of the measurement point group, and slightly adjust the weight coefficient W in front of IQR in formulas (14) and (15). After making a small adjustment, conduct a test on all sample data. If the accuracy rate still does not meet the standard, continue to adjust until it meets the standard.
5. A method for detecting abnormal data in a water supply pipe network according to claim 4, characterized in that: The default value of the weight coefficient W is 1.5.
Citation Information
Patent Citations
Water supply pipe network pipe burst and leakage preliminary positioning method
CN109869638A
Method and system for determining fault characteristics of active power distribution network
CN110750524A