Power distribution network anomaly detection and classification method, device and equipment and storage medium

Through DBSCAN and K-Means clustering algorithms, weighted least squares method and integration method, the efficiency and robustness of distribution network anomaly detection in a high dynamic environment are solved, and high-precision and efficient abnormality detection and classification are achieved.

CN120067764APending Publication Date: 2025-05-30STATE GRID JIANGSU ELECTRIC POWER CO LTD RESEARCH INSTITUTE +4
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510226264.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing distribution network abnormality detection methods have low detection efficiency in high dynamic environments, insufficient robustness and accuracy, making it difficult to adapt to complex distribution system environments.

Method used

The distribution network data is clustered by DBSCAN and K-Means clustering algorithms, and state estimation is performed by weighted least squares method, and the detection results of multiple independent anomaly detectors are integrated through an integrated method to perform event classification.

Benefits of technology

It improves the accuracy and efficiency of abnormal detection of distribution networks, enhances robustness, and can accurately detect and classify abnormal events in high noise and dynamic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067764A_ABST
    Figure CN120067764A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of power systems, in particular to a power distribution network anomaly detection and classification method, device and equipment and a storage medium, and the method comprises the following steps: S100, anomaly detection: employing DBSCAN and K-Means clustering algorithms to cluster power distribution network data, and identifying abnormal data points; s200, state estimation: estimating the voltage amplitude of the power distribution network by adopting a weighted least square method, and improving the estimation precision in combination with an abnormal data point compensation mechanism; and S300, integration and event classification: integrating the detection results of the plurality of independent anomaly detectors through an integration method, and classifying the detected events. According to the method, through the synergistic effect of multiple algorithms, clustering, weighted least square method estimation and event classification are combined, and the accuracy and efficiency of power distribution network anomaly detection are improved. According to the method, the weighted least square method is combined with anomaly detection, and the detection accuracy is improved by reasonably distributing weights for the environment with high noise level and measurement abnormal values.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power systems, and particularly to a method, device, equipment and storage medium for abnormal detection and classification of a distribution network. Background Art

[0002] With the continuous expansion of the scale and the increase in complexity of power systems, enhancing network monitoring and controllability has become an important goal in the power industry. To this end, the power industry has developed various sensor technologies, such as power quality meters, disturbance recorders, and high-resolution power flow meters, to capture power event disturbances and their characteristics. However, these traditional technologies have certain limitations. For example, they lack an accurate timestamp function and cannot report the actual root cause of events in a timely manner, resulting in low efficiency in fault detection and diagnosis. To solve this problem, distribution phasor measurement units (D-PMUs) have been widely deployed to collect and analyze phasor data with timestamps, providing important support for the real-time monitoring of power systems.

[0003] Nevertheless, existing abnormal detection methods still face many challenges. First, existing methods usually rely on the accuracy of training datasets, system information, and network models, and are tested on simulated event datasets, making it difficult to ensure the detection efficiency under different grid conditions. Second, although existing unsupervised techniques perform abnormal detection through statistical feature extraction and clustering-based methods, most detectors rely on specific statistical features, resulting in insufficient robustness and accuracy of detection results. In addition, the distribution system has dynamic behaviors (such as frequent switching and topological changes), further increasing the difficulty of abnormal detection and classification.

[0004] To improve detection accuracy and robustness, an ensemble method (EA) has been introduced to optimize performance by integrating the detection results of different statistical features. However, how to accurately detect and classify abnormalities in a highly dynamic environment, while reducing the dependence on hyperparameter tuning and effectively handling the missing connection problems in the base clustering results, remains the focus and difficulty of current distribution network fault detection. Therefore, there is an urgent need to develop an abnormal detection and classification method that can adapt to complex distribution system environments and has high accuracy and robustness. Summary of the Invention

[0005] In view of at least one of the above technical problems, the present invention provides a method, device, equipment and storage medium for abnormal detection and classification of a distribution network, and improves the accuracy of abnormal detection of the distribution network by improving the method.

[0006] According to a first aspect of the present invention, there is provided a method for abnormal detection and classification of a distribution network, including the following steps:

[0007] A method for abnormal detection and classification of a distribution network, characterized by including the following steps:

[0008] S100: Anomaly detection, using DBSCAN and K-Means clustering algorithms to cluster the distribution network data and identify abnormal data points;

[0009] S200: State estimation, using the weighted least squares method to estimate the voltage amplitude of the distribution network and combining an abnormal data point compensation mechanism to improve the estimation accuracy;

[0010] S300: Integration and event classification, integrating the detection results of multiple independent anomaly detectors through an integration method and classifying the detected events;

[0011] Among them, the elbow method is used to determine the optimal number of clusters of the K-Means clustering algorithm, and the DBSCAN algorithm is used to identify low-density points as abnormal data points. The event classification is active power events, reactive power events, complex power events, and faults.

[0012] In some embodiments of the present invention, in the step S100, the following formula is used for calculation:

[0013]

[0014]

[0015] In the formula, x represents all observations; S i represents the i-th cluster; k represents the index of the cluster; x represents a specific observation in the dataset; μ i represents the centroid of the i-th cluster, which is the average of all data points in the cluster; K represents the total number of clusters; S k also represents the k-th cluster, k is a variable representing the index of the cluster being considered; p represents the number of dimensions of the data points; x ij represents the value of the i-th data point on the j-th feature; represents the mean of the j-th feature of the k-th cluster.

[0016] In some embodiments of the present invention, the Net-LOF algorithm and the RRCF algorithm are used to score the data. The Net-LOF algorithm reduces the correlation between the features of the dataset through the feature bagging method, runs the LOF algorithm multiple times to obtain the outlier vectors, and finally takes the average of all outlier vectors as the anomaly detection result.

[0017] In some embodiments of the present invention, the LOF is based on the local density concept of a point and its n adjacent points, and its calculation process is as follows:

[0018] S110: Assume that the n-dist(p) of the data point p is the distance from the data point to its n-th nearest neighbor, and represent the set of n neighbors of p as Sn(p);

[0019] S120: Calculate the reachability distance using the definition of n-dist(.), denoted as r-dist(.) from point q to point p, and its expression is as follows:

[0020] r-dist n (p,q) = max(dist(p,q), n-dist(q))

[0021] In the formula, dist(p,q) is any distance function;

[0022] S130: Using r-dist(.), the local reachability density of data point p with respect to its n nearest neighbors, and its expression is as follows:

[0023]

[0024] In the formula, the reciprocal of the mean of the reachability distances from the n nearest points to data point p;

[0025] S140: For a given lrd n (p), define the LOF of data point p as the average of the ratio of the local reachability distances of all the nearest neighbors of p to the local reachability distance of data point p, and its expression is as follows:

[0026]

[0027] Net-LOF is the combination of single LOF and integrated LOF. Multiply the outlier vectors of single and integrated LOF and find the square root to combine the fractions, and the expression is as follows:

[0028]

[0029] In some embodiments of the present invention, in the step S200, the weighted least squares method technology is adopted, assuming a linear regression model, where X i and Y i are defined as system variables, i = 1,..., n, and the system equation is as follows:

[0030] Y i = α 0 + α 1 X i + ε i

[0031] For ε i ~N(0,σ 2 / w i ), use the minimization of the E w function to estimate the parameters α 0 and α 1 , where the constant w1 ,..., w n , is defined as follows:

[0032]

[0033] The final estimated value is as follows:

[0034]

[0035]

[0036] Wherein, and are defined as the average values after considering the assigned weights;

[0037]

[0038]

[0039] In some embodiments of the present invention, in the step S300, the detection results of each independent anomaly detector are integrated by using the average maximum value, and the anomaly scores are normalized.

[0040] In some embodiments of the present invention, when classifying the detected events, the features of voltage V, current I, active power P, and reactive power Q are extracted based on the DBSCAN algorithm, and the events are classified into active power events, reactive power events, complex power events, and faults.

[0041] According to a second aspect of the present invention, there is also provided a distribution network anomaly detection and classification device, including:

[0042] An anomaly detection module for clustering distribution network data by using DBSCAN and K-Means clustering algorithms to identify anomaly data points;

[0043] A state estimation module for estimating the voltage amplitude of the distribution network by using the weighted least squares method and improving the estimation accuracy by combining an anomaly data point compensation mechanism;

[0044] Integrated into the classification module, the detection results of multiple independent anomaly detectors are integrated by an integration method, and the detected events are classified.

[0045] According to a third aspect of the present invention, there is also provided a distribution network anomaly detection and classification device, including a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method as described above is implemented.

[0046] According to a fourth aspect of the present invention, the sea provides a distribution network anomaly detection and classification storage medium, including a storage medium on which a computer program is stored, and when the computer program is executed by a processor, the above-mentioned method is implemented.

[0047] The beneficial effects of the present invention are as follows: Through the synergistic effect of multiple algorithms, the present invention combines clustering, weighted least squares estimation, and event classification, improving the accuracy and efficiency of distribution network anomaly detection. Compared with traditional technologies, the weighted least squares method is combined with anomaly detection to estimate the voltage amplitude of the distribution system. In an environment with a high noise level and the presence of measurement outliers, the detection accuracy is improved by reasonably allocating weights. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0049] Figure 1 It is a schematic flowchart of the steps of the distribution network anomaly detection and classification method in the embodiment of the present invention;

[0050] Figure 2 It is a schematic flowchart of the calculation steps of the LOF algorithm in the embodiment of the present invention;

[0051] Figure 3 It is a schematic flowchart of the bandit architecture in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0052] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.

[0053] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.

[0054] As Figure 1 The distribution network anomaly detection and classification method shown in the figure includes the following steps:

[0055] S100: Anomaly detection, using the DBSCAN and K-Means clustering algorithms to cluster the distribution network data and identify abnormal data points;

[0056] S200: State estimation, using the weighted least squares method to estimate the voltage amplitude of the distribution network and combining an abnormal data point compensation mechanism to improve the estimation accuracy;

[0057] S300: Integration and event classification, integrating the detection results of multiple independent anomaly detectors through an integration method and classifying the detected events;

[0058] Among them, the elbow method is used to determine the optimal number of clusters of the K-Means clustering algorithm, and the DBSCAN algorithm is used to identify low-density points as abnormal data points. The events are classified as active power events, reactive power events, complex power events, and faults.

[0059] In the embodiment of the present invention, as Figure 1As shown, in step S100, DBSCAN (Density-Based Spatial Clustering of Applications with Noise) and K-Means (a distance-based clustering algorithm) techniques are used to cluster the data. It should be noted here that DBSCAN is a clustering algorithm used to group closed points and identify low-density points as anomalies, especially suitable for the identification of noise points and sparse data points. K-Means clusters the data by partitioning based on the center points, finding similar patterns in the distribution network data, and thus identifying anomaly points that are significantly different from other data. In step S200, the weighted least squares method is used to estimate by assigning different weights to different data points, and the selection of the weights is determined based on the reliability of the data. During this calculation process, an abnormal data point compensation mechanism is incorporated. When the detected abnormal data points are identified, their impact on the estimation result is adjusted according to the compensation mechanism to improve the accuracy and precision of the estimation. In step S300, the integration method is to integrate the detection results of multiple independent anomaly detectors and improve the comprehensiveness and accuracy of anomaly detection through comprehensive analysis. In addition, classification methods are used to classify the detected abnormal events to better understand the nature of the abnormal events. In a distribution system with a large number of dynamic behaviors, such as frequent switching and topological changes, the elbow method can better detect discrete events, even if some electrical quantities between events are similar. The basic idea of the elbow method is to calculate the sum of squared errors of clustering under different numbers of clusters K and plot the change of the sum of squared errors of clustering as K increases. As K increases, the sum of squared errors of clustering will gradually decrease. When K increases to a certain point, the rate of decrease of the sum of squared errors of clustering will become significantly slower. This turning point is the so-called "elbow", and the corresponding K value is the optimal number of clusters. In this embodiment, the K-Means clustering algorithm and the DBSCAN algorithm are combined to perform anomaly detection on the distribution network data. First, K-Means is used to cluster the data, and the optimal number of clusters is determined by the elbow method. Then, DBSCAN is used to identify low-density data points, which are considered abnormal data, to help the system identify and handle anomalies in the distribution network. According to the integrated detection results, the detected abnormal events are classified into four specific types. Active power events, abnormal events related to the change of active power in the system, that is, the effective power output; reactive power events, abnormal events related to the change of reactive power in the system, that is, the non-effective power, usually related to the voltage and power factor of the power grid; complex power events, anomalies involving changes in multiple electrical parameters, which may include active power, reactive power, and other complex interaction factors; fault events, such as equipment failures, line disconnections, etc., that cause problems in the distribution network.

[0060] In the above embodiments, the present invention combines clustering, weighted least squares estimation, and event classification through the synergistic effect of multiple algorithms, improving the accuracy and efficiency of abnormal detection in the distribution network. Compared with traditional technologies, the weighted least squares method is combined with abnormal detection to estimate the voltage amplitude of the distribution system. In an environment with a high noise level and the presence of measurement outliers, the accuracy of detection is improved by reasonably allocating weights.

[0061] In an embodiment of the present invention, in step S100, the following formula is used for calculation:

[0062]

[0063]

[0064] In the formula, x represents all observed values; S i represents the i-th cluster; k represents the index of the cluster; x represents a specific observed value in the dataset; μi represents the centroid of the i-th cluster, which is the average of all data points in the cluster; K represents the total number of clusters; S k also represents the k-th cluster, where k is a variable representing the index of the cluster being considered; p represents the number of dimensions of the data points; x ij represents the value of the i-th data point on the j-th feature; represents the mean of the j-th feature of the k-th cluster.

[0065] The K-Means objective is specified based on the formula The formula is the calculation of the sum of within-cluster distances and is required to calculate the optimal number of clusters.

[0066] In an embodiment of the present invention, the Net-LOF algorithm and the RRCF algorithm are used to score data. The Net-LOF algorithm reduces the correlation between dataset features through a feature bagging method and runs the LOF algorithm multiple times to obtain outlier vectors. Finally, the average value of all outlier vectors is used as the anomaly detection result. The Net-LOF algorithm is an LOF-based anomaly detection method, and LOF is a criterion used to measure whether a data point is significantly different from its neighboring points. The RRCF algorithm is an anomaly detection algorithm based on random cut forests. It evaluates the anomaly degree of data points by constructing a set of trees generated by random cuts. Each data point is "cut" in the tree and gradually forms a set of tree structures. By traversing the tree structures, it is determined whether the data point is significantly different from its neighbors, thereby identifying outlier points. This method performs anomaly detection by randomly cutting the data space and has strong robustness, especially suitable for high-dimensional data and data with complex distributions. Net-LOF obtains the anomaly score of each data point through feature bagging and running the LOF algorithm multiple times, comprehensively considering the data performance under different feature subsets, thereby reducing the bias that may be brought by the correlation between features. Feature bagging is a feature selection algorithm used to reduce the correlation between dataset features. Multiple data points can have similar changes reflected across various features.

[0067] In an embodiment of the present invention, LOF is based on the concept of local density of a point and its n adjacent points, and its calculation process is as follows:

[0068] S110: Assume that the n-dist(p) of data point p is the distance from the data point to its nth nearest neighbor, and the set of p's n neighbors is denoted as Sn(p); this distance metric determines which points are considered neighbors of data point p. The general distance function can be the Euclidean distance or the Manhattan distance, and the Euclidean distance is preferred.

[0069] S120: Use the definition of n-dist(.) to calculate the reachability distance, denoted by r-dist(.) from point q to point p, and its expression is as follows:

[0070] r-dist n (p,q) = max(dist(p,q), n-dist(q))

[0071] In the formula, dist(p,q) is any distance function;

[0072] S130: Use r-dist(.), the local reachability density of data point p with respect to its n nearest neighbors, and its expression is as follows:

[0073]

[0074] Wherein, it is the reciprocal of the mean of the reachable distances of the nearest n points to the data point p;

[0075] S140: For a given lrd n (p), the LOF of the data point p is defined as the average of the ratios of the local reachable distances of all the nearest neighbors of p to the local reachable distance of the data point p, and its expression is as follows:

[0076]

[0077] Net-LOF is a combination of single LOF and integrated LOF. The expression for combining the scores by multiplying the outlier vectors of single and integrated LOF and finding its square root is as follows:

[0078]

[0079] The new dataset with random features can now be passed to the LOF algorithm to obtain the outlier vector. By running the feature bagging algorithm multiple times, different random features and the unique outlier vectors for each set of random features can be obtained. The final result is the average of all the outlier vectors. Since this method runs the LOF algorithm multiple times on a unique feature selection, we call it integrated LOF. In this embodiment, by combining the methods of single LOF and integrated LOF, the feature bagging technique is used to reduce the correlation between features, enhance the robustness of the algorithm, and improve the accuracy of anomaly detection by comprehensively calculating the LOF value.

[0080] Here, another algorithm, Robust Random Cut Forest (RRCF), can also be used to reduce errors. RRCF uses a set of independent robust random cut trees to structure the data into a tree graph. Anomaly detection is performed by comparing the changes in the structure of the graph when adding a data point as a leaf / node. This algorithm is used to detect outliers in streaming data. It can handle high-dimensional data well and reduce the influence of irrelevant dimensions in the dataset. To help us detect scores, the scores are then scaled using a min-max scaler, which is to provide a form of consistency. This is to provide a unified form. The scaled value of the outlier O score of the array output is calculated as follows:

[0081]

[0082] Wherein, O scaled is the scaled outlier score; O max is the maximum value of the outlier score; O minIs the minimum value of the anomaly score. To evaluate whether the RRCF score indicates an outlier, we analyzed the distribution of the data point scores. The obtained RRCF scores were compared with the best-fit line of all the calculated scores. If there are outliers in the dataset, their scores will be abnormally higher than those of the normal data points; if there are no outliers in the dataset, then the scores should be evenly distributed. This score distribution was used to plot the best-fit line for all the calculated data points. Check whether there are any outliers in the RRCF scores. If the check satisfies the existence of outliers, a specific data point can be classified as an outlier by checking the distance between the RRCF score and the value of the best-fit line of this index.

[0083] In an embodiment of the present invention, in step S200, the weighted least squares method technology is adopted, assuming a linear regression model, where X i and Y i are defined as system variables, i = 1,..., n, and the system equation is as follows:

[0084] Y i = α 0 + α 1 X i + ε i

[0085] For ε i ~N(0, σ 2 / w i ), the minimization of the E w function is used to estimate the parameters α 0 and α 1 , where the constants w 1 ,..., w n are defined as follows:

[0086]

[0087] The final estimated values are as follows:

[0088]

[0089]

[0090] In the formula, and are defined as the average values after considering the assigned weights;

[0091]

[0092]

[0093] In the current work, the weighted least squares method is extracted in the general matrix formula with W, and W is defined as having the corresponding weights w1 ,...,w n , as diagonal elements. The development of the distribution system state estimation is based on the original weighted least squares method and the comparison results of outliers and filtered outliers in the result part of this embodiment. To improve the accuracy and robustness of the distribution anomaly detection and classification method, the outlier compensation mechanism is embedded as an abnormal data point x t as part of the distribution system state estimation block:

[0094] Mark xt as an abnormal data point;

[0095] x t and x j will be marked as two previously adjacent normal data points;

[0096]

[0097] Replace with x t instead of

[0098] The anomaly-aware state estimation through the weighted least squares technique, and then the detection and classification of anomalies, can improve the accuracy and stability of the system estimation in the case of imbalance or the presence of abnormal points.

[0099] In an embodiment of the present invention, in step S300, the detection results of each independent anomaly detector are integrated using the average maximum value, and the anomaly scores are normalized. The generation of anomaly detection results by a single detector is always vulnerable to certain assumptions and limitations of specific technologies. To improve the accuracy of anomaly detection, an integration method, EA, is designed in this embodiment. This method integrates the detection results of each independent base detector. The purpose of the integration technique is to improve the accuracy compared to a single detector, enhance the robustness to the sensitivity of the hyperparameters of a single method, and reduce the dependence on the assumptions and limitations that each independent base technology may have. Anomaly detection can be interpreted as identifying high noise levels, missing or bad data, and actual physical events. The details of each detection technique proposed in this embodiment are tailored to the processing requirements and are drawn as independent anomaly detection schemes by the detection strategies used. First, the data extracted by the D-PMU is scanned once using the detector respectively, and then the results are subjected to score conversion; then, the anomaly scores are normalized to reduce the dependence of the detection quality on hyperparameter tuning. The EA method starts from obtaining the anomaly score vector for the simulated observations, which needs to be output by the respective base detectors. The main challenge in this step is the nature of the detection process. The detection processes used for each anomaly score result in different score generation processes. A synchronization step is required to normalize the vectors generated by the independent detectors so that the final score values are within a coherent range for integration. In the final stage of detection, the average maximum value (MOA) technique is used to combine the scores. The result output is used as the detection criterion. In some cases, the clusters formed by the anomaly data points within a given data window may not be able to distinguish between bad or missing data and actual feeder events. Based on the existing base detectors, MOA is selected as the most effective outlier integration method. The advantage of the integration method is to improve the accuracy compared to a single base detector, and at the same time increase the robustness to the sensitivity of hyperparameters and the tuning of a single method. It will also reduce the dependence of the final result on certain assumptions and limitations that each independent base technology may have. The results of each independent technology are converted into quantitative score results. The scores are normalized to reduce the scale dependence. After synchronization, score conversion, and normalization, the MOA technique is used to create scores for each input data point.

[0100] In an embodiment of the present invention, when classifying the detected events, based on the DBSCAN algorithm, the features of voltage V, current I, active power P, and reactive power Q are extracted, and the events are classified into active power events, reactive power events, complex power events, and faults.

[0101] As Figure 3 shown, for active power events: if there is a significant change in the event signal at the active power data point, it is classified as a P event;

[0102] Reactive power event: If the event signal has a significant change in the reactive power data point, it is classified as a Q event;

[0103] Complex power event: If the event signal has significant changes in both active and reactive power data points, it is regarded as a complex event and is assigned a PQ event label;

[0104] Fault: Fault classification is completed before the anomaly classification based on physical rules and is carried out based on the collected local and spatial information. If an event in which a system parameter, namely P, Q, and frequency, changes can be detected by more than a specified number of PMUs, it can be considered a fault.

[0105] In the given test D-PMU dataset, it includes three-phase voltage amplitudes and angles, three-phase current amplitudes and angles, frequency, and rate of change of frequency. Based on these attributes, active power injection and reactive power injection are calculated. Then the processed data is sent to DBSCAN for feature extraction. Feature extraction is a prerequisite step for event classification, which obtains the main features from a segment of the signal. The feature extraction process in the present invention is implemented by a fine-tuned DBSCAN. For the post-processing of event data, after filtering out the incorrect data using basic anomaly classification, the detected event data segments are forwarded to a higher-level event classification process. The classified events can help the power system to perform more targeted processing and response. For example, if a fault event is detected, the system can immediately take isolation or restoration measures; if it is an active power event, load regulation can be carried out; if it is a complex power event, more detailed system analysis may be required to find the specific cause. Through this classification method, the system can take corresponding processing measures according to different types of events, thereby improving the monitoring efficiency and response ability of the distribution network.

[0106] Those skilled in the art should know that the embodiments of the present application can be provided as methods, devices, electronic devices, or computer storage medium products. Therefore, the embodiments of the present application can completely adopt hardware embodiments, hardware and software combined embodiments, or pure software embodiments. The following introduces the test detection process data processing device in the embodiments of the present application. The device embodiments in the following text correspond to the method embodiments in the above text. Those skilled in the art can understand the following implementation process based on the above description and will not be described in detail here.

[0107] In an embodiment of the present invention, a distribution network anomaly detection and classification device is also proposed, which is characterized by including:

[0108] Anomaly detection module, which is used to cluster the distribution network data by using DBSCAN and K-Means clustering algorithms to identify abnormal data points;

[0109] A state estimation module is used to estimate the voltage amplitude of a distribution network by using the weighted least squares method, and improve the estimation accuracy by combining an abnormal data point compensation mechanism;

[0110] Integrated into the classification module, it integrates the detection results of multiple independent anomaly detectors through an integration method and classifies the detected events.

[0111] In the following parts of the embodiments of the present invention, the embodiments of the electronic device and the computer storage medium in the embodiments of the present invention are introduced. The embodiments of the electronic device and the computer storage medium in the following text correspond to the method embodiments in the above text. Those skilled in the art can understand the implementation process in the following text based on the above description, and no detailed description will be given here.

[0112] In an embodiment of the present invention, a computer device is further proposed, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the above method is implemented.

[0113] The embodiment of the present application also provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above method is implemented.

[0114] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0115] The steps of the method or algorithm described in combination with the embodiments disclosed in this article can be directly implemented by hardware, a software module executed by a processor, or a combination of the two. The software module can be placed in a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.

[0116] Those skilled in the art of this industry should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification is only to illustrate the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.

Claims

1. A distribution network anomaly detection and classification method, characterized in that: The following steps are involved: S100: Anomaly detection, using DBSCAN and K-Means clustering algorithms to cluster distribution network data and identify abnormal data points; S200: state estimation, using weighted least squares method to estimate the voltage amplitude of the distribution network, and combining abnormal data point compensation mechanism to improve the estimation accuracy; S300: Integration and event classification, integrating the detection results of multiple independent anomaly detectors through an integration method and classifying the detected events; The elbow method is used to determine the optimal number of clusters for the K-Means clustering algorithm, and the DBSCAN algorithm is used to identify low-density points as abnormal data points. The events are classified into active power events, reactive power events, complex power events, and faults.

2. The method for detecting and classifying anomalies in a distribution network according to claim 1, characterized in that: In the step S100, the following formula is used for calculation: In the formula, x represents all observed values; S i represents the i-th cluster; k represents the cluster index; x represents a specific observation in the data set; μ i represents the centroid of the ith cluster, which is the average value of all data points in the cluster; K represents the total number of clusters; S k It also represents the kth cluster, where k is a variable indicating the index of the cluster under consideration; p represents the number of dimensions of the data points; x ij It represents the value of the i-th data point on the j-th feature; represents the mean of the jth feature of the kth cluster.

3. The method for detecting and classifying anomalies in a distribution network according to claim 1, characterized in that: The Net-LOF algorithm and the RRCF algorithm are used to score the data. The Net-LOF algorithm reduces the correlation between the features of the data set through the feature bagging method, and runs the LOF algorithm multiple times to obtain the outlier vectors. Finally, the average value of all outlier vectors is used as the anomaly detection result.

4. The method for detecting and classifying anomalies in a distribution network according to claim 3, characterized in that: The LOF is based on the local density concept of a point and its n neighboring points, and its calculation process is as follows: S110: Assume that n-dist(p) of a data point p is the distance from the data point to its nth nearest neighbor, and denote the set of n neighbors of p as Sn(p); S120: Use the definition of n-dist(.) to calculate the reachable distance, expressed as r-dist(.) from point q to point p, which is expressed as follows: r-dist n (p,q)=max(dist(p,q),n-dist(q)) Where dist(p,q) is any distance function; S130: Using r-dist(.), the local reachable density point p of the data relative to its nearest n neighbors is expressed as follows: Where, the reciprocal of the mean of the reachable distances from the nearest n points to the data point p; S140: For a given lrd n (p), the LOF of data point p is defined as the average of the ratios of the local reachable distances of all the nearest neighbors of p to the local reachable distance of data point p, which is expressed as follows: Net-LOF is a combination of single LOF and integrated LOF. The expression for combining the scores by multiplying the outlier vectors of single and integrated LOF and finding their square roots is as follows:

5. The method for detecting and classifying anomalies in a distribution network according to claim 1, characterized in that: In the step S200, a weighted least squares technique is used to assume a linear regression model, where X i and Y i is defined as a system variable, i=1,...,n, and the system equation is as follows: Y i =α0+α1X i +e i For ε i ~N(0,σ 2 / w i ), use E w The parameters α0 and α1 are estimated by minimizing the function, where the constants w1,...,w n , defined as follows: The final estimates are as follows: In the formula, and It is defined as the average value after taking into account the assigned weights; 6. The method for detecting and classifying anomalies in a distribution network according to claim 1, characterized in that: In the step S300, the detection results of the independent anomaly detectors are integrated by using the average maximum value, and the anomaly scores are normalized.

7. The method for detecting and classifying anomalies in a distribution network according to claim 1, characterized in that: When classifying the detected events, the features of voltage V, current I, active power P and reactive power Q are extracted based on the DBSCAN algorithm, and the events are classified into active power events, reactive power events, complex power events and faults.

8. A distribution network anomaly detection and classification device, characterized in that: include: The anomaly detection module is used to cluster the distribution network data using DBSCAN and K-Means clustering algorithms to identify abnormal data points; The state estimation module is used to estimate the voltage amplitude of the distribution network using the weighted least square method and improve the estimation accuracy by combining the abnormal data point compensation mechanism; Integrated into the classification module, it combines the detection results of multiple independent anomaly detectors through an ensemble method and classifies the detected events.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.

10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.