Low-quality data detection method and device based on density local outlier algorithm
By calculating the Pearson correlation coefficient and LOF value of the PMU measurement data based on the density local outlier point algorithm, the outlier data points of the power system are screened out, which solves the problem of low detection accuracy in the prior art and realizes higher accuracy low-quality data detection.
Patent Information
- Application Number
- CN202510701325.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-08-29
AI Technical Summary
In the prior art, the PMU measurement data detection method of the power system has a deviation from the actual operating state, resulting in low detection accuracy and difficult to meet the actual needs of the power system.
The density local outlier point algorithm is used to determine the power system state by calculating the Pearson correlation coefficient of the PMU measurement data, and filter out multiple outlier data points from the measurement data, and use the LOF value to determine the low-quality data.
It improves the accuracy of low-quality data detection, can more accurately fit the actual state of the power system, and improves the reliability of data quality and state perception.
Smart Images

Figure CN120561808A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data detection, and in particular to a method and device for detecting low-quality data based on a density local outlier algorithm. Background Art
[0002] As power systems expand and the number of installed electrical equipment increases, the problem of poor data detection becomes increasingly prominent. Synchronized phase measurement units (PMUs), with their advantages of excellent synchronization, high resolution, and direct phase angle measurement, are a crucial information source for online, real-time state perception of power systems. By measuring the voltage and current at each node, PMUs provide essential data for power system state estimation, fault diagnosis, and fault location.
[0003] However, some PMU measurement data collected in the field is not only subject to anomalies such as systematic errors, data loss, jumps, and deviations, but is also susceptible to external interference such as satellite synchronization signal attacks, network attacks, and extreme weather, leading to data errors. This can result in a large amount of abnormal or low-quality data. Therefore, detecting low-quality and abnormal PMU measurement data is crucial for improving data quality and state awareness. A common approach currently involves using a data classification model to identify data anomalies based on the data's numerical value after acquiring PMU measurement data, thereby determining whether the data is of low quality.
[0004] However, the currently commonly used methods have the following technical problems: during the operation of the power system, its real-time state is constantly changing, causing the detected PMU measurement data to change and fluctuate all the time. Using a single classification model to perform detection based on the numerical value of the data deviates from the actual operating state of the power system, resulting in subsequent classification detection results that are difficult to meet actual needs and low detection accuracy. Summary of the Invention
[0005] The present invention provides a low-quality data detection method, device, equipment and medium based on a density local outlier algorithm, which can solve the technical problems of the existing technology that the detection is deviated from the actual operating state and the detection accuracy is low.
[0006] A first aspect of an embodiment of the present invention provides a low-quality data detection method based on a density local outlier algorithm, the method comprising:
[0007] Obtain PMU measurement data of the power system;
[0008] Calculating a Pearson correlation coefficient of the PMU measurement data, and determining a power system state based on the Pearson correlation coefficient;
[0009] A plurality of outlier data points corresponding to the power system state are screened from the PMU measurement data, and low-quality data are determined according to LOF values of the outlier data points.
[0010] After acquiring the power system's PMU measurement data, the present invention determines the power system state based on the Pearson correlation coefficient of the PMU measurement data. Multiple outlier data points corresponding to the power system state are screened from the PMU measurement data, and low-quality data is identified based on the LOF values of the outlier data points. By determining the power system state based on the Pearson correlation coefficient of the data and performing detection processing, the system is aligned with the actual state of the power system, thereby improving detection accuracy.
[0011] In conjunction with the first aspect, in one implementation, screening the plurality of outlier data points corresponding to the power system state from the PMU measurement data includes:
[0012] If the power system state is transient, preprocessing the PMU measurement data at the same time to obtain a transient measurement data set;
[0013] Constructing a transient two-dimensional plane map using the data of the transient measurement data set, and aggregating and classifying the data points in the transient two-dimensional plane map using the K-means clustering algorithm to obtain a plurality of transient classification data sets;
[0014] Outlier data points are screened from each transient classification data set to obtain a plurality of outlier data points corresponding to the transient state.
[0015] In conjunction with the first aspect, in one implementation, constructing a transient two-dimensional plane map using data from the transient measurement data set includes:
[0016] Obtaining a first transient difference value and a second transient difference value of each node from the transient measurement data set, wherein the first transient difference value is a difference between data of each node at time t and data at time t-1, and the second transient difference value is a difference between data of each node at time t-1 and data at time t-2;
[0017] A transient two-dimensional coordinate is generated using the first transient difference as the abscissa and the second transient difference as the ordinate, and data of the transient measurement data set are added as data points to the transient two-dimensional coordinate to obtain a transient two-dimensional plane diagram.
[0018] In conjunction with the first aspect, in one implementation, filtering outlier data points from each transient classification data set to obtain multiple outlier data points corresponding to the transient includes:
[0019] Determining a transient center point of each of the transient classification data sets and removing data points in each of the transient classification data sets that deviate from the transient center point to obtain a transient cleaned data set;
[0020] Calculating the cluster radius of the transient data center point and the distance between each data point in the transient cleaned data set and the transient data center point to obtain a transient radius distance and a transient point distance respectively;
[0021] In each transient cleaned data set, data points whose transient point distance is greater than the transient radius distance are screened to obtain a plurality of outlier data points corresponding to the transient.
[0022] In conjunction with the first aspect, in one implementation, screening the plurality of outlier data points corresponding to the power system state from the PMU measurement data includes:
[0023] If the power system state is steady, pre-processing the PMU measurement data acquired by the same PMU device within a preset time period to obtain a steady-state measurement data set;
[0024] constructing a steady-state two-dimensional plane graph using the data of the steady-state measurement data set, and aggregating and classifying the data points in the steady-state two-dimensional plane graph using the K-means clustering algorithm to obtain a plurality of steady-state classification data sets;
[0025] Outlier data points are screened from each steady-state classification data set to obtain a plurality of outlier data points corresponding to the steady state.
[0026] In conjunction with the first aspect, in one implementation, constructing a steady-state two-dimensional plane map using data from the steady-state measurement data set includes:
[0027] Obtaining a steady-state difference value of each node from the steady-state measurement data set, wherein the steady-state difference value is a difference between data of each node at time t and data at time t-1;
[0028] A steady-state two-dimensional coordinate is generated using the steady-state difference as the vertical coordinate and zero as the horizontal coordinate, and the data of the steady-state measurement data set are added as data points to the steady-state two-dimensional coordinate to obtain a steady-state two-dimensional plane diagram.
[0029] In conjunction with the first aspect, in one implementation, screening outlier data points from each steady-state classification data set to obtain multiple outlier data points corresponding to the steady state includes:
[0030] determining a steady-state central point of each of the steady-state classification data sets and removing data points in each of the steady-state classification data sets that deviate from the steady-state central point to obtain a steady-state cleaned data set;
[0031] Calculating the cluster radius of the steady-state data center point and the distance between each data point in the steady-state cleaned data set and the steady-state data center point to obtain a steady-state radius distance and a steady-state point distance respectively;
[0032] In each of the steady-state cleaned data sets, data points whose steady-state point distance is greater than the steady-state radius distance are screened to obtain a plurality of outlier data points corresponding to the steady state.
[0033] In conjunction with the first aspect, in one implementation, calculating the Pearson correlation coefficient of the PMU measurement data includes:
[0034] Obtaining the voltage amplitude, historical amplitude, and voltage average of two adjacent nodes from the PMU measurement data;
[0035] The Pearson correlation coefficient is calculated using the voltage amplitude, the historical amplitude, and the voltage average.
[0036] In conjunction with the first aspect, in one implementation, determining the power system state according to the Pearson correlation coefficient includes:
[0037] Determining whether the Pearson correlation coefficient is greater than a preset coefficient;
[0038] If the Pearson correlation coefficient is less than a preset coefficient, determining that the power system state is a steady state;
[0039] If the Pearson correlation coefficient is greater than a preset coefficient, it is determined that the power system state is transient.
[0040] A second aspect of an embodiment of the present invention provides a low-quality data detection device based on a density local outlier algorithm, the device comprising:
[0041] An acquisition module is used to obtain PMU measurement data of the power system;
[0042] a determination module, configured to calculate a Pearson correlation coefficient of the PMU measurement data and determine a power system state based on the Pearson correlation coefficient;
[0043] The detection module is used to screen a plurality of outlier data points corresponding to the power system state from the PMU measurement data, and determine low-quality data of the PMU measurement data according to the LOF values of the outlier data points.
[0044] Compared to existing technologies, the low-quality data detection method and device based on a density local outlier algorithm provided by embodiments of the present invention have the following advantages: after acquiring the PMU measurement data of the power system, the present invention can determine the power system status based on the Pearson correlation coefficient of the PMU measurement data; filter multiple outlier data points corresponding to the power system status from the PMU measurement data, and determine low-quality data based on the LOF values of the outlier data points. The power system status is determined by the Pearson correlation coefficient of the data and then detected and processed to match the actual state of the power system, thereby improving detection accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 1 is a flow chart of a method for detecting low-quality data based on a density local outlier algorithm provided by one embodiment of the present invention;
[0046] Figure 2 is a schematic diagram of transient and steady-state data provided by an embodiment of the present invention;
[0047] Figure 3 This is an operational flow chart of a low-quality data detection method based on a density local outlier algorithm provided by one embodiment of the present invention;
[0048] Figure 4 This is a transient process pruning effect diagram provided by an embodiment of the present invention;
[0049] Figure 5 This is a transient process LOF calculation result diagram provided by an embodiment of the present invention;
[0050] Figure 6 Schematic diagram of LOF values of various nodes in a transient process provided by one embodiment of the present invention;
[0051] Figure 7 It is a transient process low-quality data type graph provided by an embodiment of the present invention;
[0052] Figure 8 It is a graph showing low-quality data appearing at a node in a steady-state phase provided by an embodiment of the present invention;
[0053] Figure 9 This is a diagram showing the pruning effect during the steady-state phase provided by an embodiment of the present invention;
[0054] Figure 10 This is a diagram of LOF calculation results in the steady state phase provided by one embodiment of the present invention;
[0055] Figure 11 is a comparison diagram of simulation results provided by an embodiment of the present invention;
[0056] Figure 121 is a schematic structural diagram of a low-quality data detection device based on a density local outlier algorithm provided by one embodiment of the present invention. DETAILED DESCRIPTION
[0057] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0058] As power systems expand and the number of installed electrical equipment increases, the problem of poor data detection becomes increasingly prominent. Synchronized phase measurement units (PMUs), with their advantages of excellent synchronization, high resolution, and direct phase angle measurement, are a crucial information source for online, real-time state perception of power systems. By measuring the voltage and current at each node, PMUs provide essential data for power system state estimation, fault diagnosis, and fault location.
[0059] However, some PMU measurement data collected in the field is not only subject to anomalies such as systematic errors, data loss, jumps, and deviations, but is also susceptible to external interference such as satellite synchronization signal attacks, network attacks, and extreme weather, leading to data errors. This can result in a large amount of abnormal or low-quality data. Therefore, detecting low-quality and abnormal PMU measurement data is crucial for improving data quality and state awareness. A common approach currently involves using a data classification model to identify data anomalies based on the data's numerical value after acquiring PMU measurement data, thereby determining whether the data is of low quality.
[0060] However, the currently commonly used methods have the following technical problems: during the operation of the power system, its real-time state is constantly changing, causing the detected PMU measurement data to change and fluctuate all the time. Using a single classification model to perform detection based on the numerical value of the data deviates from the actual operating state of the power system, resulting in subsequent classification detection results that are difficult to meet actual needs and low detection accuracy.
[0061] In order to solve the above problems, a low-quality data detection method and device based on a density local outlier algorithm provided in an embodiment of the present application will be introduced and explained in detail through the following specific embodiments.
[0062] In order to solve the technical problem that the noise reduction processing of the existing technology removes the signal characteristics of the original sampling signal, resulting in low noise reduction accuracy, refer to Figure 1 , shows a flow chart of a low-quality data detection method based on a density local outlier algorithm provided by an embodiment of the present invention.
[0063] As an example, the low-quality data detection method based on the density local outlier algorithm may include:
[0064] S11. Acquire PMU measurement data of the power system.
[0065] The PMU measurement data of the power system can be obtained and then classified to identify low-quality or abnormal data.
[0066] In actual operation, the detection can be carried out when the power system is in operation, and it is only necessary to calculate the data of all PMU devices in the power system at the current moment and a period of time before.
[0067] Low-quality or outlier data in PMU measurement data refers to data points that significantly deviate from the normal data distribution. These data points clearly deviate from the normal pattern and their generation mechanism appears to be unusual, manifesting as extremely large or extremely small numerical anomalies. These points can arise from a variety of reasons and interfere with the stable operation of the power grid and data analysis. The causes of low-quality or outlier data can be attributed to the following aspects:
[0068] First, equipment failure or anomaly: Sensors and measuring equipment in the power grid may be aged, damaged, or improperly calibrated, resulting in deviations or anomalies in the collected data, thereby generating low-quality data.
[0069] Second, data transmission errors: During the power grid data transmission process, due to communication failures, network congestion or data format conversion errors, the data may be damaged or changed during the transmission process, resulting in low-quality data.
[0070] Third, external interference: The power grid system may be affected by factors such as weather, natural disasters, and human interference, resulting in abnormal fluctuations in measurement data and low-quality data.
[0071] Fourth, data recording or processing errors: During the process of data recording, storage or processing, human errors, software defects or improper algorithms may cause outliers in the data, resulting in low-quality data.
[0072] This low-quality or abnormal data can negatively impact grid condition monitoring, fault diagnosis, and optimized dispatch. Given the complexity of data analysis, researchers typically focus on the general patterns and generation mechanisms of data. Therefore, they do not want low-quality or abnormal data to interfere with subsequent data fitting processes, thereby avoiding analytical bias or errors. Therefore, effective detection and processing of quality or abnormal data within PMU measurement data is necessary to improve data quality and grid reliability.
[0073] S12. Calculate the Pearson correlation coefficient of the PMU measurement data, and determine the power system state according to the Pearson correlation coefficient.
[0074] In one embodiment, when the power grid is operating normally, its state is relatively stable and predictable, and the electrical parameters of each node of the power system, such as voltage and current, generally fluctuate within a certain range. However, when the power grid encounters an abnormal event, such as a short circuit fault, equipment failure removal, sudden switching of large-capacity loads, or lightning strikes, the system will enter a dynamic transient process. During this process, the electrical parameters of the power grid, such as voltage and current, will change dramatically, and the interactions between the nodes will become complex and difficult to directly predict. The state of the power system can be determined by PMU measurement data detected by synchronized phasor measurement units, and then corresponding detection and processing can be performed based on the state to match the actual state of the power system.
[0075] In one embodiment, the power system state includes steady state and transient state. To achieve real-time online analysis of the grid's operating status, simply detecting the moment an abnormal event occurs is insufficient. After an abnormal event occurs, the grid undergoes a transition from transient to steady state, during which the grid's state is constantly changing. Therefore, in addition to having a detection indicator that can promptly detect the occurrence of an abnormal event, a detection indicator is also required that can accurately determine when the transient transition process ends, that is, when the system returns from transient state to steady state.
[0076] During transient conditions, due to the close electrical connections between power grid nodes, the measured values of adjacent nodes often fluctuate in the same direction. This fluctuation significantly increases the correlation between the measured values of the two nodes, demonstrating a strong sense of synchronization. This synchronization makes it possible to monitor power grid status through correlation analysis.
[0077] As the transient process transitions to steady state, the electrical parameters of each node in the power grid gradually stabilize, and the correlation between the measured values of adjacent nodes will continue to decrease. This is because, under steady-state conditions, the electrical parameters of each node are mainly affected by its local load and grid structure, and the direct connection with other nodes is relatively weak.
[0078] When transient fluctuations completely subside and the grid enters steady-state operation, the measured values of adjacent nodes no longer exhibit significant fluctuations in the same direction. At this point, the correlation between the measurements of the two nodes is minimized, leaving only uncertain errors due to factors such as measurement errors and random load fluctuations. This change in correlation provides an important basis for determining whether the grid has transitioned from transient to steady-state operation.
[0079] In an optional embodiment, calculating the Pearson correlation coefficient of the PMU measurement data may include the following sub-steps:
[0080] S121. Obtain the voltage amplitude, historical amplitude, and voltage average value of two adjacent nodes from the PMU measurement data.
[0081] S122. Calculate a Pearson correlation coefficient using the voltage amplitude, the historical amplitude, and the voltage average value.
[0082] The Pearson correlation coefficient (r) measures the linear relationship between two continuous variables. Its value ranges from -1 to 1, with values close to 1 or -1 indicating a strong correlation and values close to 0 indicating no correlation. Its calculation formula is as follows:
[0083]
[0084] Among them, x and y represent the voltage amplitudes of two adjacent nodes, and n data within a period of time before and at the target time are extracted. Represents the average value of n data.
[0085] In an optional embodiment, determining the power system state according to the Pearson correlation coefficient may include the following sub-steps:
[0086] S123: Determine whether the Pearson correlation coefficient is greater than a preset coefficient.
[0087] S124. If the Pearson correlation coefficient is less than a preset coefficient, it is determined that the power system state is a steady state.
[0088] S125. If the Pearson correlation coefficient is greater than a preset coefficient, it is determined that the power system state is transient.
[0089] Therefore, by monitoring the correlation changes between adjacent node measurements in real time and setting the r value to greater than 0.1, the grid is in a transient transition process. This can effectively determine when the grid's operating state has suddenly changed, and whether the transient process has ended and the grid has returned to steady-state operation. When the r value is less than 0.1, the grid has entered steady-state; when it is greater than 0.1, the grid is still in the transient process, thus enabling the identification of transient and steady-state processes in the power system.
[0090] Reference Figure 2 , which shows a schematic diagram of transient and steady-state data provided by an embodiment of the present invention.
[0091] In actual power systems, due to factors such as equipment failure, communication congestion, and climate change, PMU data is frequently missing or abnormal, resulting in poor data such as jumps and deviations. In this paper, these missing data and poor data points are collectively referred to as low-quality data or abnormal data.
[0092] In actual operation, most low-quality and abnormal data in PMU measurements is a single-point random occurrence, with occasional long-term continuous low-quality data. Therefore, transients can be divided into single-point transients and continuous multi-point steady states, and steady states can be divided into single-point steady states and continuous multi-point steady states.
[0093] A single point refers to a single low-quality or abnormal data point, while multiple points refer to the simultaneous presence of multiple low-quality or abnormal data points. First, after distinguishing between transient and steady states, each can be detected using its own method. Second, both single and multiple points can be detected simultaneously.
[0094] S13 , screening a plurality of outlier data points corresponding to the power system state from the PMU measurement data, and determining low-quality data according to LOF values of the outlier data points.
[0095] After the power system state is determined based on the PMU measurement data, the corresponding data can be filtered in different ways according to the power system state, thereby obtaining multiple outlier data points corresponding to different power system states. Subsequently, the LOF algorithm is used to calculate the outlier factors of multiple outlier data points in the set of candidate outlier data points to further determine the specific identity of the outlier points.
[0096] Specifically, the LOF value of each outlier data point in the PMU measurement data can be calculated, and then the LOF value can be used to determine whether the outlier data point is abnormal data or low-quality data.
[0097] When screening outlier data points, the density-based LOF algorithm determines the outlier degree by calculating the ratio of the average local data density of a data point p to its k nearest neighbors. To obtain the local data density, the minimum hypersphere radius r containing the k nearest neighbors is first determined, and then divided by k to obtain the local data density of the data point. Normal data points located in high-density areas have local data densities close to those of their nearest neighbors, resulting in an outlier degree close to 1. In contrast, outlier data points located in low-density areas have local data densities lower than the average density of their nearest neighbors, resulting in an outlier degree greater than 1. A higher outlier degree indicates that the local data density of data point p is lower than the average local data density of its nearest neighbors, making p more likely to be an outlier and thus identifying low-quality data. The LOF algorithm utilizes local data information and considers density differences between different data points to effectively identify outlier data points p1 and p2, thereby identifying low-quality data within the PMU measurement data.
[0098] Among them, the expression of LOF of point p is:
[0099]
[0100] This formula represents the neighborhood points N of point p k The average ratio of the local reachability density (LRD) of point p to the LRD of point p. Under normal circumstances, the LOF value of point p approaches 1, indicating that point p has a similar point density to its neighborhood and belongs to the same cluster as its neighborhood. When the LOF value of point p is greater than 1, the density of point p is lower than that of its neighborhood. The larger the outlier factor, the higher the degree of outlier of the data point, and the more likely it is an outlier, that is, low-quality data.
[0101] In actual operation, the specific concept of density-based LOF algorithm is defined as follows:
[0102] (1) Point-to-point distance:
[0103] d(p,o) refers to the distance from data point p to data point o.
[0104] (2) kth distance:
[0105] The kth distance d of data point p k (p), defined as:
[0106] d k (p) = d(p,o);
[0107] Simply put, it is to spread outward from point p until it circles the kth neighboring point.
[0108] (3) kth distance neighborhood:
[0109] The kth distance neighborhood N of the data point p k (p) refers to the set of all points within the kth distance of point p, including the points on the kth distance.
[0110] (4) kth reachable distance:
[0111] reach_dist k (p,o)=max{d k (o),d(p,o)};
[0112] The kth reachable distance from data point o to data point p refers to the larger value between the kth distance of point o and the true distance between o and p.
[0113] (5) Local Reachability Density (LRD):
[0114] The LRD of point p represents the point N in the kth neighborhood of point p. k The LRD of point p is the reciprocal of the average reachable distance from point (p) to p.
[0115]
[0116] (6) Local Outlier Factor (LOF):
[0117] The expression of LOF of point p is:
[0118]
[0119] This formula represents the neighborhood points N of point p k The average of the ratios of the local reachability density LRD of (p) to the LRD of point p.
[0120] The density-based LOF algorithm needs to traverse the outlier factor of each data point when calculating, so the complexity of the first step of determining the neighborhood of all nodes is O(n). Then, the time complexity of calculating the LOF value of all points in the entire data set will reach O(n). 2 When processing high-dimensional, large-scale data sets, the computational efficiency of this quadratic form is too low, which is not conducive to real-time detection of PMU measurement data. Therefore, it is necessary to improve the computational complexity of LOF.
[0121] Finally, the LOF value is calculated for the outlier data points in this "outlier candidate set". Points with LOF values greater than the threshold are outlier data points, that is, low-quality data points detected, i.e. low-quality data. This threshold can be adjusted according to actual needs.
[0122] The above operation method can reduce the time complexity of the algorithm while ensuring the detection accuracy. The K-means clustering algorithm can prune at least half of the data set. Therefore, including the calculation of the pruning process, the total time complexity is O(MT+n 2 / 4).
[0123] Subsequently, low-quality data can be eliminated and high-quality data can be filled into the gaps left after elimination, thereby improving the overall data quality.
[0124] In one embodiment, the power system state is a transient state. As an example, the step of filtering out a plurality of outlier data points corresponding to the power system state from the PMU measurement data may include the following sub-steps:
[0125] S21. If the power system state is transient, pre-process the PMU measurement data at the same time to obtain a transient measurement data set.
[0126] In one embodiment, when performing outlier detection on low-quality data or abnormal data in a transient power system state, it is difficult to distinguish and detect outliers using data from a single PMU measurement device due to the large fluctuations in transient data and the small timing of PMU measurements.
[0127] Therefore, the spatial correlation of power systems is indeed a complex and important concept, which is mainly determined by the electrical coupling relationship and characteristic parameters of transmission lines. This correlation stems from the direct electrical connection formed when different power plants and stations transmit power through transmission lines. Characteristic parameters include resistance, inductance, capacitance, etc., which determine the electrical characteristics of the power system. Electrical coupling relationship refers to the electrical interdependence and mutual influence formed by the connection between different nodes in the power system through transmission lines. This coupling relationship means that the state change of one node (such as changes in voltage and current) will directly affect the state of the adjacent nodes connected to it. Therefore, in the power system, a fault or abnormal state of one node may quickly propagate to the entire system through the electrical coupling relationship, leading to larger-scale faults or instability.
[0128] In a regional power grid, there is a strong spatial correlation between nodes. Therefore, the data from multiple PMUs at the same time can be used to jointly detect outliers of low-quality or abnormal data.
[0129] In one approach, preprocessing can involve data cleaning, which removes low-quality or abnormal data such as outliers as noise. However, from the perspective of data mining, some outliers may contain valuable information, helping to discover interesting phenomena and driving advances in practical applications. Therefore, during data processing, it is necessary to balance the consideration of retaining or removing outliers based on the specific situation to fully utilize the information in the data.
[0130] S22 . Construct a transient two-dimensional plane graph using the data of the transient measurement data set, and aggregate and classify the data points in the transient two-dimensional plane graph using the K-means clustering algorithm to obtain multiple transient classification data sets.
[0131] A planar graph of data points is constructed using the data from the transient measurement dataset as data points, resulting in a transient two-dimensional planar graph. Subsequently, a K-means clustering algorithm can be used to aggregate and classify the multiple data points within the transient two-dimensional planar graph, resulting in multiple transient classification datasets. The K-means clustering algorithm is a distance-based clustering algorithm that divides the dataset into K clusters. Each transient classification dataset can be considered a cluster.
[0132] In one embodiment, classic machine learning algorithms such as support vector machines, logistic regression, and decision trees can also be used for classification, that is, using samples of known categories to train a classifier, and then classify samples of unknown categories. In contrast, clustering is to divide samples into different categories based on the intrinsic relationship between data in the absence of sample category labels, ensuring high similarity between samples of the same category and low similarity between samples of different categories. Classification problems fall under the category of supervised learning, while clustering falls under the category of unsupervised learning. Through these algorithms, the present invention can process and analyze data under different learning modes to achieve more accurate category division.
[0133] The K-means clustering algorithm is a distance-based clustering algorithm that divides a data set into K clusters such that each data point belongs to the closest cluster and the center of the cluster is the average of all data points. The algorithm is implemented through iterative optimization, and each iteration step updates the center point of the cluster until the convergence condition is reached.
[73] At the beginning of the algorithm, the dataset is divided into K clusters, and K data points are randomly selected as the initial cluster centers. Each data point is then assigned to the cluster center closest to it, and each data point can only belong to one cluster. Next, the cluster centers are updated based on the assigned data points. This is achieved by calculating the average value of the data points belonging to each cluster. This process is repeated until the class centers no longer change significantly or the preset number of iterations is reached. The specific steps are as follows.
[0134] (1) Define the cost function as:
[0135]
[0136] Where x i represents the i-th sample, c i is x i The cluster to which it belongs, represents the center point of the cluster, and M is the total number of samples.
[0137] (2) For each sample x i , assign it to the closest cluster:
[0138]
[0139] Where T=0,1,2,… is the number of iteration steps.
[0140] (3) For each cluster K, recalculate the center of the cluster:
[0141]
[0142] (4) Repeat steps 2 and 3 above until J converges.
[0143] The advantages of the K-means clustering algorithm are its simplicity, high computational efficiency, and ability to process large datasets. However, it also has some disadvantages, such as being sensitive to the choice of initial category centers, the possibility of falling into local optimal solutions, and the need to pre-set the number of categories K. The choice of parameter K in the K-means clustering algorithm has a significant impact on the clustering results. If K is too large, the categories may be overly subdivided, causing some samples that should belong to the same category to be separated; if K is too small, the categories may be too general, causing some samples that should belong to different categories to be merged. Therefore, in practical applications, it is necessary to select an appropriate K value based on the specific problem and data characteristics.
[0144] The present invention demonstrates excellent scalability and efficiency when processing large datasets. Its computational complexity is near-linear, specifically O(MKT), where M represents the number of data objects, K is the number of clusters, and T is the number of iterations. Although the algorithm often reaches a local optimum, in most cases this local optimum is sufficient to meet clustering requirements.
[0145] Since the LOF algorithm is needed to calculate the LOF value of the outlier data point later, the LOF algorithm analyzes the positional relationship of the data points on the two-dimensional plane diagram. Therefore, the data at time t measured by a PMU device is used as the x-axis value, and the data at time t-1 is used as the y-axis value. At this time, at the same time, multiple PMU measurements can generate multiple data points on the two-dimensional plane. However, it is not feasible to directly use the data measured by the PMU for LOF algorithm detection. Since the node data measured by the PMU device is distributed within a certain range, when the low-quality data on these nodes falls within this range, it cannot be detected. In order to solve the above problem, as an example, the construction of a transient two-dimensional plane diagram using the data of the transient measurement data set can include the following sub-steps:
[0146] S221. Acquire a first transient difference value and a second transient difference value of each node from the transient measurement data set, wherein the first transient difference value is a difference between data of each node at time t and data at time t-1, and the second transient difference value is a difference between data of each node at time t-1 and data at time t-2;
[0147] S222: Generate a transient two-dimensional coordinate using the first transient difference as the abscissa and the second transient difference as the ordinate, and add the data of the transient measurement data set as data points to the transient two-dimensional coordinate to obtain a transient two-dimensional plane diagram.
[0148] Due to the spatial correlation between each node in the power system, the spatial fluctuations of each PMU measurement data are similar during the transient phase of the power system. Therefore, the data difference between the previous and next moments is used as the horizontal and vertical coordinates of the LOF algorithm data point. To achieve real-time low-quality outlier detection of PMU measurements, the difference between the PMU measurement data at the latest moment t and the data at the previous moment is the first transient difference, which is used as the horizontal coordinate of the data center point. Similarly, the difference between the PMU measurement data at each node at time t-1 and the data at the previous moment can be used as the second difference, which is also used as the vertical coordinate of the data center point.
[0149] Then, a coordinate system is generated using the horizontal and vertical coordinates to obtain a transient two-dimensional coordinate system. The data of the transient measurement data set are then added as data points to the transient two-dimensional coordinate system to obtain a transient two-dimensional plane diagram.
[0150] S23 , screening outlier data points from each transient classification data set to obtain a plurality of outlier data points corresponding to the transient state.
[0151] In one embodiment, some data points in a transient classification data set may not appear abnormal in the global scope, but exhibit significant outlier characteristics in the local scope. In order to address the limitations of the global outlier detection method in identifying local outliers, a density-based local outlier detection method has emerged. The local outlier detection method can be used to screen outlier data points in a transient classification data set to obtain multiple outlier data points corresponding to the transient. This method is based on the nearest neighbor assumption, which believes that data points with lower neighbor density are more likely to be outliers, while data points located in high-density neighbor areas are regarded as normal points. For a given data set, the kth nearest neighbor distance of a data point can be regarded as the radius of a hyperplane or hypersphere centered on it, or the boundary of a certain linear region. The inverse of the distance metric can be used as a basic quantitative indicator of density to more accurately identify local outliers.
[0152] As an example, the step of filtering outlier data points from each transient classification data set to obtain multiple outlier data points corresponding to transients may include the following sub-steps:
[0153] S231 , determining a transient data center point of each transient classification data set and removing data points in each transient classification data set that deviate from the transient data center point to obtain a transient cleaned data set.
[0154] S232: Calculate the cluster radius of the transient data center point and the distance between each data point in the transient cleaned data set and the transient data center point, and obtain a transient radius distance and a transient point distance respectively.
[0155] S233: Filter data points whose transient point distance is greater than the transient radius distance in each transient cleaned data set to obtain a plurality of outlier data points corresponding to the transient.
[0156] In the detection of local outlier data points, data points in high-density areas are generally considered normal because they are surrounded by many other data points, forming a relatively dense data distribution. In contrast, data points in low-density areas, that is, those with few or no other data points around them, are considered outliers. Whether it is a high-density or low-density area, the center of the area is surrounded by a cluster of data with similar data size characteristics. In the LOF calculation, this part of the data is dispensable for outlier detection. Therefore, in order to improve the operating efficiency of the LOF algorithm, it is necessary to prune and remove this data.
[0157] In one operation mode, after the original data is divided into K clusters using the K-means clustering algorithm to obtain multiple transient cleaned data sets, the center point of each cluster (ie, transient cleaned data set) can be calculated to obtain the transient data center point.
[0158] Then we can calculate the average distance from all data points in the cluster to the center point, and use this value as the radius R of the cluster to obtain the transient radius distance;
[0159] For each point in each cluster category, if the distance from a point to its cluster center is not less than the radius R of the preset cluster, the point is included in the "outlier candidate set".
[0160] However, low-quality PMU data detection methods for transient and steady-state processes have their own unique characteristics. In particular, during steady-state conditions, the LOF algorithm dataset consists solely of measurement data collected by a single device over a period of time. During transient conditions, the LOF algorithm dataset consists of the difference between the data from the previous and next moments. Under normal circumstances, these data are similar, belonging to the same type and can be grouped together. Therefore, there is no need to classify these datasets.
[0161] When no classification is required, there is only one cluster, that is, K=1, so the computational complexity is reduced to O(MT).
[0162] At the same time, when calculating the cluster radius R, we need to remove the influence of outliers on the size of R. To do this, we calculate the median of the distances from all points to the center point. If the distance D of a point to the center point is greater than 4 times, it is considered that the point will greatly affect the size of the cluster radius and is removed. The remaining points are used to calculate the cluster radius R.
[0163] To this end, in specific operations, the K-means clustering algorithm can be used to determine the transient center point of each transient classification data set; then, the data points that are too far away from the center point in each transient classification data set are removed. In specific operations, the data points that are too far away from the center point in the data set can be removed according to the following formula:
[0164]
[0165] In the above formula, when m takes the value of 1, the data point is retained to calculate the radius R of the cluster; when m = 0, it is removed.
[0166] After removing low-quality data points, the average distance to the center point is calculated using the remaining data points and recorded as the radius R of the cluster to obtain the transient radius distance.
[0167] For the points in each cluster category, if the distance from a point to its cluster center is greater than the preset cluster radius R, the point is classified into the "outlier candidate set", thereby obtaining multiple outlier data points corresponding to the transient state to determine the transient low-quality data.
[0168] As an example, screening the plurality of outlier data points corresponding to the power system state from the PMU measurement data may include the following sub-steps:
[0169] S31. If the power system state is steady, pre-process the PMU measurement data acquired by the same PMU device within a preset time period to obtain a steady-state measurement data set.
[0170] For a steady-state power system, when using the LOF algorithm to detect low-quality data points, since the power system is a dynamic system with strong inertia, the strong temporal correlation between the measured data of the same PMU device over a period of time can be fully utilized.
[0171] In steady state, when performing real-time low-quality data detection on a node's PMU measurements using the LOF algorithm, calculations are performed by simply inputting the node's current PMU measurement data and the previous period. In practice, preprocessing can also involve data cleaning, which is similar to step S21 above. To avoid repetition, this will not be detailed here; please refer to the above description for details.
[0172] S32. Construct a steady-state two-dimensional plane graph using the data of the steady-state measurement data set, and aggregate and classify the data points in the steady-state two-dimensional plane graph using the K-means clustering algorithm to obtain multiple steady-state classification data sets.
[0173] In one embodiment, a planar graph of data points can be constructed using the data from the steady-state measurement dataset as data points, resulting in a steady-state two-dimensional planar graph. Subsequently, a K-means clustering algorithm can be used to cluster and classify the multiple data points within the steady-state two-dimensional planar graph, resulting in multiple steady-state classified datasets. The K-means clustering algorithm is a distance-based clustering algorithm that divides the dataset into K clusters. Each steady-state classified dataset can be considered a cluster.
[0174] The operation is the same as that of the above step S22. In order to avoid repetition, it will not be described here. Please refer to the above description for details.
[0175] As an example, the step of constructing a steady-state two-dimensional plane map using the data of the steady-state measurement data set may include the following sub-steps:
[0176] S321. Obtain a steady-state difference value of each node from the steady-state measurement data set, wherein the steady-state difference value is a difference value between data of each node at time t and data at time t-1.
[0177] S322 , generating a steady-state two-dimensional coordinate using the steady-state difference as the vertical coordinate and zero as the horizontal coordinate, and adding the data of the steady-state measurement data set as data points to the steady-state two-dimensional coordinate to obtain a steady-state two-dimensional plane diagram.
[0178] In one embodiment, the calculation of the steady-state difference is the same as the calculation of the first transient-state difference, using the data difference between the previous and next moments as the vertical coordinate, and the horizontal coordinate is set to 0.
[0179] Then, a steady-state two-dimensional coordinate is generated with the steady-state difference as the vertical coordinate and the zero value as the horizontal coordinate, and the data of the steady-state measurement data set are added as data points to the steady-state two-dimensional coordinate to obtain a steady-state two-dimensional plane diagram.
[0180] S33. Filter outlier data points from each steady-state classification data set to obtain multiple outlier data points corresponding to the steady state.
[0181] Since the low-quality data detection is performed on the PMU measurement in real time, it is only necessary to calculate whether the LOF value of the data difference between the latest moment and the previous moment exceeds the set threshold, so that multiple outlier data points corresponding to the steady state can be screened out.
[0182] As an example, the step of filtering outlier data points from each steady-state classification data set to obtain a plurality of outlier data points corresponding to the steady state may include the following sub-steps:
[0183] S331 , determining the steady-state data center point of each of the steady-state classification data sets and removing data points in each of the steady-state classification data sets that deviate from the steady-state data center point to obtain a steady-state cleaned data set.
[0184] S332. Calculate the cluster radius of the steady-state data center point and the distance between each data point in the steady-state cleaned data set and the steady-state data center point to obtain a steady-state radius distance and a steady-state point distance, respectively.
[0185] S333: Filter data points whose steady-state point distance is greater than the steady-state radius distance in each steady-state cleaned data set to obtain multiple outlier data points corresponding to the steady state.
[0186] In one operation mode, the operation process of steps S331-S333 is the same as the operation process of steps S231-S233. For details, please refer to the above analysis description. In order to avoid repetition, it will not be repeated here.
[0187] Reference Figure 3 , shows an operational flow chart of a low-quality data detection method based on a density local outlier algorithm provided by an embodiment of the present invention.
[0188] Specifically, the operation of the low-quality data detection method based on the density local outlier algorithm includes the following steps:
[0189] The first step is to obtain PMU measurement data.
[0190] The second step is to determine the operating state of the power system, including transient and steady state.
[0191] In the third step, when the operating state of the power system is transient, the data measured by a single node PMU over a period of time is preprocessed into a LOF dataset, the LOF dataset is pruned to obtain an outlier candidate set, and LOF calculation is performed on the outlier candidate set.
[0192] In the fourth step, when the power system is in a steady state, the PMU measurement data of multiple nodes at the same time are preprocessed into a LOF dataset, the LOF dataset is pruned to obtain an outlier candidate set, and LOF calculation is performed on the outlier candidate set.
[0193] The fifth step is to determine whether the data point is low-quality data or abnormal data based on the LOF value.
[0194] Based on the above process, the proposed method was tested and validated using the IEEE 39 power system. The proposed algorithm was implemented and run in MATLAB 2020a. Power management units (PMUs) were deployed at all 39 nodes in the power system for online monitoring. The generation and transmission voltages within the power system were set to 13.8 kV and 345 kV, respectively, with a rated capacity of 100 MVA. The sampling interval was set, and Gaussian noise was added to the simulation data.
[0195] For transient process simulation verification: the difference between the PMU measurement data of the 39 nodes in the simulation system at the latest time t and the data at the previous time is used as the horizontal coordinate of the center point of the data set, and the difference between the PMU measurement data of each node at time t-1 and the data at the previous time is used as the vertical coordinate of the center point of the data set.
[0196] The dataset is pruned by calculating the cluster center and cluster radius R. The LOF algorithm is then used to detect outliers.
[0197] Reference Figure 4-5 , respectively showing a transient process pruning effect diagram provided by an embodiment of the present invention and a transient process LOF calculation result diagram provided by an embodiment of the present invention.
[0198] When a short circuit occurs in the system at 4.0s and the switch is reclosed at 4.1s, the pruning process and its detection results are as follows: Figure 4 and Figure 5 shown.
[0199] from Figure 4 As can be seen from the figure, except for the outliers, the density of other data points is not much different, and they belong to the same cluster. Therefore, by calculating the cluster center point and the cluster radius R, the data set can be pruned. Figure 4 The presence or absence of the green data points in the circle formed by the center point and radius R has little impact on the detection of outliers and can be completely eliminated. The remaining outlier candidate set is used to detect outliers.
[0200] Reference Figure 6 , which shows a schematic diagram of the LOF values of each node in the transient process provided by an embodiment of the present invention. The LOF value obtained is as follows Figure 6 shown.
[0201] like Figure 6 As shown in the figure, the LOF value of a normal data point is around 1. When the LOF value of a data point is greater than the threshold of 2, it is considered that low-quality data has appeared on that node. For example, the LOF value of node 6 in the figure is 4.93325, which exceeds the threshold of 2, so an outlier is detected on node 5. The LOF values of nodes 10 and 18 are 3.41233 and 4.29833, respectively, both exceeding the threshold of 2, so low-quality data is detected on these nodes at the same time.
[0202] Reference Figure 7 , shows a transient process low-quality data type diagram provided by an embodiment of the present invention. Graph analysis is performed on the node data where abnormal values are detected, such as Figure 7As shown in the figure, the algorithm can effectively and accurately detect low-quality data even close to the occurrence of a grid anomaly. Furthermore, when PMU measurements at multiple nodes show abnormal values at the same time, at 4.20s, the algorithm can accurately detect outliers.
[0203] For steady-state simulation verification, under steady-state conditions, the LOF algorithm performs real-time low-quality detection of PMU measurement data at a specific node. Simply input the current moment and its short-term historical PMU measurement data for that node. Similar to transient calculations, the vertical axis uses the difference between the previous and next moment's data as the coordinate, while the horizontal axis is fixed at 0. Given the real-time low-quality detection of PMU measurements, the key step is to calculate the LOF value of the difference between the latest and previous moment's data and determine whether it exceeds a preset threshold.
[0204] Reference Figure 8 , shows a graph of low-quality data appearing at a node in the steady-state phase provided by an embodiment of the present invention. Assume that in the steady-state phase, at 7.01s, node 6 has data anomalies, such as Figure 8 .
[0205] Reference Figure 9 , shows the pruning effect diagram of the steady state stage provided by an embodiment of the present invention. When testing it, the 70 data points before 7.01s are subtracted from the data of the previous moment as the vertical coordinates of the 70 data points, and the horizontal coordinates are all set to 0 to obtain the LOF data set. After pruning this data set, the following is obtained: Figure 9 results.
[0206] Reference Figure 10 , which shows a diagram of LOF calculation results in the steady-state stage provided by an embodiment of the present invention.
[0207] After pruning the dataset, Figure 10 The remaining outlier candidate set is divided into three regions. Two of these regions have a large and concentrated amount of data. The remaining region contains only a single data point, which is relatively far away from the other two regions. This point is the outlier. This point is the difference between the measurement value of node 6 at 7.01s and the value at the previous moment. Its LOF value is 11.2638, far exceeding the LOF threshold of 2, making it easily detected as an outlier.
[0208] Simulation verification of the effect of the improved LOF algorithm: In order to compare and analyze how much the improved LOF algorithm improves the computational efficiency, the running time of the LOF algorithm before and after the improvement is simulated and verified. Figure 11 , which shows a comparison diagram of simulation results provided by an embodiment of the present invention; the simulation results are as follows Figure 11 shown.
[0209] Figure 11 The experimental data clearly shows the difference in outlier detection before and after the improvement of the LOF method. The LOF method before the improvement directly performs outlier detection on the entire data set, while the improved LOF algorithm uses clustering "pruning" technology to preprocess the original data set before performing outlier detection. Figure 11 The runtime data in Figure 2 shows that the improved LOF algorithm has a significantly lower time complexity than the original LOF method, further demonstrating its efficiency advantage. While maintaining a PMU acquisition rate of 10 milliseconds, sufficient node measurement data can be input during transient processes, or a longer period of data from a single PMU device can be input during steady-state processes to improve detection accuracy and robustness.
[0210] From the simulation case of the present invention, it can be concluded that the proposed low-quality data detection method based on the density-based local outlier algorithm analyzes the low-quality data type of PMUs and fully utilizes the strong spatial correlation between the voltage and current phasors between nodes during transient processes, as well as the certain temporal continuity and stability of individual PMU measurement data during steady-state processes to perform low-quality data detection based on the density-based LOF algorithm. The simulation results of the example show that this method can fully utilize the spatial and temporal correlation of PMU measurement data during transient and steady-state processes to detect low-quality data. It can effectively and accurately detect low-quality data in the measurement, providing strong protection for the safe and stable operation of the power system.
[0211] In this embodiment, a method for detecting low-quality data based on a density local outlier algorithm is provided. The method has the following beneficial effects: after acquiring PMU measurement data from a power system, the method can determine the power system state based on the Pearson correlation coefficient of the PMU measurement data; multiple outlier data points corresponding to the power system state are screened from the PMU measurement data to determine low-quality data based on the LOF values of the outlier data points. By determining the power system state based on the Pearson correlation coefficient of the data and performing detection processing, the method is able to accurately reflect the actual state of the power system, thereby improving detection accuracy.
[0212] The embodiment of the present invention also provides a low-quality data detection device based on a density local outlier algorithm, see Figure 12 , shows a structural schematic diagram of a low-quality data detection device based on a density local outlier algorithm provided by an embodiment of the present invention.
[0213] Wherein, as an example, the low-quality data detection device based on the density local outlier algorithm may include:
[0214] An acquisition module 201 is used to acquire PMU measurement data of the power system;
[0215] a determination module 202 for calculating a Pearson correlation coefficient of the PMU measurement data and determining a power system state based on the Pearson correlation coefficient;
[0216] The detection module 203 is configured to filter outlier data points corresponding to the power system state from the PMU measurement data, and determine low-quality data of the PMU measurement data based on LOF values of the outlier data points.
[0217] Optionally, screening a plurality of outlier data points corresponding to the power system state from the PMU measurement data includes:
[0218] If the power system state is transient, preprocessing the PMU measurement data at the same time to obtain a transient measurement data set;
[0219] Constructing a transient two-dimensional plane map using the data of the transient measurement data set, and aggregating and classifying the data points in the transient two-dimensional plane map using the K-means clustering algorithm to obtain a plurality of transient classification data sets;
[0220] Outlier data points are screened from each transient classification data set to obtain a plurality of outlier data points corresponding to the transient state.
[0221] Optionally, constructing a transient two-dimensional plane map using data from the transient measurement data set includes:
[0222] Obtaining a first transient difference value and a second transient difference value of each node from the transient measurement data set, wherein the first transient difference value is a difference between data of each node at time t and data at time t-1, and the second transient difference value is a difference between data of each node at time t-1 and data at time t-2;
[0223] A transient two-dimensional coordinate is generated using the first transient difference as the abscissa and the second transient difference as the ordinate, and data of the transient measurement data set are added as data points to the transient two-dimensional coordinate to obtain a transient two-dimensional plane diagram.
[0224] Optionally, filtering outlier data points from each transient classification data set to obtain a plurality of outlier data points corresponding to the transient state includes:
[0225] Determining a transient center point of each of the transient classification data sets and removing data points in each of the transient classification data sets that deviate from the transient center point to obtain a transient cleaned data set;
[0226] Calculating the cluster radius of the transient data center point and the distance between each data point in the transient cleaned data set and the transient data center point to obtain a transient radius distance and a transient point distance respectively;
[0227] In each transient cleaned data set, data points whose transient point distance is greater than the transient radius distance are screened to obtain a plurality of outlier data points corresponding to the transient.
[0228] Optionally, screening a plurality of outlier data points corresponding to the power system state from the PMU measurement data includes:
[0229] If the power system state is steady, pre-processing the PMU measurement data acquired by the same PMU device within a preset time period to obtain a steady-state measurement data set;
[0230] constructing a steady-state two-dimensional plane graph using the data of the steady-state measurement data set, and aggregating and classifying the data points in the steady-state two-dimensional plane graph using the K-means clustering algorithm to obtain a plurality of steady-state classification data sets;
[0231] Outlier data points are screened from each steady-state classification data set to obtain a plurality of outlier data points corresponding to the steady state.
[0232] Optionally, constructing a steady-state two-dimensional plane map using data from the steady-state measurement data set includes:
[0233] Obtaining a steady-state difference value of each node from the steady-state measurement data set, wherein the steady-state difference value is a difference between data of each node at time t and data at time t-1;
[0234] A steady-state two-dimensional coordinate is generated using the steady-state difference as the vertical coordinate and zero as the horizontal coordinate, and the data of the steady-state measurement data set are added as data points to the steady-state two-dimensional coordinate to obtain a steady-state two-dimensional plane diagram.
[0235] Optionally, filtering outlier data points from each steady-state classification data set to obtain a plurality of outlier data points corresponding to the steady state includes:
[0236] determining a steady-state central point of each of the steady-state classification data sets and removing data points in each of the steady-state classification data sets that deviate from the steady-state central point to obtain a steady-state cleaned data set;
[0237] Calculating the cluster radius of the steady-state data center point and the distance between each data point in the steady-state cleaned data set and the steady-state data center point to obtain a steady-state radius distance and a steady-state point distance respectively;
[0238] In each of the steady-state cleaned data sets, data points whose steady-state point distance is greater than the steady-state radius distance are screened to obtain a plurality of outlier data points corresponding to the steady state.
[0239] Optionally, calculating the Pearson correlation coefficient of the PMU measurement data includes:
[0240] Obtaining the voltage amplitude, historical amplitude, and voltage average of two adjacent nodes from the PMU measurement data;
[0241] The Pearson correlation coefficient is calculated using the voltage amplitude, the historical amplitude, and the voltage average.
[0242] Optionally, determining the power system state according to the Pearson correlation coefficient includes:
[0243] Determining whether the Pearson correlation coefficient is greater than a preset coefficient;
[0244] If the Pearson correlation coefficient is less than a preset coefficient, determining that the power system state is a steady state;
[0245] If the Pearson correlation coefficient is greater than a preset coefficient, it is determined that the power system state is transient.
[0246] Those skilled in the art can clearly understand that, for the sake of convenience and brevity, the specific working process of the device described above can refer to the corresponding process in the aforementioned method embodiment and will not be repeated here.
[0247] Furthermore, an embodiment of the present application also provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and runnable on the processor, wherein when the processor executes the program, the low-quality data detection method based on the density local outlier algorithm as described in the above embodiment is implemented.
[0248] Furthermore, an embodiment of the present application also provides a computer-readable storage medium, which stores a computer-executable program, and the computer-executable program is used to enable a computer to execute the low-quality data detection method based on the density local outlier algorithm as described in the above embodiment.
[0249] In the description of the embodiments of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "upper" and "lower" is based on the orientation or positional relationship shown in the accompanying drawings. It is only for the convenience of describing the embodiments of the present invention and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be understood as a limitation of the present invention. When an element such as a layer, region or substrate is referred to as being "on" or "above" another element, it can be directly on the other element, or there can be an intermediate element. In contrast, when an element is referred to as being "directly on" or "above" another element, there are no intermediate elements. It should also be understood that when an element is referred to as being "under" or "below" another element, it can be directly under or below the other element, or there can be an intermediate element. In contrast, when an element is referred to as being "directly under" or "below" another element, there are no intermediate elements. Unless otherwise specified or limited, the terms "mounted," "connected," and "connected" should be interpreted broadly. For example, they can refer to fixed, removable, or integral connections; mechanical or electrical connections; direct or indirect connections through an intermediary; and internal communication between two components. Those skilled in the art will understand the specific meanings of these terms in the present invention based on specific circumstances.
[0250] Those skilled in the art will appreciate that the embodiments of the present application may also provide computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0251] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), apparatuses and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0252] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0253] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0254] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A low-quality data detection method based on density local outlier algorithm, characterized in that: The method comprises: Obtain PMU measurement data of the power system; Calculating a Pearson correlation coefficient of the PMU measurement data, and determining a power system state based on the Pearson correlation coefficient; A plurality of outlier data points corresponding to the power system state are screened from the PMU measurement data, and low-quality data are determined according to LOF values of the outlier data points.
2. The low-quality data detection method based on the density local outlier algorithm according to claim 1 is characterized in that: The step of screening a plurality of outlier data points corresponding to the power system state from the PMU measurement data includes: If the power system state is transient, preprocessing the PMU measurement data at the same time to obtain a transient measurement data set; Constructing a transient two-dimensional plane map using the data of the transient measurement data set, and aggregating and classifying the data points in the transient two-dimensional plane map using the K-means clustering algorithm to obtain a plurality of transient classification data sets; Outlier data points are screened from each transient classification data set to obtain a plurality of outlier data points corresponding to the transient state.
3. The low-quality data detection method based on density local outlier algorithm according to claim 2, characterized in that: The constructing of a transient two-dimensional plane diagram using data from the transient measurement data set includes: Obtaining a first transient difference value and a second transient difference value of each node from the transient measurement data set, wherein the first transient difference value is a difference between data of each node at time t and data at time t-1, and the second transient difference value is a difference between data of each node at time t-1 and data at time t-2; A transient two-dimensional coordinate is generated using the first transient difference as the abscissa and the second transient difference as the ordinate, and data of the transient measurement data set are added as data points to the transient two-dimensional coordinate to obtain a transient two-dimensional plane diagram.
4. The low-quality data detection method based on density local outlier algorithm according to claim 2, characterized in that: The step of filtering outlier data points from each transient classification data set to obtain a plurality of outlier data points corresponding to the transient state includes: Determining a transient center point of each of the transient classification data sets and removing data points in each of the transient classification data sets that deviate from the transient center point to obtain a transient cleaned data set; Calculating the cluster radius of the transient data center point and the distance between each data point in the transient cleaned data set and the transient data center point to obtain a transient radius distance and a transient point distance respectively; In each transient cleaned data set, data points whose transient point distance is greater than the transient radius distance are screened to obtain a plurality of outlier data points corresponding to the transient.
5. The low-quality data detection method based on density local outlier algorithm according to claim 1, characterized in that: The step of screening a plurality of outlier data points corresponding to the power system state from the PMU measurement data includes: If the power system state is steady, pre-processing the PMU measurement data acquired by the same PMU device within a preset time period to obtain a steady-state measurement data set; constructing a steady-state two-dimensional plane graph using the data of the steady-state measurement data set, and aggregating and classifying the data points in the steady-state two-dimensional plane graph using the K-means clustering algorithm to obtain a plurality of steady-state classification data sets; Outlier data points are screened from each steady-state classification data set to obtain a plurality of outlier data points corresponding to the steady state.
6. The low-quality data detection method based on density local outlier algorithm according to claim 5, characterized in that: The step of constructing a steady-state two-dimensional plane diagram using the data of the steady-state measurement data set includes: Obtaining a steady-state difference value of each node from the steady-state measurement data set, wherein the steady-state difference value is a difference between data of each node at time t and data at time t-1; A steady-state two-dimensional coordinate is generated using the steady-state difference as the vertical coordinate and zero as the horizontal coordinate, and the data of the steady-state measurement data set are added as data points to the steady-state two-dimensional coordinate to obtain a steady-state two-dimensional plane diagram.
7. The low-quality data detection method based on density local outlier algorithm according to claim 5, characterized in that: The step of screening outlier data points from each steady-state classification data set to obtain a plurality of outlier data points corresponding to the steady state includes: determining a steady-state central point of each of the steady-state classification data sets and removing data points in each of the steady-state classification data sets that deviate from the steady-state central point to obtain a steady-state cleaned data set; Calculating the cluster radius of the steady-state data center point and the distance between each data point in the steady-state cleaned data set and the steady-state data center point to obtain a steady-state radius distance and a steady-state point distance respectively; In each of the steady-state cleaned data sets, data points whose steady-state point distance is greater than the steady-state radius distance are screened to obtain a plurality of outlier data points corresponding to the steady state.
8. The low-quality data detection method based on the density local outlier algorithm according to any one of claims 1 to 7, characterized in that: Calculating the Pearson correlation coefficient of the PMU measurement data includes: Obtaining the voltage amplitude, historical amplitude, and voltage average of two adjacent nodes from the PMU measurement data; The Pearson correlation coefficient is calculated using the voltage amplitude, the historical amplitude, and the voltage average.
9. The low-quality data detection method based on density local outlier algorithm according to any one of claims 1 to 7, characterized in that: Determining the power system state according to the Pearson correlation coefficient includes: Determining whether the Pearson correlation coefficient is greater than a preset coefficient; If the Pearson correlation coefficient is less than a preset coefficient, determining that the power system state is a steady state; If the Pearson correlation coefficient is greater than a preset coefficient, it is determined that the power system state is transient.
10. A low-quality data detection device based on density local outlier algorithm, characterized in that: The device comprises: An acquisition module is used to obtain PMU measurement data of the power system; a determination module, configured to calculate a Pearson correlation coefficient of the PMU measurement data and determine a power system state based on the Pearson correlation coefficient; The detection module is used to screen a plurality of outlier data points corresponding to the power system state from the PMU measurement data, and determine low-quality data of the PMU measurement data according to the LOF values of the outlier data points.