Method and device for detecting abnormal equipment in power grid based on kmeans
By using k-means-based grid abnormal equipment detection method in the smart grid, by constructing and judging the cluster to which the data points belong, the problems of low detection efficiency and inaccurate results in the prior art are solved, and a higher detection accuracy is achieved.
Patent Information
- Application Number
- CN202210125407.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-10
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-02-10
AI Technical Summary
The existing k-means and their improved methods have low detection efficiency and inaccurate detection results in the detection of abnormal equipment in smart grids.
By acquiring the power grid data set, select K data points with the farthest distance between each other as the cluster center, and then select K-1 nearest neighbors to form K basic clusters. Then select the K points with the farthest distance from the basic cluster to form the basic point set, calculate the similarity between other data points and points in the basic point set, and judge the cluster to which the data points belong.
The accurate detection of abnormal equipment of smart grids has been improved, more accurate clustering results have been obtained, and the accuracy of detection has been enhanced.
Smart Images

Figure CN114462538B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent power grid abnormal device detection, and particularly to a method and device for detecting abnormal power grid devices based on kmeans. Background Art
[0002] Due to the needs of informatization and intelligence, in the process of the evolution of the traditional power grid physical system to an intelligent power grid, people introduce 3C (Computing, Communication, Control) technologies to achieve self-awareness, precise control, remote collaboration, and optimal scheduling of the intelligent power grid system, making the system more flexible, efficient, economical, and intelligent. This leads to the close integration of the power grid physical system and the information system, and further makes the operating environment of the intelligent power grid system change from closed and isolated to open and interconnected. The organic integration of the intelligent power grid information system and the physical system, while improving the operating efficiency of the intelligent power grid, also provides new attack channels for attackers, making the intelligent power grid more likely to face attacks from malicious insiders or rival national competitors. A series of information security incidents in recent years have fully confirmed the vulnerability of the intelligent power grid. There is an urgent need for a new method for detecting abnormal devices in the intelligent power grid. Based on the measurement data of intelligent power grid devices, it can detect abnormal power grid devices and provide help for the intelligent power grid to defend against security attacks. The existing k-means and its improved methods simply determine the cluster to which a point belongs based on the maximum similarity between the point and the cluster center. Therefore, the existing technology has problems of low detection efficiency and inaccurate detection results for abnormal devices. Summary of the Invention
[0003] The present invention provides a method and device for detecting abnormal power grid devices based on kmeans, which improves the accurate detection of abnormal intelligent power grid devices.
[0004] An embodiment of the present invention provides a method for detecting abnormal power grid devices based on kmeans, including the following steps:
[0005] Obtain a power grid data set, and select K-1 nearest neighbors of each of the K center points of the power grid data set to form K first basic clusters; K is a positive integer;
[0006] Perform clustering on the power grid data set based on the first basic clusters to obtain M first clustering clusters; during the clustering process, sequentially determine the first basic cluster to which the first data point in the power grid data set belongs or determine that the first data point is an abnormal data point, and find the corresponding abnormal device according to the abnormal data point; M is a positive integer;
[0007] Compare the aggregation degree of the first clustering cluster obtained after each clustering with the first clustering cluster obtained after the previous clustering. When the difference in the aggregation degree between the two is less than the first preset threshold, end the clustering. When the difference is greater than or equal to the first preset threshold, construct the first basic cluster based on the first clustering cluster and re-cluster based on the first basic cluster.
[0008] Further, clustering the power grid dataset based on the first basic cluster includes the following steps:
[0009] Construct K basic point sets according to the K first basic clusters;
[0010] For the first data points in the power grid dataset, calculate the predicted selection rate of each first data point relative to each first basic cluster in turn, and record the first basic cluster with the largest predicted selection rate as the second basic cluster;
[0011] Calculate the average similarity of the second basic cluster and the average value of the second probabilities that the first data points belong to the second basic cluster relative to each basic point set. The average similarity is the average similarity between the center point of the second basic cluster and other data points in the second basic cluster;
[0012] Judge whether the first data point belongs to the second basic cluster or judge the first data point as an abnormal data point according to the average similarity and the average value of the second probabilities.
[0013] Further, judging whether the first data point is an abnormal data point according to the average similarity and the average value of the second probabilities is specifically:
[0014] Judge whether the average similarity is less than or equal to a preset multiple of the average value of the second probabilities. If so, judge that the first data point belongs to the second basic cluster; if not, judge that the first data point is an abnormal data point.
[0015] Further, constructing K basic point sets according to the K first basic clusters is specifically:
[0016] Each time, select a second data point from each of the K basic clusters, and make the distance between the second data points selected this time the largest. Combine the K second data points selected each time to form a basic point set, and select K times in total to obtain K basic point sets.
[0017] Further, calculating the predicted selection rate of each first data point relative to each first basic cluster in turn is specifically:
[0018] Each time, select a first data point from the current power grid dataset, and calculate the first probability that the first data point belongs to each first basic cluster with respect to each basic point set; the power grid dataset is the current power grid dataset obtained by deleting the data of each first basic cluster after this clustering.
[0019] According to the calculation result of the first probability, count the number of basic point sets that make the first probability the largest.
[0020] According to the first probability and the number, calculate the predicted selection rate of the first data point belonging to each first basic cluster.
[0021] Furthermore, according to the formula calculate the first probability that the first data point belongs to each first basic cluster with respect to each basic point set; in the formula, x t is the first data point, C j {j = 1, 2, …, k} is the first basic cluster, G s (s = 1, 2, …, k) is the basic point set, is the corresponding data point in the basic point set G s , represents the normalized similarity between x t and .
[0022] Furthermore, according to the formula calculate the predicted selection rate of the first data point belonging to each first basic cluster; in the formula #{p(C j |G s )|s = 1, 2, …, k} is the number of basic point sets that make the first probability the largest, k represents the number of the first basic clusters or basic point sets, p(C j |G i ) is the first probability, C j {j = 1, 2, …, k} is the first basic cluster, G s (s = 1, 2, …, k) is the basic point set.
[0023] Furthermore, according to the K center points of the power grid dataset, select K - 1 nearest neighbors of each center point to form K first basic clusters, specifically:
[0024] Generate the first basic cluster: calculate the mean point of the current power grid dataset, select a data point closest to the mean point from the current power grid dataset as the center point, then select K - 1 nearest neighbors of the center point and the center point from the current power grid dataset to form the first basic cluster, and delete the power grid data in the first basic cluster from the current power grid dataset.
[0025] Repeat the process of generating the first basic clusters until K first basic clusters are obtained.
[0026] Further, when clustering the power grid data set based on the K first basic clusters, judge the number of elements in each first basic cluster. When the number of elements in the first basic cluster is less than the second preset threshold, set each data point in the first basic cluster as a first data point, and assign the first data point to other first basic clusters or judge the first data point as an abnormal data point, and then delete the first basic cluster to obtain M first clustering clusters.
[0027] Another embodiment of the present invention provides a power grid abnormal device detection device based on kmeans, including a basic cluster construction module, an abnormal detection module, and a clustering result detection module;
[0028] The basic cluster construction module is used to obtain a power grid data set, and select K - 1 nearest neighbors of each center point according to the K center points of the power grid data set to form K first basic clusters;
[0029] The abnormal detection module is used to cluster the power grid data set based on the first basic clusters to obtain M first clustering clusters; during the clustering process, sequentially judge the first basic cluster to which the first data point in the power grid data set belongs or judge the first data point as an abnormal data point, and find the corresponding abnormal device according to the abnormal data point; K and M are positive integers;
[0030] The clustering result detection module is used to compare the aggregation degree of the first clustering clusters obtained after each clustering with the first clustering clusters obtained after the previous clustering. When the difference between the two aggregation degrees is less than the first preset threshold, end the clustering. When the difference between the two is greater than or equal to the first preset threshold, construct the first basic clusters according to the first clustering clusters and re - cluster based on the first basic clusters.
[0031] The embodiments of the present invention have the following beneficial effects:
[0032] The present invention provides a method and device for detecting abnormal power grid equipment based on kmeans. The method first selects k data points with the farthest Euclidean distances from each other as cluster centers, and then sequentially selects k - 1 nearest neighbors of each cluster center to form k basic clusters. Then, k points with the farthest distances from each other are sequentially selected from the basic clusters to form a basic point set, and the similarity between other data points and the points in the k basic point sets is calculated. Furthermore, the probability that the data point belongs to each cluster is obtained, and they are organically fused. By comprehensively considering the similarity between the point and each cluster, the cluster to which the data point belongs is determined. Existing k-means and its improved methods simply determine the cluster to which a point belongs based on the maximum similarity between the point and the cluster center. Compared with existing k-means and its improved methods, the present invention determines the cluster to which a data point belongs by comprehensively considering the similarities between the data point and multiple points in multiple clusters, and can obtain a more accurate clustering result, improving the accuracy of detecting abnormal power grid equipment. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 FIG. is a schematic flowchart of a method for detecting abnormal power grid equipment based on kmeans according to an embodiment of the present invention;
[0034] Figure 2 FIG. is a schematic structural diagram of a device for detecting abnormal power grid equipment based on kmeans according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0035] The technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0036] As Figure 1 shown, a method for detecting abnormal power grid equipment based on kmeans according to an embodiment of the present invention includes:
[0037] Step S101: Obtain a power grid data set, and select k - 1 nearest neighbors of each center point according to k center points of the power grid data set to form k first basic clusters.
[0038] As one of the embodiments, step S101 includes the following sub-steps:
[0039] Sub-step S1011: Generate the first basic cluster: Calculate the mean point of the current power grid dataset, select a data point closest to the mean point from the current power grid dataset as the center point, then select K - 1 nearest neighbors of the center point and the center point from the current power grid dataset to form the first basic cluster, and delete the power grid data in the first basic cluster from the current power grid dataset.
[0040] Sub-step S1012: Repeat the process of generating the first basic cluster until K first basic clusters are obtained.
[0041] Step S102: Cluster the power grid dataset based on the first basic cluster to obtain M first clustering clusters; during the clustering process, sequentially determine the first basic cluster to which the first data point in the power grid dataset belongs or determine that the first data point is an abnormal data point, and find the corresponding abnormal device according to the abnormal data point; K and M are positive integers.
[0042] As one of the embodiments, step S102 includes the following sub-steps:
[0043] Sub-step S1021: Construct K basic point sets according to the K first basic clusters.
[0044] As one of the embodiments, sub-step S1021 is specifically: Each time, select a second data point from each of the K basic clusters, and make the distance between the selected second data points the largest each time. Combine the K selected second data points to form a basic point set, and select K times in total to obtain K basic point sets.
[0045] Sub-step S1022: For the first data point in the power grid dataset, sequentially calculate the predicted selection rate of each first data point relative to each first basic cluster, and record the first basic cluster with the largest predicted selection rate as the second basic cluster.
[0046] As one of the embodiments, sequentially calculating the predicted selection rate of each first data point relative to each first basic cluster is specifically:
[0047] Each time, select a first data point from the current power grid dataset, and calculate the first probability that the first data point belongs to each first basic cluster relative to each basic point set; the power grid dataset is the current power grid dataset obtained after deleting the data of each first basic cluster after this clustering.
[0048] According to the calculation result of the first probability, count the number of basic point sets that make the first probability the largest.
[0049] According to the first probability and the number, calculate the predicted selection rate of the first data point belonging to each first basic cluster.
[0050] As one of the embodiments, according to the formula calculate the first probability that the first data point belongs to each first basic cluster with respect to each basic point set; where x t is the first data point, C j {j = 1, 2, …, k} is the first basic cluster, G s (s = 1, 2, …, k) is the basic point set, is the corresponding data point in the basic point set G s , represents the normalized similarity between x t and .
[0051] As one of the embodiments, according to the formula calculate the predicted selection rate that the first data point belongs to each first basic cluster; where #{p(C j |G s )|s = 1, 2, …, k} is the number of basic point sets that maximize the first probability, k represents the number of the first basic clusters or basic point sets, p(C j |G i ) is the first probability, C j {j = 1, 2, …, k} is the first basic cluster, G s (s = 1, 2, …, k) is the basic point set.
[0052] Sub-step S1023: Calculate the average similarity of the second basic cluster and the average value of the second probabilities that the first data point belongs to the second basic cluster with respect to each basic point set, where the average similarity is the average similarity between the center point of the second basic cluster and other data points in the second basic cluster.
[0053] Sub-step S1024: Determine whether the first data point belongs to the second basic cluster or determine that the first data point is an abnormal data point according to the average similarity and the average value of the second probabilities. Specifically, determine whether the average similarity is less than or equal to a preset multiple of the average value of the second probabilities. If so, determine that the first data point belongs to the second basic cluster; if not, determine that the first data point is an abnormal data point.
[0054] Step S103: Compare the aggregation degree of the first clustering cluster obtained after each clustering with the first clustering cluster obtained after the previous clustering. When the difference in their aggregation degrees is less than the first preset threshold, end the clustering. When the difference is greater than or equal to the first preset threshold, construct the first basic cluster according to the first clustering cluster and re-cluster based on the first basic cluster.
[0055] As one of the embodiments, when clustering the power grid data set based on the K first basic clusters, the number of elements in each first basic cluster is judged. When the number of elements in the first basic cluster is less than the second preset threshold, each data point in the first basic cluster is set as a first data point, and the first data point is assigned to other first basic clusters or the first data point is judged as an abnormal data point, and then the first basic cluster is deleted. The first basic clusters that have completed clustering are recorded as the first clustering clusters, that is, M first clustering clusters are obtained.
[0056] As one of the detailed embodiments, it includes the following steps:
[0057] Step A101: Denote the obtained power grid data set as X = {x 1 , x 2 , …, x N}, calculate the mean point of X T , and first select the data point with the closest Euclidean distance to the mean point in X as the center point, and then select K - 1 nearest neighbors with the closest Euclidean distance to the center point in X to form a first basic cluster BC . 1 .
[0058] After each first basic cluster is constructed, the data points identical to those in the first basic cluster are removed from the power grid data set to obtain the current power grid data set, that is, X = X - {BC 1}, and then the above process of constructing the first basic cluster is repeated until K first basic clusters are constructed.
[0059] Step A102: Select any point j from the constructed first basic clusters BC such that has the maximum sum of Euclidean distances from other points already selected from the first basic cluster. Let , and iterate this process to obtain the basic point set
[0060] Each time, select a second data point j from each of the K basic clusters BC (j = 1, 2, …, k) and make the Euclidean distances between the second data points selected this time the largest (i.e., the sum of Euclidean distances is the largest). Combine the K second data points selected each time to form a basic point set, and select K times in total to obtain K basic point sets
[0061] Construct the following matrix according to the first basic cluster and the basic point set:
[0062]
[0063] Step A103: Determine whether it is empty. If it is empty, it means that all data points in the power grid dataset have been clustered. If it is not empty, for the first data point in the power grid dataset, calculate the predicted selection rate of each of the first data points relative to each first basic cluster in turn, and record the first basic cluster with the largest predicted selection rate as the third basic cluster. The first data point is the data point selected after deleting the data points of each first basic cluster from the power grid data.
[0064] Step A1031: Select the first data point x t , calculate the similarity between x t and the points in each first basic cluster BC j (j = 1, 2, …, k), denoted as where 1 ≤ s ≤ k. Normalize the calculated similarity, that is Denote the first basic cluster BC j {j = 1, 2, …, k} as C j .
[0065] Step A1032: Based on the similarity between x t and the points in the basic point set G s (s = 1, 2, …, k), calculate the normalized similarity of x t belonging to the first basic cluster C j {j = 1, 2, …, k}, that is, calculate the first probability of the first data point belonging to each first basic cluster relative to each basic point set The calculation results are shown in the following table:
[0066] <![CDATA[C 1 > <![CDATA[C 2 > … <![CDATA[C j > … <![CDATA[C k > <![CDATA[G 1 > <![CDATA[p(C 1 |G 1 )]]> <![CDATA[p(C 2 |G 1 )]]> … <![CDATA[p(C j |G 1 )]]> <![CDATA[p(C k |G 1 )]]> <![CDATA[G 2 > <![CDATA[p(C 1 |G 2 )]]> <![CDATA[p(C 2 |G 2 )]]> <![CDATA[p(C j |G 2 )]]> <![CDATA[p(C k |G 2 )]]> <![CDATA[G s > <![CDATA[p(C 1 |G s )]]> <![CDATA[p(C 2 |G s )]]> <![CDATA[p(C j |G s )]]> … <![CDATA[p(C k |G s )]]> … … … … … … <![CDATA[G k > <![CDATA[p(C 1 |G k )]]> <![CDATA[p(C 2 |G k )]]> … <![CDATA[p(C j |G k )]]> … <![CDATA[p(C k |G k )]]>
[0067] In the formula, represents the normalized similarity between x t and .
[0068] If the first probability (i.e., the normalized similarity) of the first data point x t belonging to the basic point set G s and belonging to the first basic cluster C 1 is calculated as: 0.125, 0.25, 0.375, 0.5, 0.625, 0.75, 0.875, then the first probability (i.e., the normalized similarity) of the first data point x t belonging to the basic point set G s and belonging to the cluster C 1 is calculated as: Similarly, for each basic point set Gs (s = 1, 2, …, k) for each first basic cluster C j {j = 1, 2, …, k}, calculate and calculate #{p(C j |G s )|s = 1, 2, …, k} is the number of the basic point sets that maximize the said first probability.
[0069] Step A1033: According to the formula calculate the prediction selection rate ps(C t ) of the said first data point x belonging to each first basic cluster. j )
[0070] For example, the first probability (i.e., the normalized similarity) of the first data point relative to the basic point set G s (s = 1, 2, 3) belonging to the first basic cluster C j {j = 1, 2, 3} is shown in the following table:
[0071] <![CDATA[C 1 > <![CDATA[C 2 > <![CDATA[C 3 > Maximum value <![CDATA[G 1 > 0.1000 0.4000 0.5000 0.500 <![CDATA[G 2 > 0.1538 0.3846 0.4615 0.4615 <![CDATA[G 3 > 0.1875 0.4375 0.3750 0.4375
[0072] Then p(C 1 ) = 0;
[0073] ps(C 1 ) = 0,
[0074] Denote the first basic cluster with the maximum prediction selection rate as the second basic cluster.
[0075] Step A104: Calculate the average similarity between the center point of the said third basic cluster and other points, denoted as Calculate the average value As of the second probabilities that the first data point belongs to the said third basic cluster relative to each basic point set, that is, calculate the average value of the normalized similarities between the first data point and the points in the second basic cluster as the second basic cluster.
[0076] Step A105: Judge whether the first data point belongs to the second basic cluster or judge that the first data point is an abnormal data point according to the average similarity and the average value of the second probabilities. Specifically, when judging whether the average similarity is less than or equal to a preset multiple of the average value of the second probabilities, if so, judge that the first data point belongs to the second basic cluster; if not, judge that the first data point is an abnormal data point. That is, if holds, then the first data point belongs to the second basic cluster, if not, then x t is an abnormal data point.
[0077] Step A106: After all the first data points are judged, the number of elements in each first basic cluster obtained by clustering is judged. When the number of elements in the first basic cluster is less than the second preset threshold, each data point of the first basic cluster is set as a first data point, and the first data point is assigned to another first basic cluster or the first data point is judged as an abnormal data point, and then the first basic cluster is deleted. The remaining first basic clusters are recorded as first cluster clusters, that is, M first cluster clusters are obtained.
[0078] Step A107: Compare the degree of clustering of the first cluster obtained after each clustering with the first cluster obtained after the previous clustering. When the difference in the degree of clustering between the two is less than a first preset threshold, terminate the clustering. When the difference between the two is greater than or equal to the first preset threshold, construct a first basic cluster based on the first cluster, and re-cluster based on the first basic cluster.
[0079] Specifically, for the M first clusters C obtained by clustering t (1≤t≤m), according to the formula Calculate the degree of clustering, and judge whether the difference between the degree of clustering of the first cluster obtained after this clustering and the first cluster obtained after the previous clustering is less than the first preset threshold according to the formula WSS-WSS′<α; if so, end the clustering, if not, then cluster the M first clusters C obtained by clustering t (1≤t≤m), calculate its mean point The first cluster C t Zhongyu The point with the closest Euclidean distance is selected as the cluster center, and then the m-1 points with the closest Euclidean distance to the cluster center are selected as neighbors in the corresponding first cluster cluster to form m first basic clusters. Steps A102-A106 are repeated until WSS-WSS′<α, and clustering is terminated; where WSS′ is the degree of aggregation of the first cluster cluster obtained after the last clustering, and WSS is the degree of aggregation of the first cluster cluster obtained after this clustering. The basic cluster is the cluster before clustering, and the cluster cluster is the cluster after clustering is completed.
[0080] In the embodiments of the present invention, the similarity between a data point and multiple points of multiple clusters is integrated to determine the cluster to which the data point belongs, and a relatively accurate clustering result can be obtained, improving the accuracy of detecting abnormal power grid equipment. In the embodiments of the present invention, k data points with the farthest Euclidean distances from each other are sequentially selected as cluster centers, and then k - 1 nearest neighbors of each cluster center are sequentially selected to form k basic clusters. Then, k points with the farthest distances from each other are sequentially selected from the basic clusters to form a basic point set, and then the similarity between other data points and the points in the k basic point sets is calculated. Furthermore, the probability that the data point belongs to each cluster is obtained, and they are organically integrated. By synthesizing the similarity between the point and each cluster, the cluster to which the data point belongs is determined. In the existing k-means and its improved methods, only the maximum similarity between a point and the cluster center is simply used to determine the cluster to which the point belongs. Compared with the existing k-means and its improved methods, in the embodiments of the present invention, the similarity between a data point and multiple points of multiple clusters is integrated to determine the cluster to which the point belongs, and the cluster to which the data point belongs can be determined more accurately. At the same time, more complex clusters, such as spherical clusters with different densities, can also be discovered, and thus abnormal equipment can be accurately detected.
[0081] As Figure 2 shown, another embodiment of the present invention provides a power grid abnormal equipment detection device based on kmeans, including a basic cluster construction module, an abnormal detection module, and a clustering result detection module;
[0082] The basic cluster construction module is used to obtain a power grid data set, and according to the K central points of the power grid data set, select K - 1 nearest neighbors of each central point to form K first basic clusters;
[0083] The abnormal detection module is used to cluster the power grid data set based on the first basic clusters to obtain M first clustering clusters; during the clustering process, it is sequentially determined whether the first data point in the power grid data set belongs to the first basic cluster or determines that the first data point is an abnormal data point, and the corresponding abnormal equipment is found according to the abnormal data point; K and M are positive integers;
[0084] The clustering result detection module is used to compare the aggregation degree of the first clustering clusters obtained after each clustering with the first clustering clusters obtained after the previous clustering. When the difference between their aggregation degrees is less than a first preset threshold, the clustering ends. When the difference is greater than or equal to the first preset threshold, the first basic clusters are constructed according to the first clustering clusters, and clustering is restarted based on the first basic clusters.
[0085] For the convenience and conciseness of description, the power grid abnormal equipment detection device based on kmeans in the embodiments of this device includes all the implementation manners in the embodiments of the power grid abnormal equipment detection method based on kmeans described above, which will not be elaborated here.
[0086] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided by the present invention, the connection relationships between the modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines. Those of ordinary skill in the art can understand and implement without creative efforts.
[0087] The above is the preferred embodiment of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present invention.
[0088] Those of ordinary skill in the art can understand that to implement all or part of the processes in the above embodiments, it can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above embodiments. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
Claims
1. A method for detecting abnormal devices in a power grid based on kmeans, characterized in that, it includes the following steps: Obtain a power grid data set, and select K-1 nearest neighbors of each of the K central points of the power grid data set to form K first basic clusters; K is a positive integer; Cluster the power grid data set based on the first basic clusters to obtain M first clustering clusters; during the clustering process, sequentially calculate the first basic cluster to which the first data point in the power grid data set belongs or determine that the first data point is an abnormal data point, and find the corresponding abnormal device according to the abnormal data point; M is a positive integer; Compare the aggregation degree of the first clustering cluster obtained after each clustering with the first clustering cluster obtained after the previous clustering. When the difference in the aggregation degree between the two is less than the first preset threshold, the clustering ends. When the difference between the two is greater than or equal to the first preset threshold, construct the first basic cluster according to the first clustering cluster, and re-cluster based on the first basic cluster; Among them, clustering the power grid data set based on the first basic cluster specifically means: constructing K basic point sets according to the K first basic clusters; for the first data point in the power grid data set, sequentially calculate the prediction selection rate of each of the first data points relative to each first basic cluster, and record the first basic cluster with the largest prediction selection rate as the second basic cluster; calculate the average similarity of the second basic cluster and the average value of the second probability that the first data point belongs to the second basic cluster relative to each basic point set, where the average similarity is the average similarity between the central point of the second basic cluster and other data points in the second basic cluster; determine whether the first data point belongs to the second basic cluster or determine that the first data point is an abnormal data point according to the average similarity and the average value of the second probability; Among them, sequentially calculating the prediction selection rate of each of the first data points relative to each first basic cluster specifically means: each time, select a first data point from the current power grid data set, and calculate the first probability that the first data point belongs to each first basic cluster relative to each basic point set; the power grid data set is the current power grid data set obtained by deleting the data of each first basic cluster after this clustering; according to the calculation result of the first probability, count the number of basic point sets that make the first probability the largest; calculate the prediction selection rate of the first data point belonging to each first basic cluster according to the first probability and the number; Among them, constructing K basic point sets according to the K first basic clusters specifically means: each time, select a second data point from each of the K basic clusters, and make the distance between the second data points selected this time the largest. Combine the K second data points selected each time to form a basic point set, and select K times in total to obtain K basic point sets.
2. The method for detecting abnormal devices in a power grid based on kmeans according to claim 1, characterized in that, Determining whether the first data point is an abnormal data point according to the average similarity and the average value of the second probability specifically means: Determine whether the average similarity is less than or equal to a preset multiple of the average value of the second probability. If so, determine that the first data point belongs to the second basic cluster; if not, determine that the first data point is an abnormal data point.
3. The kmeans-based power grid abnormal equipment detection method according to claim 1, characterized in that According to the formula , calculate the first probability that the first data point belongs to each first basic cluster with respect to each basic point set; in the formula is the first data point,[[]] is the first basic cluster,[[]] is the basic point set,[[]] is the corresponding data point in the basic point set , and represents and the normalized similarity between them.
4. The kmeans-based power grid abnormal equipment detection method according to claim 3, characterized in that , Calculate the predicted selection rate of the first data point belonging to each first basic cluster according to the formula ; where = , is the number of the basic point sets that maximize the first probability, k represents the number of the first basic clusters or basic point sets, is the first probability, is the first basic cluster, is the basic point set.
5. The kmeans-based power grid abnormal equipment detection method according to claim 4, characterized in that Based on the K center points of the power grid data set, select K-1 nearest neighbors of each center point to form K first basic clusters, specifically: Generate the first basic cluster: Calculate the mean point of the current power grid data set, select a data point closest to the mean point from the current power grid data set as the center point, and then select K-1 nearest neighbors of the center point from the current power grid data set and form the first basic cluster together with the center point, and delete the power grid data in the first basic cluster from the current power grid data set; Repeat the process of generating the first basic cluster until K first basic clusters are obtained.
6. The kmeans-based power grid abnormal equipment detection method according to any one of claims 1 to 5, characterized in that When clustering the power grid data set based on the K first basic clusters, judge the number of elements in each first basic cluster. When the number of elements in the first basic cluster is less than the second preset threshold, set each data point in the first basic cluster as the first data point, and allocate the first data point to other first basic clusters or judge the first data point as an abnormal data point, and then delete the first basic cluster to obtain M first clustering clusters.
7. A kmeans-based power grid abnormal equipment detection device, characterized in that comprises a basic cluster construction module, an abnormal detection module and a clustering result detection module; The basic cluster construction module is used to obtain the power grid data set, and based on the K center points of the power grid data set, select K-1 nearest neighbors of each center point to form K first basic clusters; The abnormal detection module is used to cluster the power grid data set based on the first basic cluster to obtain M first clustering clusters; during the clustering process, sequentially judge the first basic cluster to which the first data point in the power grid data set belongs or judge the first data point as an abnormal data point, and find the corresponding abnormal equipment according to the abnormal data point; K and M are positive integers; The clustering result detection module is used to compare the aggregation degree of the first clustering cluster obtained after each clustering with the first clustering cluster obtained after the previous clustering. When the difference in the aggregation degree between the two is less than the first preset threshold, end the clustering. When the difference between the two is greater than or equal to the first preset threshold, construct the first basic cluster according to the first clustering cluster and re-cluster based on the first basic cluster; Among them, clustering the power grid data set based on the first basic cluster specifically includes: constructing K basic point sets according to the K first basic clusters; for the first data points in the power grid data set, successively calculating the predicted selection rate of each first data point relative to each first basic cluster, and recording the first basic cluster with the largest predicted selection rate as the second basic cluster; calculating the average similarity of the second basic cluster and the average value of the second probabilities that the first data points belong to the second basic cluster relative to each basic point set, where the average similarity is the average similarity between the center point of the second basic cluster and other data points in the second basic cluster; judging whether the first data point belongs to the second basic cluster or judging that the first data point is an abnormal data point according to the average similarity and the average value of the second probabilities. Among them, successively calculating the predicted selection rate of each first data point relative to each first basic cluster specifically includes: each time selecting a first data point from the current power grid data set, and calculating the first probability that the first data point belongs to each first basic cluster relative to each basic point set; the power grid data set is the current power grid data set obtained after deleting the data of each first basic cluster after this clustering; according to the calculation results of the first probabilities, counting the number of basic point sets that make the first probabilities the largest; calculating the predicted selection rate of the first data point belonging to each first basic cluster according to the first probabilities and the number. Among them, constructing K basic point sets according to the K first basic clusters specifically includes: each time selecting a second data point from each of the K basic clusters, and making the distance between the second data points selected this time the largest, combining the K second data points selected each time to form a basic point set, and selecting K times in total to obtain K basic point sets.
Citation Information
Patent Citations
Method and system for reducing inter-symbol interference effects in transmission over a serial link with mapping of each word in a cluster of received words to a single transmitted word
CA2454452A1
Load curve morphological clustering algorithm based on improved kmeans
CN110796173A