Implementation method of a power grid alarm data analysis system based on clustering algorithm
By applying K-Means and DBSCAN clustering algorithms in the grid alarm data analysis system, the problem of large amount of grid alarm data and difficulty in extracting key information is solved, data analysis and processing capabilities are improved, false alarm rates and false alarm rates are reduced, and the safety and stability of the power system are ensured.
Patent Information
- Application Number
- CN202210366598.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-08
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-04-08
AI Technical Summary
The power grid alarm data is large in size and difficult to read, and it is difficult to extract key information. It is difficult to effectively analyze and process massive alarm data in the existing technology, resulting in low security analysis and prediction accuracy, and a large false alarm rate and false alarm rate.
The power grid alarm data analysis system based on clustering algorithm is adopted, and the timestamps of the alarm data are clustered through the K-Means algorithm. The DBSCAN algorithm clusters the attribute keywords, obtains the event types of alarm data in different time periods, and improves the analysis and processing capabilities of the power grid alarm data.
It effectively improves the analysis and processing capabilities of power grid alarm data, reduces the difficulty of data processing, improves the accuracy and effectiveness of alarm data, reduces the false alarm rate and false alarm rate, and ensures the safety and stability of the power system.
Smart Images

Figure CN114692771B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for implementing a power grid alarm data analysis system based on a clustering algorithm, belonging to the technical field of automatic analysis of power grid data. Background Art
[0002] At present, the global cybersecurity situation is becoming increasingly severe. Personal information and business data have suffered large-scale leakage and illegal use, and malicious cyberattacks on critical information infrastructure occur frequently. Critical information infrastructure is a country's vital asset. Once damaged, malfunctioning, or data is leaked, it will not only likely lead to property losses but also seriously affect the stable operation of the economy and society. As the infrastructure in fields such as finance, energy, power, and communication becomes more and more dependent on information networks, cyberattacks on critical information infrastructure are constantly escalating.
[0003] With the continuous expansion of load-side resources, the security protection scope of power monitoring systems is constantly expanding, and access security becomes particularly important. As the security protection scope continues to expand, the generated security data is increasing exponentially, and the number of alarms is constantly increasing. How to ensure the timeliness and effectiveness of alarms and improve the accuracy and timeliness of alarms is also a major problem in security protection supervision. A large number of scholars have conducted in-depth research on intelligent alarm technologies. From the existing research results, it mainly focuses on two aspects: one is to analyze discrete abnormal data using machine learning algorithms or artificial intelligence algorithms, extract several discrete abnormal data from a large amount of normal operation data, and use typical modeling methods to analyze security events as a basis. The other is to perform hierarchical classification based on basic data such as rules, models, and reasoning analysis to obtain the characteristics of current alarm information, and to a certain extent, achieve attack prediction while analyzing faults. However, directly associating only the device-reported information with the security analysis and prediction results will result in low accuracy of security analysis and prediction, with a relatively high false alarm rate and false positive rate. Therefore, it is urgently necessary to further solve the internal relationship of power grid alarm data, start from the data itself, solve the multiple attributes of the data, analyze the correlation between massive alarm data through various data processing methods, thereby improving the ability of the power grid alarm system to process data and ensuring the safety and stability of the power system. Summary of the Invention
[0004] The object of the present invention is to propose a method for implementing a power grid alarm data analysis system based on a clustering algorithm in view of the problems of large volume, difficult to read, and difficult to extract key information of power grid alarm data. This method analyzes the original alarm log data, extracts multiple attributes, uses the K-Means algorithm to perform clustering analysis on the timestamps of alarm data, and uses the DBSCAN algorithm to cluster the attribute keywords to obtain the event types of alarm data in different time periods, which well improves the ability of the existing system to analyze and process alarm data and constructs a more efficient power grid data analysis model.
[0005] The technical solution adopted by the present invention to solve its technical problems is: a method for implementing a power grid alarm data analysis system based on a clustering algorithm. From the perspective of multiple attributes of data, this method performs multiple clusterings on the original data set. The method includes the following steps:
[0006] Step 1: Perform K-Means clustering based on the latest occurrence time of alarm data. By partitioning the attributes of the original alarm data set, obtain the latest occurrence time of the alarm. Use relevant tools to convert the latest occurrence time into the form of a timestamp, and select an appropriate value of K to perform K-Means clustering, where K represents the number of categories obtained after K-Means clustering;
[0007] Step 2: For the original data set in Step 1, obtain the start time of the alarm, also convert it into the form of a timestamp, select the optimal value of K, perform the second K-Means clustering, and correct the result of the first clustering to obtain the final K-Means clustering data;
[0008] Step 3: On the basis of Step 2, for each attribute, formulate a keyword vectorization rule, convert all data into the form of mathematical vectors, and use the DBSCAN algorithm to cluster the feature keywords to obtain the final DBSCAN clustering data;
[0009] Step 4: On the basis of Step 3, comprehensively combine the results of K-Means and DBSCAN clustering, describe the relevance of the original alarm data, and give a specific classification description of the alarm information within a specific time period.
[0010] Further, the present invention performs K-Means clustering based on the latest occurrence time of alarm data. By partitioning the attributes of the original alarm data set, obtain the latest occurrence time of the alarm. Use relevant tools to convert the latest occurrence time into the form of a timestamp, and select an appropriate value of K to perform K-Means clustering, including:
[0011] Step 1-1, Attribute partitioning: Partition the original alarm log data set, including alarm level, alarm content, alarm device, reporting device, alarm start time, latest occurrence time, alarm count, reporting status, log type, log subtype, and alarm status, so that each log data in the original data set can be concretely described using these 11 attributes;
[0012] Step 1-2, Timestamp conversion: Extract the attribute of the latest occurrence time in the data set, and use an online timestamp conversion tool to convert the latest occurrence time of all entries into the form of a 10-digit timestamp and store it correspondingly in the original file;
[0013] Step 1-3, K value selection: According to the time span of all data in the dataset, formulate the expected time span for each category after clustering. Define the accuracy rate of the clustering algorithm as the ratio of the number of categories that meet the time span to the total number of categories after clustering with a specific K value. The value range of the fixed K value is 10% to 30% of the total number of entries in the original dataset, and it is stipulated that the K value is a positive integer. Experimentally analyze the relationship between the K value and the accuracy rate, and obtain the optimal K value for a specific dataset from the relationship curve graph;
[0014] Step 1-4, K-Means clustering: Take the dataset and the selected optimal K value as inputs, and perform clustering by the K-Means algorithm to obtain the K-Means clustering result based on the latest occurrence time of the alarms.
[0015] Furthermore, in step 2 of the present invention, for the original dataset in step 1, obtain the alarm start time, also convert it into the form of a timestamp, select the optimal K value, perform the second K-Means clustering, and correct the result of the first clustering to obtain the final K-Means clustering data, including:
[0016] Step 2-1, Extract the alarm start time in the dataset, use an online tool to generate the corresponding timestamp, and experimentally analyze the relationship between the K value and the accuracy rate when clustering according to the alarm start time. Finally, determine the optimal K value;
[0017] Step 2-2, Based on the selected optimal K value, perform K-Means clustering on the dataset after converting the alarm start time to obtain the K-Means clustering result based on the alarm start time;
[0018] Step 2-3, Summarize the results of the two clusterings, use the result of the second clustering as an aid to correct the result of the first clustering, and finally obtain the optimal result after K-Means clustering.
[0019] Furthermore, in step 3 of the present invention, on the basis of step 2, for each attribute, formulate a keyword vectorization rule, convert all data into the form of mathematical vectors, and use the DBSCAN algorithm to cluster the feature keywords to obtain the final DBSCAN clustering data, including:
[0020] Step 3-1, Analyze the keywords included in each attribute for the alarm level, alarm content, alarm times, reporting status, log type, log subtype, and alarm status. For example, the keywords for the alarm level are important and unimportant, and the keywords for the alarm content are USB, port, or serial port, etc. For each attribute, formulate a keyword vectorization rule, and convert the attribute into different forms of mathematical vectors through different keywords;
[0021] Step 3-2: Select the threshold MinPts of the number of data objects in the neighborhood of the DBSCAN algorithm as the subtype category of the alarm log data. With MinPts determined, experimentally study the relationship between the neighborhood radius ε of DBSCAN and the noise rate, where the noise rate is defined as the ratio of the number of noise points that appear after clustering to the number of all entries before clustering. Noise points are entries that have not been successfully clustered. Finally, determine the optimal neighborhood radius ε.
[0022] Step 3-3: Input the vectorized alarm data set, the optimal neighborhood radius ε, and the threshold MinPts of the number of data objects in the neighborhood into DBSCAN for clustering operations to obtain the DBSCAN clustering result.
[0023] Furthermore, in step 4 of the present invention, based on step 3, the results of K-Means and DBSCAN clustering are comprehensively used to describe the relevance of the original alarm data, and a specific classification description of the alarm information within a specific time period is given, including:
[0024] Step 4-1: Construct a relevance result description template, specifically that the xth log to the xth log in a certain city occurred an x event from x time to x time.
[0025] Step 4-2: Give a relevance description of the results after K-Means and DBSCAN clustering, and segment and give the specific detailed alarm data information according to the latest occurrence time and log subtype.
[0026] Beneficial effects:
[0027] 1. The present invention starts from the alarm log data itself, and through the partitioning-based clustering algorithm K-Means and the density-based clustering algorithm DBSCAN, performs multi-dimensional feature processing and clustering operations on the original alarm data, ensuring that the dimensions of data processing are comprehensive enough. In addition, for the original alarm data, attribute partitioning is carried out from multiple dimensions, and keyword vectorization operations are performed on each attribute. Compared with other methods, the difficulty of data processing is greatly reduced.
[0028] 2. The present invention conducts multiple rounds of experiments on the key parameter K value in the K-Means algorithm and the key parameter neighborhood radius ε in the DBSCAN algorithm, and finally gives the curve graph of the K value and the accuracy of K-Means, as well as the curve graph of the neighborhood radius ε and the noise rate of DBSCAN, and determines the optimal K value and the optimal neighborhood radius ε. Compared with other methods, the clustering algorithm is more reasonable and efficient, and can analyze and study alarm data more effectively.
[0029] 3. The present invention combines the results of the K-Means and DBSCAN algorithms, formulates a detailed description of the correlation results, can segmentally describe the specific information of the alarm data, and the accuracy rate, noise rate, and compression rate all meet the expected requirements. It is also more convenient for system migration for similar data in the later stage, playing a good guiding role for subsequent analysis and research in this direction. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 is the flowchart of the method of the present invention.
[0031] Figure 2 is the structural block diagram of clustering based on the timestamp in the present invention.
[0032] Figure 3 is the structural block diagram of clustering based on keyword features in the present invention.
[0033] Figure 4 is the structural block diagram of the description of the correlation result in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] The following will disclose the embodiments of the present invention with diagrams. For the sake of clarity, many practical details will be described together in the following description. However, it should be understood that these practical details are not used to limit the present invention. That is to say, in some embodiments of the present invention, these practical details are not necessary.
[0035] As Figures 1-4 shown, the present invention provides a method for implementing a power grid alarm data analysis system based on a clustering algorithm. From the perspective of multi-attributes of data, multiple clusterings are performed on the original data set. The method includes the following steps:
[0036] Step 1: Perform K-Means clustering based on the latest occurrence time of the alarm data. By partitioning the attributes of the original alarm data set, use relevant tools to convert the latest occurrence time into the form of a timestamp, and perform K-Means clustering. As Figure 2 shown, it specifically includes:
[0037] (1) Attribute partitioning: Partition the attributes of the original alarm log data set, including alarm level, alarm content, alarm device, reporting device, alarm start time, latest occurrence time, alarm times, reporting status, log type, log subtype, and alarm status, so that each log data in the original data set can be concretely described using these 11 attributes;
[0038] (2) Timestamp conversion: Extract the attribute of the latest occurrence time in the data set, use an online timestamp conversion tool to convert the latest occurrence time of all entries into the form of a 10-digit timestamp, and store them correspondingly in the original file;
[0039] (3) Selection of K value: According to the time span of all data in the dataset, determine the expected time span for each category after clustering, and define the accuracy rate of the clustering algorithm as the ratio of the number of categories that meet the time span to the total number of categories after clustering with a specific K value. The value range of the fixed K value is 10% to 30% of the total number of entries in the original dataset, and it is stipulated that the K value is a positive integer. Experimentally analyze the relationship between the K value and the accuracy rate, and obtain the optimal K value for a specific dataset from the relationship curve graph;
[0040] (4) K-Means clustering: Take the dataset and the selected optimal K value as inputs, and perform clustering using the K-Means algorithm to obtain the K-Means clustering result based on the latest occurrence time of the alarms.
[0041] Step 2: For the original dataset in Step 1, obtain the alarm start time, also convert it into the form of a timestamp, select the optimal K value, perform the second K-Means clustering, and correct the result of the first clustering to obtain the final K-Means clustering data. Specifically:
[0042] Step 2-1: Extract the alarm start time in the dataset, use an online tool to generate the corresponding timestamp, and experimentally analyze the relationship between the K value and the accuracy rate when clustering based on the alarm start time. Finally, determine the optimal K value;
[0043] Step 2-2: Based on the selected optimal K value, perform K-Means clustering on the dataset after converting the alarm start time to obtain the K-Means clustering result based on the alarm start time;
[0044] Step 2-3: Summarize the results of the two clusterings, use the result of the second clustering as an aid to correct the result of the first clustering, and finally obtain the optimal result after K-Means clustering.
[0045] Step 3: On the basis of Step 2, for each attribute, formulate a keyword vectorization rule, convert all data into the form of mathematical vectors, and use the DBSCAN algorithm to cluster the feature keywords to obtain the final DBSCAN clustering data, such as Figure 3 , specifically including:
[0046] (3-1) For the alarm level, alarm content, alarm times, reporting status, log type, log subtype, and alarm status, analyze the keywords included in each attribute. For example, the keywords for the alarm level are important and unimportant, and the keywords for the alarm content are USB, port, or serial port, etc. For each attribute, formulate a keyword vectorization rule, and convert the attribute into different forms of mathematical vectors through different keywords;
[0047] (3-2) Select the threshold MinPts of the number of data objects in the neighborhood of the DBSCAN algorithm as the subtype category of the alarm log data. Under the condition that MinPts is determined, experimentally study the relationship between the neighborhood radius ε and the noise rate of DBSCAN, where the noise rate is defined as the ratio of the number of noise points that appear after clustering to the number of all entries before clustering. Noise points are entries that have not been successfully clustered. Finally, determine the optimal neighborhood radius ε;
[0048] (3-3) Input the vectorized alarm data set, the optimal neighborhood radius ε, and the threshold MinPts of the number of data objects in the neighborhood into DBSCAN for clustering operation to obtain the DBSCAN clustering result.
[0049] Step 4: On the basis of Step 3, comprehensively combine the results of K-Means and DBSCAN clustering to describe the relevance of the original alarm data, and give a specific classification description of the alarm information within a specific time period, such as Figure 4 , specifically including:
[0050] Step 4-1: Construct a template for describing the relevance result, specifically that the xth log to the xth log in a certain city occurred an x event from x time to x time;
[0051] Step 4-2: Describe the relevance of the results after K-Means and DBSCAN clustering. According to the latest occurrence time and the log subtype, give the specific detailed alarm data information in segments to improve the ability of the power grid alarm system to process data, and at the same time ensure the safety and stability of the power system.
[0052] Through the analysis of the power grid alarm log data, this invention extracts multi-dimensional attribute information, performs K-Means clustering and DBSCAN clustering based on the timestamp and attribute keyword features. In addition, multiple rounds of experimental studies are also carried out on the key parameters in the clustering algorithm to select the optimal parameters. At the same time, the system results are evaluated through indicators such as accuracy rate, compression rate, and noise rate. Finally, the specific detailed alarm data information is given in segments to improve the ability of the existing system to analyze and process alarm data and construct a more efficient power grid data analysis model.
[0053] The above is only the embodiment of the present invention and is not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the scope of the claims of the present invention.
Claims
1. A method for implementing a power grid alarm data analysis system based on a clustering algorithm, characterized in that: It includes the following steps: Step 1: Perform K-Means clustering based on the latest occurrence time of alarm data. By partitioning the attributes of the original alarm data set, use relevant tools to convert the latest occurrence time into the form of a timestamp, and perform K-Means clustering, where K represents the number of categories obtained after K-Means clustering; Step 2: For the original data set in Step 1, obtain the alarm start time, perform the second K-Means clustering, and correct the results of the first clustering to obtain the final K-Means clustering data, including: Step 2-1, Extract the alarm start time in the data set, analyze the relationship between the K value and the accuracy rate through multiple experiments, and finally determine the optimal K value; Step 2-2, Based on the selected optimal K value, perform K-Means clustering on the data set after converting the alarm start time to obtain the K-Means clustering result based on the alarm start time; Step 2-3, Summarize the results of the two clusterings, use the result of the second clustering as an aid to correct the result of the first clustering, and finally obtain the optimal result after K-Means clustering; Step 3: For each attribute, formulate a keyword vectorization rule, convert all data into the form of mathematical vectors, and use the DBSCAN algorithm to cluster the feature keywords to obtain the final DBSCAN clustering data; Step 4: Integrate the results of K-Means and DBSCAN clustering, describe the relevance of the original alarm data, and give a specific classification description of the alarm information within a specific time period.
2. The method for implementing a power grid alarm data analysis system based on a clustering algorithm according to claim 1, characterized in that: The said Step 1 includes: Step 1-1, Attribute partitioning: Partition the attributes of the original alarm log data set so that each log data in the original data set can be specifically described using multi-dimensional attributes; Step 1-2, Timestamp conversion: Extract the attribute of the latest occurrence time in the data set, and use an online timestamp conversion tool to convert the latest occurrence time of all entries into the form of a 10-digit timestamp; Step 1-3, K value selection: According to the time span of all data in the data set, formulate the expected time span of each category after clustering, define the accuracy rate of the clustering algorithm, fix the value range of the K value, analyze the relationship between the K value and the accuracy rate through experiments, and obtain the optimal K value for a specific data set from the relationship curve graph; Step 1-4, K-Means clustering: Use the data set and the selected optimal K value as inputs, and perform clustering by the K-Means algorithm to obtain the K-Means clustering result based on the latest occurrence time of the alarm.
3. The method for implementing a power grid alarm data analysis system based on a clustering algorithm according to claim 1, characterized in that: The said Step 3 includes: Step 3-1, For other multiple attributes, analyze the keywords included in each attribute. The keywords for the alarm level are important and unimportant, and the keywords for the alarm content are USB, port, or serial port. For each attribute, formulate a keyword vectorization rule, and convert the attribute into different forms of mathematical vectors through different keywords; Step 3-2: Select the threshold MinPts of the number of data objects in the neighborhood of the DBSCAN algorithm as the subtype category of the alarm log data. With MinPts determined, experimentally study the relationship between the neighborhood radius ε of DBSCAN and the noise rate, where the noise rate is defined as the ratio of the number of noise points that appear after clustering to the number of all entries before clustering, and the noise points are the entries that have not been successfully clustered. Finally, determine the optimal neighborhood radius ε. Step 3-3: Input the vectorized alarm data set, the optimal neighborhood radius ε, and the threshold MinPts of the number of data objects in the neighborhood into DBSCAN for clustering operation to obtain the DBSCAN clustering result.
4. The method for implementing a power grid alarm data analysis system based on a clustering algorithm according to claim 1, characterized in that: The said Step 4 includes: Step 4-1: Construct a template for describing the relevance result, specifically that the xth log to the xth log in a certain city occurred x events within the time period from x time to x time. Step 4-2: Conduct a relevance description of the results after K-Means and DBSCAN clustering, and give detailed specific information of the alarm data section by section according to the latest occurrence time and the log subtype.
Citation Information
Patent Citations
Alarm transaction extraction method, device and equipment and computer storage medium
CN111984634A
Alarm suppression method and device
CN114091704A