Power monitoring system protocol classification method, device and equipment and readable storage medium

Through a convolutional autoencoder, the protocol data in the network traffic of the power monitoring system is extracted and similarity analysis is performed, and the number of similar protocol data and thresholds are clustered, which solves the problem of low classification accuracy in traditional methods and achieves more efficient protocol data classification.

CN120541559APending Publication Date: 2025-08-26ELECTRIC POWER RES INST CHINA SOUTHERN POWER GRID CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510711113.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

The traditional power monitoring system protocol classification method has the problem of low classification accuracy.

Method used

The convolutional autoencoder is used to collect protocol data to be classified from the network traffic of the power monitoring system. Through feature extraction and similarity analysis, the number of similar protocol data and preset thresholds are used for clustering to realize the classification of protocol data.

Benefits of technology

It improves the classification accuracy of protocol data, adapts to the classification needs of unstructured message data, and does not rely on the selection of the initial center, improving the accuracy and speed of classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120541559A_ABST
    Figure CN120541559A_ABST
Patent Text Reader

Abstract

The invention relates to a power monitoring system protocol classification method, device and equipment and a readable storage medium. The method comprises the steps of collecting multiple pieces of to-be-classified protocol data from network traffic of a current monitoring system, inputting each piece of to-be-classified protocol data into a convolutional auto-encoder for feature extraction to obtain protocol features corresponding to each piece of to-be-classified protocol data, obtaining similarity degrees among the to-be-classified protocol data based on each protocol feature, and determining the similarity degrees of the to-be-classified protocol data according to the similarity degrees. On the basis of the similarity degree, the number of similar protocol data corresponding to the to-be-classified protocol data is obtained, and the similar protocol data are other to-be-classified protocol data with the similarity degree with the current to-be-classified protocol data larger than or equal to a preset similarity degree threshold value; and finally, based on the quantity of the similar protocol data corresponding to the to-be-classified protocol data and a preset protocol quantity threshold, clustering the to-be-classified protocol data, and obtaining a classification result according to a clustering result. By adopting the method, the classification accuracy of the protocol data can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of network security for power monitoring systems, and in particular to a method, apparatus, computer equipment, computer-readable storage medium, and computer program product for classifying protocols for power monitoring systems. Background Art

[0002] Power monitoring systems widely use a variety of communication protocols. With system expansion and the diversification of equipment manufacturers, new or variant protocols continue to emerge. Traditional protocol classification methods achieve this by selecting an accurate initial center and removing the influence of noise.

[0003] However, the current traditional protocol classification method has the problem of low classification accuracy. Summary of the Invention

[0004] Based on this, it is necessary to provide a power monitoring system protocol classification method, device, computer equipment, computer-readable storage medium and computer program product that can improve classification accuracy in response to the above technical problems.

[0005] In a first aspect, the present application provides a method for classifying protocols of a power monitoring system, including:

[0006] Collecting multiple protocol data to be classified from network traffic of a current monitoring system;

[0007] Input each protocol data to be classified into the trained convolutional autoencoder for feature extraction to obtain the protocol features corresponding to each protocol data to be classified;

[0008] Based on the characteristics of each protocol, the degree of similarity between each protocol data to be classified is obtained, and based on the degree of similarity, the number of similar protocol data corresponding to each protocol data to be classified is obtained; the similar protocol data is other protocol data to be classified whose degree of similarity with the current protocol data to be classified is greater than or equal to a preset similarity threshold; the current protocol data to be classified is any one of the protocol data to be classified, and the other protocol data to be classified is the protocol data to be classified in the protocol data to be classified except the current protocol data to be classified;

[0009] Based on the number of similar protocol data corresponding to each protocol data to be classified and a preset protocol number threshold, the protocol data to be classified are clustered, and a classification result of each protocol data to be classified is obtained according to the clustering result.

[0010] In one embodiment, clustering the protocol data to be classified based on the number of similar protocol data corresponding to each protocol data to be classified and a preset protocol number threshold includes:

[0011] Based on the number of similar protocol data corresponding to each protocol data to be classified and a preset protocol number threshold, a density distribution type corresponding to each protocol data to be classified is obtained; the density distribution type includes: core protocol type, boundary protocol type and noise protocol type;

[0012] Randomly obtain a target protocol data from each protocol data to be classified, and use the target protocol data as the cluster center of the cluster set corresponding to the target protocol data; wherein the target protocol data is the protocol data to be classified that has not been clustered and the corresponding density distribution type is the core protocol type;

[0013] Adding similar protocol data corresponding to the target protocol data into the cluster set, and returning to the step of randomly obtaining a target protocol data from each protocol data to be classified, until the target protocol data is no longer included in the protocol data to be classified;

[0014] The protocol data to be classified whose density distribution type is the noise protocol type is added to the noise cluster set.

[0015] In one embodiment, adding similar protocol data corresponding to the target protocol data into a cluster set includes:

[0016] Obtain current similar protocol data; the current similar protocol data is any one of the similar protocol data corresponding to the target protocol data;

[0017] When the density distribution type of the current similar protocol data is a boundary protocol type, the current similar protocol data is added to the cluster set;

[0018] When the density distribution type of the current similar protocol data is the core protocol type, the current similar protocol data is added to the cluster set, and the similar protocol data corresponding to the current similar protocol data is added as the similar protocol data corresponding to the target protocol data.

[0019] In one embodiment, similar protocol data corresponding to the target protocol data is added to the cluster set, and the step of randomly obtaining a target protocol data from each protocol data to be classified is returned to be executed until the target protocol data is no longer included in the protocol data to be classified, further comprising:

[0020] Obtaining the remaining protocol data to be classified; the density distribution type of the remaining protocol data to be classified is a boundary protocol type;

[0021] Obtaining the similarity between the remaining protocol data to be classified and the target protocol data as the cluster center of each cluster set;

[0022] The remaining protocol data to be classified are added to the cluster set corresponding to the target protocol data with the highest similarity.

[0023] In an exemplary embodiment, the protocol quantity threshold includes a first protocol quantity threshold and a second protocol quantity threshold, and the first protocol quantity threshold is greater than the second protocol quantity threshold; based on the quantity of similar protocol data corresponding to each to-be-classified protocol data and a preset protocol quantity threshold, obtaining the density distribution type corresponding to each to-be-classified protocol data includes:

[0024] When the number of similar protocol data corresponding to the current protocol data to be classified is greater than or equal to the first protocol number threshold, determining that the density distribution type of the current protocol data to be classified is a core protocol type;

[0025] When the number of similar protocol data corresponding to the current protocol data to be classified is less than the first protocol number threshold and greater than the second protocol number threshold, determining that the density distribution type of the current protocol data to be classified is a boundary protocol type;

[0026] When the number of similar protocol data corresponding to the current protocol data to be classified is less than or equal to the second protocol number threshold, it is determined that the density distribution type of the current protocol data to be classified is a noise protocol type.

[0027] In one embodiment, the trained convolutional autoencoder includes a convolutional layer and a pooling layer;

[0028] Each protocol data to be classified is input into the trained convolutional autoencoder for feature extraction to obtain the protocol features corresponding to each protocol data to be classified, including:

[0029] Input each protocol data to be classified into the convolution layer for feature extraction to obtain the feature maps corresponding to the protocol data to be classified;

[0030] Each feature map is input into the pooling layer for feature dimensionality reduction to obtain the protocol features corresponding to each protocol data to be classified.

[0031] In one embodiment, the convolutional autoencoder is trained by the following steps:

[0032] Obtain multiple sample protocol data;

[0033] Input each sample protocol data into the convolutional autoencoder to be trained for feature extraction, and obtain the sample protocol features corresponding to each sample protocol data;

[0034] Input each sample protocol feature into the decoder corresponding to the convolutional autoencoder to be trained for data reconstruction, and obtain the reconstructed protocol data corresponding to each sample protocol data;

[0035] According to the sample protocol data and the reconstruction protocol data, the parameters of the convolutional autoencoder to be trained are trained to obtain a trained convolutional autoencoder.

[0036] In a second aspect, the present application further provides a power monitoring system protocol classification device, comprising:

[0037] A data acquisition module is used to collect a plurality of protocol data to be classified from the network traffic of the current monitoring system;

[0038] The feature extraction module is used to input each protocol data to be classified into the trained convolutional autoencoder for feature extraction, and obtain the protocol features corresponding to each protocol data to be classified;

[0039] A similarity acquisition module is configured to acquire the similarity between each protocol data to be classified based on each protocol feature, and based on the similarity, acquire the number of similar protocol data corresponding to each protocol data to be classified; similar protocol data is other protocol data to be classified whose similarity to the current protocol data to be classified is greater than or equal to a preset similarity threshold; the current protocol data to be classified is any one of the protocol data to be classified, and the other protocol data to be classified is the protocol data to be classified other than the current protocol data to be classified in the protocol data to be classified;

[0040] The data classification module is used to cluster the protocol data to be classified based on the number of similar protocol data corresponding to each protocol data to be classified and a preset protocol number threshold, and obtain the classification result of each protocol data to be classified according to the clustering result.

[0041] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0042] Collecting multiple protocol data to be classified from network traffic of a current monitoring system;

[0043] Input each protocol data to be classified into the trained convolutional autoencoder for feature extraction to obtain the protocol features corresponding to each protocol data to be classified;

[0044] Based on the characteristics of each protocol, the degree of similarity between each protocol data to be classified is obtained, and based on the degree of similarity, the number of similar protocol data corresponding to each protocol data to be classified is obtained; the similar protocol data is other protocol data to be classified whose degree of similarity with the current protocol data to be classified is greater than or equal to a preset similarity threshold; the current protocol data to be classified is any one of the protocol data to be classified, and the other protocol data to be classified is the protocol data to be classified in the protocol data to be classified except the current protocol data to be classified;

[0045] Based on the number of similar protocol data corresponding to each protocol data to be classified and a preset protocol number threshold, the protocol data to be classified are clustered, and a classification result of each protocol data to be classified is obtained according to the clustering result.

[0046] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the following steps:

[0047] Collecting multiple protocol data to be classified from network traffic of a current monitoring system;

[0048] Input each protocol data to be classified into the trained convolutional autoencoder for feature extraction to obtain the protocol features corresponding to each protocol data to be classified;

[0049] Based on the characteristics of each protocol, the degree of similarity between each protocol data to be classified is obtained, and based on the degree of similarity, the number of similar protocol data corresponding to each protocol data to be classified is obtained; the similar protocol data is other protocol data to be classified whose degree of similarity with the current protocol data to be classified is greater than or equal to a preset similarity threshold; the current protocol data to be classified is any one of the protocol data to be classified, and the other protocol data to be classified is the protocol data to be classified in the protocol data to be classified except the current protocol data to be classified;

[0050] Based on the number of similar protocol data corresponding to each protocol data to be classified and a preset protocol number threshold, the protocol data to be classified are clustered, and a classification result of each protocol data to be classified is obtained according to the clustering result.

[0051] In a fifth aspect, the present application further provides a computer program product, comprising a computer program, which, when executed by a processor, implements the following steps:

[0052] Collecting multiple protocol data to be classified from network traffic of a current monitoring system;

[0053] Input each protocol data to be classified into the trained convolutional autoencoder for feature extraction to obtain the protocol features corresponding to each protocol data to be classified;

[0054] Based on the characteristics of each protocol, the degree of similarity between each protocol data to be classified is obtained, and based on the degree of similarity, the number of similar protocol data corresponding to each protocol data to be classified is obtained; the similar protocol data is other protocol data to be classified whose degree of similarity with the current protocol data to be classified is greater than or equal to a preset similarity threshold; the current protocol data to be classified is any one of the protocol data to be classified, and the other protocol data to be classified is the protocol data to be classified in the protocol data to be classified except the current protocol data to be classified;

[0055] Based on the number of similar protocol data corresponding to each protocol data to be classified and a preset protocol number threshold, the protocol data to be classified are clustered, and a classification result of each protocol data to be classified is obtained according to the clustering result.

[0056] The above-mentioned power monitoring system protocol classification method, device, computer equipment, computer-readable storage medium and computer program product collect multiple protocol data to be classified from the network traffic of the current monitoring system, input each protocol data to be classified into the trained convolutional autoencoder for feature extraction, obtain the protocol features corresponding to each protocol data to be classified, obtain the similarity between each protocol data to be classified based on each protocol feature, and obtain the number of similar protocol data corresponding to each protocol data to be classified based on the similarity, where the similar protocol data is other protocol data to be classified whose similarity with the current protocol data to be classified is greater than or equal to a preset similarity threshold, the current protocol data to be classified is any one of the protocol data to be classified, and the other protocol data to be classified is the protocol data to be classified in each protocol data to be classified except the current protocol data to be classified. Finally, based on the number of similar protocol data corresponding to each protocol data to be classified and a preset protocol number threshold, the protocol data to be classified are clustered, and the classification results of each protocol data to be classified are obtained according to the clustering results. By acquiring the protocol data to be classified in real time and analyzing the similarity of the protocol data to be classified, the protocol data to be classified are finally clustered according to the number of similar protocol data to be classified and the preset number threshold. This clustering method does not rely on the selection of the initial center, is suitable for the classification needs of unstructured message data, and improves the classification accuracy of protocol data. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments of the present application or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying any creative work.

[0058] Figure 1 This is an application environment diagram of a power monitoring system protocol classification method in one embodiment;

[0059] Figure 2 1 is a flow chart of a method for classifying protocols of a power monitoring system according to an embodiment;

[0060] Figure 3 1 is a flow chart of a method for classifying protocols of a power monitoring system according to another embodiment;

[0061] Figure 4A structural block diagram of a protocol classification device for a power monitoring system according to an embodiment;

[0062] Figure 5 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0063] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0064] The power monitoring system protocol classification method provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown, the power monitoring system includes multiple power equipment terminals. Power equipment terminal i102 can communicate with server 104 via a network, using a wide range of communication protocols. A data storage system can store data that server 104 needs to process. The data storage system can be integrated with server 104, or it can be located in the cloud or on other network servers.

[0065] The server 104 collects a plurality of protocol data to be classified from the network traffic of the current monitoring system, inputs each protocol data to be classified into the trained convolutional autoencoder for feature extraction, obtains the protocol features corresponding to each protocol data to be classified, obtains the similarity between each protocol data to be classified based on the protocol features, and obtains the number of similar protocol data corresponding to each protocol data to be classified based on the similarity, wherein the similar protocol data are other protocol data to be classified whose similarity with the current protocol data to be classified is greater than or equal to a preset similarity threshold, the current protocol data to be classified is any one of the protocol data to be classified, and the other protocol data to be classified are the protocol data to be classified other than the current protocol data to be classified in the protocol data to be classified, and finally, based on the number of similar protocol data corresponding to each protocol data to be classified and a preset protocol number threshold, the protocol data to be classified are clustered, and the classification results of each protocol data to be classified are obtained according to the clustering results.

[0066] The power equipment terminal i102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart car devices, and projectors. Portable wearable devices can include smart watches, smart bracelets, and head-mounted devices. Head-mounted devices can include virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, and the like. The server 104 can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server providing cloud computing services.

[0067] In an exemplary embodiment, Figure 2 As shown, a power monitoring system protocol classification method is provided, which is applied to Figure 1 The server 104 in the example is used as an example to illustrate the process, including the following steps S201 to S204.

[0068] Step S201 : collecting a plurality of protocol data to be classified from the network traffic of the current monitoring system.

[0069] In step S202, each protocol data to be classified is input into the trained convolutional autoencoder for feature extraction to obtain protocol features corresponding to each protocol data to be classified.

[0070] Among them, the network traffic of the current monitoring system can be understood as the protocol data stream flowing through the communication channel after the communication channel is established between the power terminal equipment and the server. The protocol data stream is composed of multiple communication protocols; the protocol data to be classified can be understood as the protocol messages in the network traffic, including IEC 60870-5-104, DNP3, private protocols and other types of protocol messages. The type of protocol message is unknown during collection; the protocol feature can be understood as a compact expression form of the protocol data to be classified, with more concise content.

[0071] Optionally, the server 104 obtains a plurality of protocol data to be classified from the protocol data stream flowing through the communication channel established with the power equipment terminal i. Assume that the protocol data stream is ,in Represents a protocol message for a power monitoring system. To adapt the deep learning model to the analysis of power protocol traffic, the message data needs to be converted into a fixed-length tensor representation. Assuming the maximum length of the protocol data is L, the protocol data can be represented as a matrix (i.e., the aforementioned protocol data to be classified): Among them, each protocol message After padding or truncation, the length is unified to L.

[0072] The matrix is ​​used as the input of the trained convolutional autoencoder to extract protocol features. The convolutional layer of the convolutional autoencoder is used to extract the feature map of the matrix. The feature map is then input into the pooling layer of the convolutional autoencoder to perform aggregation operations on the local area of ​​the feature map and compress it into a single value, thereby reducing the size of the feature map and obtaining the protocol features corresponding to each protocol data to be classified.

[0073] Based on the above implementation method, by converting the protocol data to be classified into a unified tensor representation, the data length of different types of protocol data is guaranteed to be unified, the difficulty of feature extraction is reduced, and the convolutional autoencoder is used for feature extraction to ensure the stability and compactness of the extracted protocol features, thereby accelerating the classification speed and real-time performance of the protocol data.

[0074] Step S203: Based on the characteristics of each protocol, obtain the similarity between each protocol data to be classified, and based on the similarity, obtain the number of similar protocol data corresponding to each protocol data to be classified; similar protocol data are other protocol data to be classified whose similarity with the current protocol data to be classified is greater than or equal to a preset similarity threshold; the current protocol data to be classified is any one of the protocol data to be classified, and the other protocol data to be classified are the protocol data to be classified in the protocol data to be classified except the current protocol data to be classified.

[0075] The similarity degree can be understood as the similarity of data content between the protocol data to be classified, and the similar protocol data can be understood as non-current classification protocol data whose similarity with the current classification protocol data meets preset conditions.

[0076] Exemplarily, the server 104 calculates the similarity between each protocol data to be classified based on the characteristics of each protocol, and for any current protocol data to be classified, determines from the non-current protocol data to be classified other protocol data to be classified whose similarity with the current protocol data to be classified is greater than or equal to a preset similarity threshold, and uses the other protocol data to be classified as similar protocol data to the current protocol data to be classified, and obtains the number of similar data corresponding to each protocol data to be classified.

[0077] Among them, since the communication protocols used in the power monitoring system have different data structures and feature distributions, and step S201 uses a convolutional autoencoder to extract features, each protocol data point All have been converted into one -dimensional feature vector: .in, Represents the characteristic dimension of the protocol data, N is the total number of data points, and calculates every two protocol data to be classified , The distance between them is used to determine their similarity. The Euclidean distance is used to measure the similarity between the protocol data to be classified:

[0078]

[0079] It is determined by the K-distance plot, that is, calculating the k-nearest neighbor distance of each point and drawing a curve. The neighborhood radius is selected according to the inflection point, and the neighborhood radius is used to compare the similarity between the protocol data to be classified represented by the Euclidean distance. The closer the Euclidean distance, the higher the similarity between the protocol data to be classified. The protocol data to be classified whose Euclidean distance is less than or equal to the neighborhood radius is regarded as the similar protocol data of the current protocol data to be classified, and the corresponding number of similar protocol data is obtained.

[0080] Based on the above-mentioned embodiment, the Euclidean distance is used to characterize the similarity between the protocol data to be classified, which fits the data characteristics of the protocol data to be classified. At the same time, when the convolutional autoencoder is used for feature extraction, the extracted features are also adapted to this calculation method. The similar protocol data corresponding to each protocol data to be classified determined in this way is more accurate, thereby improving the clustering effect of the protocol data to be classified.

[0081] Step S204 , clustering the protocol data to be classified based on the number of similar protocol data corresponding to each protocol data to be classified and a preset protocol number threshold, and obtaining a classification result for each protocol data to be classified according to the clustering result.

[0082] The protocol quantity threshold may be understood as a pre-set value used to determine the coreness of the current protocol to be classified.

[0083] Optionally, the server 104 determines the protocol data to be classified whose number of similar protocol data is greater than or equal to the first protocol number threshold as the cluster center corresponding to the cluster set, adds the similar protocol data of the protocol data to be classified to the cluster set, and repeats the above process until there is no protocol data to be classified whose number of similar protocol data is greater than or equal to the first protocol number threshold; adds the protocol data to be classified whose number of similar protocol data is less than or equal to the second protocol number threshold to the noise cluster set, and the first protocol number threshold is greater than the second protocol number threshold. Finally, multiple cluster sets and a noise protocol cluster set representing abnormal protocols can be obtained.

[0084] Based on the above-mentioned implementation method, by comparing the number of similar protocol data corresponding to the protocol data to be classified and the pre-set protocol number threshold, different protocol data to be classified are clustered in a targeted manner, taking into account the data characteristics between different protocol data to be classified, which not only improves the accuracy of classification but also speeds up the classification.

[0085] In the above-mentioned power monitoring system protocol classification method, multiple protocol data to be classified are collected from the network traffic of the current monitoring system, and each protocol data to be classified is input into the trained convolutional autoencoder for feature extraction to obtain the protocol features corresponding to each protocol data to be classified. The similarity between each protocol data to be classified is obtained based on each protocol feature, and based on the similarity, the number of similar protocol data corresponding to each protocol data to be classified is obtained. The similar protocol data is other protocol data to be classified whose similarity with the current protocol data to be classified is greater than or equal to a preset similarity threshold. The current protocol data to be classified is any one of the protocol data to be classified, and the other protocol data to be classified is the protocol data to be classified in each protocol data to be classified except the current protocol data to be classified. Finally, based on the number of similar protocol data corresponding to each protocol data to be classified and a preset protocol number threshold, the protocol data to be classified are clustered, and the classification result of each protocol data to be classified is obtained according to the clustering result. By acquiring the protocol data to be classified in real time and analyzing the similarity of the protocol data to be classified, the protocol data to be classified are finally clustered according to the number of similar protocol data to be classified and the preset number threshold. This clustering method does not rely on the selection of the initial center, is suitable for the classification needs of unstructured message data, and improves the classification accuracy of protocol data.

[0086] In one embodiment, clustering the protocol data to be classified based on the number of similar protocol data corresponding to each protocol data to be classified and a preset protocol number threshold includes:

[0087] Based on the number of similar protocol data corresponding to each protocol data to be classified and a preset protocol number threshold, a density distribution type corresponding to each protocol data to be classified is obtained; the density distribution type includes: core protocol type, boundary protocol type and noise protocol type;

[0088] Randomly obtain a target protocol data from each protocol data to be classified, and use the target protocol data as the cluster center of the cluster set corresponding to the target protocol data; wherein the target protocol data is the protocol data to be classified that has not been clustered and the corresponding density distribution type is the core protocol type;

[0089] Adding similar protocol data corresponding to the target protocol data into the cluster set, and returning to the step of randomly obtaining a target protocol data from each protocol data to be classified, until the target protocol data is no longer included in the protocol data to be classified;

[0090] The protocol data to be classified whose density distribution type is the noise protocol type is added to the noise cluster set.

[0091] Among them, the density distribution type can be understood as the distribution of other protocol data to be classified around the protocol data to be classified, including the number of data and the data position. The other protocol data to be classified here refers to the similar protocol data corresponding to the protocol data to be classified; the core protocol type can be understood as the density distribution type of the protocol data to be classified whose number of similar protocol data corresponding to the protocol data to be classified meets the core judgment condition; the boundary protocol type can be understood as the density distribution type of the protocol data to be classified whose number of similar protocol data corresponding to the protocol data to be classified meets the boundary judgment condition; the noise protocol type can be understood as the density distribution type of the remaining protocol data to be classified except for the protocol data to be classified whose density distribution type is the core protocol type and the boundary protocol type.

[0092] Target protocol data can be understood as protocol data with a density distribution type of core protocol types and that is to be clustered. The cluster center can be understood as the extension center of a cluster set. The data in the cluster set is extended from this cluster center, that is, the data in the cluster set is related to the cluster center. Noise cluster sets can be understood as abnormal protocol traffic, such as undefined protocols or malicious attack traffic.

[0093] Optionally, based on the number of similar protocol data corresponding to each protocol data to be classified and a pre-set protocol number threshold, the server 104 determines that the density distribution type of the protocol data to be classified that meets the core judgment condition is the core protocol type, the density distribution type of the protocol data to be classified that meets the boundary judgment condition is the boundary protocol type, and the density distribution type of the remaining protocol data to be classified is the noise protocol type. A target protocol data is determined from the protocol data to be classified that does not participate in clustering among the protocol data to be classified whose density distribution type is the core protocol type, and the target protocol data is used as the cluster center of the cluster set corresponding to the target protocol data. The similar protocol data corresponding to the target protocol data is added to the cluster set, and the step of determining a target protocol data from the protocol data to be classified that does not participate in clustering among the protocol data to be classified that has a density distribution type of the core protocol type is returned to execute until all the protocol data to be classified whose density distribution type is the core protocol type is traversed. The protocol data to be classified whose density distribution type is the noise protocol type is added to the noise cluster set.

[0094] Based on the above-mentioned implementation method, the density distribution type of each protocol data to be classified is determined based on the number of similar protocol data corresponding to the protocol data to be classified and a pre-set protocol number threshold, and data clustering is performed according to the corresponding clustering method based on the density distribution type of the protocol data to be classified, thereby ensuring the accuracy and robustness of clustering and accelerating the speed of clustering.

[0095] In one embodiment, adding similar protocol data corresponding to the target protocol data into a cluster set includes:

[0096] Obtain the current similar protocol data; the current similar protocol data is any one of the similar protocol data corresponding to the target protocol data; when the density distribution type of the current similar protocol data is a boundary protocol type, add the current similar protocol data to the cluster set; when the density distribution type of the current similar protocol data is a core protocol type, add the current similar protocol data to the cluster set, and add the similar protocol data corresponding to the current similar protocol data as the similar protocol data corresponding to the target protocol data.

[0097] Exemplarily, the server 104 obtains any current similar protocol data from the similar protocol data corresponding to the target protocol data, and when the density distribution type of the current similar protocol data is a boundary protocol type, adds the current similar protocol data into the cluster set; when the density distribution type of the current similar protocol data is a core protocol type, adds the current similar protocol data into the cluster set, and adds the similar protocol data corresponding to the current similar protocol data as an extension of the target protocol data, as the similar protocol data corresponding to the target protocol data, and adds them into the cluster set at the same time.

[0098] Based on the above implementation method, by further clustering and classifying the similar protocol data corresponding to the target protocol data, the cluster set corresponding to the target protocol data is processed according to the density distribution type corresponding to the similar protocol data. The similar protocol data with a density distribution type of the boundary protocol type is directly added to the cluster set. The similar protocol data with a density distribution type of the core protocol type needs to be used as the similar protocol data of the target protocol data and added to the cluster set to achieve set expansion; the classification of density distribution can clarify the relationship between protocol data, optimize the clustering results, ensure that the clustering results are of practical significance, achieve set expansion, and effectively increase the coverage and representativeness of the clustering, bring about a significant improvement in processing capabilities, and improve the integrity of classification.

[0099] In an exemplary embodiment, similar protocol data corresponding to the target protocol data is added to the cluster set, and the step of randomly obtaining a target protocol data from each protocol data to be classified is returned to be executed until the target protocol data is no longer included in the protocol data to be classified, further comprising:

[0100] Obtain the remaining protocol data to be classified; the density distribution type of the remaining protocol data to be classified is the boundary protocol type; obtain the similarity between the remaining protocol data to be classified and the target protocol data as the cluster center of each cluster set; add the remaining protocol data to be classified into the cluster set corresponding to the target protocol data with the highest similarity.

[0101] Optionally, after traversing the clustering of the target protocol data and adding all the protocol data to be classified whose density distribution type is the noise protocol type into the noise cluster set, there are still some remaining protocol data to be classified whose density distribution type is the noise protocol type. The similarity between the remaining protocol data to be classified and the target protocol data as the cluster center of each cluster set is obtained, and the remaining protocol data to be classified is added to the cluster set corresponding to the target protocol data with the highest similarity.

[0102] Based on the above implementation, by performing secondary clustering processing on the remaining protocol data to be classified after one round of clustering is completed, the integrity of data processing of the protocol data to be classified is ensured, the clustering coverage of the clustering results is enhanced, and the classification accuracy of data classification is improved.

[0103] In one embodiment, the protocol quantity threshold includes a first protocol quantity threshold and a second protocol quantity threshold, and the first protocol quantity threshold is greater than the second protocol quantity threshold. Based on the number of similar protocol data corresponding to each to-be-classified protocol data and a preset protocol quantity threshold, obtaining the density distribution type corresponding to each to-be-classified protocol data includes:

[0104] When the number of similar protocol data corresponding to the current protocol data to be classified is greater than or equal to the first protocol number threshold, determining that the density distribution type of the current protocol data to be classified is a core protocol type;

[0105] When the number of similar protocol data corresponding to the current protocol data to be classified is less than the first protocol number threshold and greater than the second protocol number threshold, determining that the density distribution type of the current protocol data to be classified is a boundary protocol type;

[0106] When the number of similar protocol data corresponding to the current protocol data to be classified is less than or equal to the second protocol number threshold, it is determined that the density distribution type of the current protocol data to be classified is a noise protocol type.

[0107] The first protocol quantity threshold can be understood as a value that can be set according to actual conditions. The same is true for the second protocol quantity threshold. The first protocol quantity threshold must always be greater than the second protocol quantity threshold.

[0108] Exemplarily, when the number of similar protocol data corresponding to the current protocol data to be classified is greater than or equal to the first protocol number threshold, the density distribution type of the current protocol data to be classified is determined to be a core protocol type; when the number of similar protocol data corresponding to the current protocol data to be classified is less than the first protocol number threshold and greater than the second protocol number threshold, the density distribution type of the current protocol data to be classified is determined to be a boundary protocol type; when the number of similar protocol data corresponding to the current protocol data to be classified is less than or equal to the second protocol number threshold, the density distribution type of the current protocol data to be classified is determined to be a noise protocol type.

[0109] Core protocol type: If a protocol data to be classified is i of The neighborhood contains at least (i.e., the aforementioned first protocol quantity threshold) points, the density distribution type of the protocol data to be classified is the core protocol type:

[0110]

[0111] Boundary protocol type: If a protocol data to be classified At a core point neighborhood, but the number of sample points in its own neighborhood is less than , then the density distribution type of the protocol data to be classified is the boundary protocol type:

[0112]

[0113] Noise protocol type: If a protocol data point is neither a core point nor in the neighborhood of any core point (i.e., the number of similar protocol data corresponding to the current protocol data to be classified is less than or equal to the second protocol number threshold), then the density distribution type of the protocol data to be classified is the noise protocol type:

[0114]

[0115] Based on the aforementioned implementation method, the first protocol quantity threshold and the second protocol quantity threshold are used to judge the size relationship between similar protocol data corresponding to the current protocol data to be classified, and the density distribution type of the current protocol data to be classified is determined based on the size relationship. The classification is based on the dynamic changes in the number of similar protocol data, and can adapt to changes in protocol behavior in different data environments, thereby ensuring the real-time and reliability of the classification.

[0116] In one embodiment, the trained convolutional autoencoder includes a convolutional layer and a pooling layer;

[0117] Each protocol data to be classified is input into the trained convolutional autoencoder for feature extraction to obtain the protocol features corresponding to each protocol data to be classified, including:

[0118] Each protocol data to be classified is input into the convolution layer for feature extraction to obtain the feature maps corresponding to the protocol data to be classified; each feature map is input into the pooling layer for feature dimensionality reduction to obtain the protocol features corresponding to the protocol data to be classified.

[0119] Among them, the convolution layer can be understood as a structural component that uses convolution calculation to extract features. It processes the input data (such as images) through convolution operations, and uses several filters (or convolution kernels) to perform sliding window operations with the input to generate a feature map. The pooling layer can be understood as a structural component that reduces the size of the feature map and is used for feature dimensionality reduction. By downsampling the feature map, the size of the feature map is reduced while retaining key information.

[0120] Optionally, each protocol data to be classified is input into the convolutional layer of the convolutional autoencoder for feature extraction. The convolutional layer gradually mines the deep spatial features of the data from different byte streams, and performs an inner product operation on the message area within the receptive field through the convolution operation to obtain the feature map of the message data. The calculation formula of the convolution operation is as follows: .in represents the convolution kernel, Represents input.

[0121] The convolutional autoencoder uses a two-layer convolutional network to extract hierarchical features of the protocol data to be classified. The first convolution layer uses 32 filters, each with a kernel size of 3, to extract local features from the input data. The second convolution layer uses 64 filters with a kernel size of 3 to further extract high-level protocol features. K convolution kernels (W) are initialized, each with a bias b. After convolution with the one-dimensional sequence data input x, k feature maps h are generated. The activation function is calculated as follows: .

[0122] Each feature map is input into the pooling layer of the convolutional autoencoder. The pooling layer aggregates the local area of ​​the feature map obtained by the convolution operation and compresses the binary message feature information into a single value. The pooling calculation formula is H pool =max(H).

[0123] Based on the above implementation method, the convolution layer of the convolutional autoencoder is used to perform convolution calculation on the protocol data to be classified, and hierarchical features are extracted to realize deep feature mining, and the feature map corresponding to each protocol data to be classified is obtained. Then, each feature map is input into the pooling layer of the convolutional autoencoder for feature dimensionality reduction operation, which reduces the size of the feature map and reduces the subsequent calculation complexity. At the same time, feature dimensionality reduction can make the protocol features corresponding to each protocol data to be classified more compact and stable, which is conducive to the subsequent protocol clustering analysis.

[0124] In one embodiment, a convolutional autoencoder is trained by the following steps: obtaining a plurality of sample protocol data; inputting each sample protocol data into the convolutional autoencoder to be trained for feature extraction to obtain sample protocol features corresponding to each sample protocol data; inputting each sample protocol feature into the decoder corresponding to the convolutional autoencoder to be trained for data reconstruction to obtain reconstructed protocol data corresponding to each sample protocol data; training the parameters of the convolutional autoencoder to be trained based on the sample protocol data and the reconstructed protocol data to obtain a trained convolutional autoencoder.

[0125] Among them, the decoder can be understood as a component that uses extracted features to reconstruct data. The decoder includes an anti-pooling layer and a deconvolution layer.

[0126] Exemplarily, the server 104 obtains a plurality of sample protocol data, inputs each sample protocol data into the convolutional autoencoder to be trained for feature extraction, obtains the sample protocol features corresponding to each sample protocol data, inputs each sample protocol feature into the decoder corresponding to the convolutional autoencoder to be trained for data reconstruction, first inputs each sample protocol feature into the depooling layer in the decoder to restore the compressed sample protocol features to higher-dimensional sample protocol features, and then inputs the higher-dimensional sample protocol features into the deconvolution layer in the decoder for data reconstruction, and uses transposed convolution to map the low-dimensional features back to the original protocol data space. The encoded binary message feature map is restored to the size of the original input through an upsampling operation, thereby realizing binary message reconstruction. Each binary message feature map h is convolved with the transpose of its corresponding convolution kernel and the results are summed, and then the bias c is added, and the formula is as follows:

[0127]

[0128] The deconvolution layer consists of three layers. The first deconvolution layer uses 64 filters to restore the feature size. The second deconvolution layer uses 32 filters to restore the protocol data structure. The third deconvolution layer uses 1 filter to output the reconstructed protocol data corresponding to each sample protocol data.

[0129] The mean squared error (MSE) is used as the loss function, and the Adam optimizer is used to train the parameters of the convolutional autoencoder to be trained:

[0130] The principle of minimizing the average quadratic reconstruction error function is the target value (i.e. the aforementioned reconstruction protocol data) minus the actual value The square sum of the sample protocol data corresponding to the reconstructed protocol data is then averaged, and the formula is as follows:

[0131]

[0132] Where n represents the number of all data points, and the loss function is used to optimize the convolutional autoencoder to be trained so that the sample features it extracts can retain the core information of the protocol.

[0133] The Adam optimizer is used for training. Adam adaptively adjusts the learning rate through first-order moment estimation and second-order moment estimation:

[0134]

[0135] in , They are first-order moment estimation and second-order moment estimation respectively; is the learning rate; , is the momentum parameter; Indicates the parameter value at the current moment, Represents the neighborhood radius.

[0136] Based on the above-mentioned implementation method, the sample features are de-pooled through the de-pooling layer in the decoder, and the dimension of the compressed sample features is increased, thereby improving the similarity between the reconstructed protocol data and the corresponding sample protocol data. The deconvolution layer in the decoder is used to transpose the sample features after the dimension increase to achieve data reconstruction, which is conducive to parameter training of the feature extraction convolution autoencoder.

[0137] The average quadratic reconstruction error is adopted as the loss function and the Adam optimizer is used for training to improve the convergence speed and stability of the convolutional autoencoder.

[0138] In an exemplary embodiment, Figure 3 As shown, a specific implementation of the power monitoring system protocol classification method is provided, including steps S1 to S15 (the following are all data in a specific example, and do not limit the application to be implemented only in this form).

[0139] S1: Collect protocol data from the network traffic of the current monitoring system and perform preprocessing.

[0140] In the power monitoring system, the protocol data flow consists of multiple communication protocols. Assume that the protocol data flow is ,in Represents a protocol message of a power monitoring system. To adapt the deep learning model to the analysis of power protocol traffic, the message data needs to be converted into a fixed-length tensor representation. Assuming the maximum length of the protocol data is L, the protocol data can be represented as a matrix: Among them, each protocol message After padding or truncation, it is unified to length L. This matrix is ​​used as the input of the convolutional autoencoder for protocol feature extraction.

[0141] S2: Convolutional layer performs power monitoring protocol feature extraction.

[0142] In the communication environment of power monitoring systems, protocol data is typically transmitted as binary byte streams. Different protocols (such as IEC 60870-5-104, DNP3, Modbus TCP, and DL / T 476) have different data structures and field layouts. Therefore, to extract the deep-level features of these protocols and achieve automatic protocol classification and abnormal traffic detection, this application uses a convolutional autoencoder (CAE) to learn protocol features.

[0143] The encoder part of CAE uses a one-dimensional convolutional layer to extract the deep features of the protocol data. The convolutional layer gradually mines the deep spatial features of the data from different byte streams. Through the convolution operation, the inner product operation is performed on the message area within the receptive field to obtain the feature map of the message data. The calculation formula of the convolution operation is as follows: .in represents the convolution kernel, Represents input.

[0144] CAE uses a two-layer convolutional network to extract hierarchical features of protocol data. The first convolution layer uses 32 filters, each with a kernel size of 3, to extract local features from the input data. The second convolution layer uses 64 filters with a kernel size of 3 to further extract high-level protocol features. K convolution kernels (W) are initialized, each with a bias b. After convolution with the one-dimensional sequence data input x, k feature maps h are generated. The activation function is calculated as follows: .

[0145] S3: The pooling layer performs feature dimensionality reduction to improve the stability of power monitoring protocol classification.

[0146] In power monitoring systems, protocol data often contains a large amount of redundant information, such as fixed fields, padding bits, or repetitive patterns in protocol headers. To reduce computational complexity and improve the stability of the classification model, this application introduces a pooling layer after the convolutional layer to perform feature dimensionality reduction. This makes the extracted power protocol features more compact and stable, facilitating subsequent protocol clustering analysis.

[0147] The pooling layer aggregates the local area of ​​the feature map obtained by the convolution operation, compressing the binary message feature information into a single value, thereby reducing the size of the feature map and reducing the computational complexity of the model. The pooling calculation formula is H pool = max(H). Max pooling reduces data size by taking the maximum value of a local region while retaining key features and improving model computational efficiency. The pooling window size is 2, meaning that a maximum value is taken for every two data points, thus reducing feature dimensionality and improving feature extraction stability.

[0148] S4: The unpooling layer performs feature recovery and reconstructs the power monitoring protocol data.

[0149] In the protocol classification process of power monitoring systems, a convolutional autoencoder (CAE) is used not only for feature extraction but also for protocol data reconstruction via a decoder to ensure that the extracted protocol features accurately reflect the original protocol structure. In the decoder, the unpooling layer, as the inverse operation of the pooling layer, restores the reduced features, making the reconstructed data closer to the original protocol data format.

[0150] De-pooling uses the position relationship matrix to restore the information compressed by the pooling layer to a higher dimension: H unpool =Unpoool(H pool ). Unpooling ensures that the decoder can recover more protocol information, making the final reconstructed protocol data as close to the original data format as possible.

[0151] S5: The deconvolution layer reconstructs the protocol data and restores the power monitoring protocol message.

[0152] Since power protocol messages usually have a fixed format structure, such as the ASDU (Application Service Data Unit) structure of IEC104 and the function code and data field of Modbus TCP, it is necessary to restore the spatial characteristics and hierarchical information of the protocol message during the reconstruction process to ensure that the decoded protocol data can accurately reflect the original protocol format.

[0153] The decoder uses transposed convolution to map low-dimensional features back to the original protocol data space. Upsampling restores the encoded binary message feature map to the original input size, thereby reconstructing the binary message. Each binary message feature map h is convolved with the transposed convolution kernel of its corresponding convolution kernel, and the sum of the results is then added with the bias c. The formula is as follows:

[0154]

[0155] The deconvolution layer consists of three layers. The first deconvolution layer recovers the feature size through 64 filters. The second deconvolution layer recovers the protocol data structure through 32 filters. The third deconvolution layer outputs the protocol data through 1 filter.

[0156] S6: Convolutional autoencoders are trained based on the Adam (Adaptive Moment Estimation) optimization algorithm to improve the ability to extract features from power monitoring protocols.

[0157] In the protocol classification task of power monitoring systems, the core goal of a convolutional autoencoder (CAE) is to accurately extract key features of protocol data, ensuring that feature vectors of different protocol types (such as IEC 104, Modbus TCP, DNP3, DL / T 476, etc.) can be effectively distinguished while maintaining the integrity of the protocol data structure. Therefore, to optimize the feature extraction capabilities of the CAE, this application uses the mean squared error (MSE) as the loss function and uses the Adam optimizer for training to improve the model's convergence speed and stability.

[0158] The principle of minimizing the average quadratic reconstruction error function is the target value Subtract the actual value The sum of the squares and the mean is then calculated, and the formula is as follows: , n represents the number of all data points. The loss function is used to optimize CAE so that the feature Z it extracts can retain the core information of the protocol.

[0159] The Adam optimizer is used for training. Adam adaptively adjusts the learning rate through first-order moment estimation and second-order moment estimation:

[0160]

[0161] in , They are first-order moment estimation and second-order moment estimation respectively; is the learning rate; , is the momentum parameter; Indicates the parameter value at the current moment, Represents the neighborhood radius. After training, the encoder outputs the low-dimensional feature Z as the input for DBSCAN (Density-Based Spatial Clustering of Applications with Noise).

[0162] S7: Calculate the similarity measure between protocol feature vectors to optimize the classification of power monitoring protocols.

[0163] Before performing DBSCAN density clustering, it is necessary to calculate the similarity between protocol feature vectors to measure the relative distance between different protocol data points, so as to accurately classify the protocol categories. Since the communication protocols used in power monitoring systems (such as IEC 60870-5-104, DNP3, Modbus TCP, DL / T 476, etc.) have different data structures and feature distributions, similarity measurement is crucial for the accuracy of protocol classification. Since this application uses convolutional autoencoders (CAEs) to extract features from protocol data, each protocol data point is All have been converted into one -dimensional feature vector: .in, Represents the characteristic dimension of the protocol data, and N is the total number of data points. DBSCAN needs to calculate every two protocol data points , The distance between them is used to determine their similarity. This application uses Euclidean distance to measure the similarity between protocol data points and constructs an adjacency matrix as the input of DBSCAN clustering. The formula is as follows:

[0164] S8: Set the key parameters of the DBSCAN algorithm.

[0165] In the classification of power monitoring system protocols, DBSCAN relies on the neighborhood radius (Eps, ) and the minimum number of sample points ( ) These two key parameters determine the density distribution of protocol data and implement adaptive protocol clustering. Therefore, a reasonable setting and It is crucial for the accuracy of protocol classification and clustering effect.

[0166] Neighborhood radius : Define a protocol data point As the density range of the core point, that is, within the radius All data points within are considered as its neighbors. Minimum number of sample points :Specify that a protocol data point requires at least neighbors to be considered as a core point. It can be determined by K-distance plot, that is, calculating the k nearest neighbor distance of each point and drawing a curve, and selecting . Generally set as feature dimension 2 to 4 times of .

[0167] S9: Identify core points, border points, and noise points to optimize power monitoring protocol classification.

[0168] During the DBSCAN clustering process, protocol data points are classified as core points, border points, or noise points based on their density distribution within the neighborhood. In power monitoring systems, protocol messages come from different communication protocols (such as IEC 60870-5-104, DNP3, Modbus TCP, DL / T 476, etc.). DBSCAN uses this classification to automatically identify protocols and detect abnormal protocol traffic or unknown protocol data.

[0169] Core point: If a certain protocol data point of The neighborhood contains at least point, then this point is the core point:

[0170]

[0171] Boundary point: If a certain protocol data point At a core point neighborhood, but the number of sample points in its own neighborhood is less than , then it is a boundary point:

[0172]

[0173] Noise point: If a protocol data point is neither a core point nor in the neighborhood of any core point, it is marked as a noise point: Noise points usually indicate abnormal protocol traffic, such as undefined protocols or malicious attack traffic.

[0174] S10: Randomly select an unvisited protocol data point , initialize the power monitoring protocol clustering.

[0175] like As the core point, a new cluster is constructed from it. .like If it is a boundary point, it will not be processed for the time being and wait for the expansion of the core point. If it is a noise point, it is directly marked as abnormal data.

[0176] S11: Add all density-reachable agreement data points to the same cluster.

[0177] If a data point At the core of If the sigmoid is in the same neighborhood, it belongs to the same cluster. If it is also a core point, its neighborhood will continue to expand to form a larger cluster.

[0178] S12: Traverse all protocol data points until all core points are expanded.

[0179] Each core point will expand a cluster, and eventually form multiple protocol categories C1, C2, ..., C K ,completely classify the protocol data in the power monitoring system.

[0180] S13: Processing boundary points.

[0181] If a data point belongs to the neighborhood of multiple core points, it is classified into the corresponding cluster according to the density relationship.

[0182] S14: Identify the protocol type and abnormal protocol data.

[0183] After DBSCAN calculation, the final output protocol categories are C1, C2, ..., C K :Each cluster represents a power monitoring system protocol type, such as Modbus, IEC104, DL / T 476, etc. Noise point set : Contains protocol data that is not classified into any cluster, which may be abnormal traffic or unknown protocols.

[0184] S15: Output clustering results.

[0185] Compared with the existing technology, this application has the following technical advantages:

[0186] 1. The convolutional autoencoder is used to extract the deep protocol features of the power monitoring system protocol data, and the extracted protocol features are subjected to data feature dimensionality reduction. This makes the protocol features corresponding to each of the extracted protocol data to be classified more compact and stable. At the same time, it reduces the computing power resources consumed by the subsequent analysis and processing of the protocol features, thereby accelerating the classification of the protocol data and ensuring real-time classification.

[0187] 2. Combined with DBSCAN to handle noise and non-spherical clusters, on large-scale data sets, DBSCAN improves processing efficiency by avoiding the use of distance calculations between all points. It has a relatively loose parameter selection and strong adaptability, making it perform well in different data scenarios, thereby improving classification robustness.

[0188] 3. Perform data preprocessing on the collected protocol data and unify the data length of the protocol data, thereby reducing the technical difficulty of feature extraction and speeding up the feature extraction and classification speeds.

[0189] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0190] Based on the same inventive concept, an embodiment of the present application further provides a power monitoring system protocol classification device for implementing the above-mentioned power monitoring system protocol classification method. The implementation solution provided by this device is similar to the implementation solution described in the above-mentioned method. Therefore, the specific limitations in one or more power monitoring system protocol classification device embodiments provided below can be referred to the limitations of the power monitoring system protocol classification method above, and will not be repeated here.

[0191] In an exemplary embodiment, Figure 4 As shown, a power monitoring system protocol classification device is provided, including: a data acquisition module 401, a feature extraction module 402, a similarity acquisition module 403 and a data classification module 404, wherein:

[0192] The data collection module 401 is used to collect a plurality of protocol data to be classified from the network traffic of the current monitoring system.

[0193] The feature extraction module 402 is used to input each protocol data to be classified into the trained convolutional autoencoder to perform feature extraction, so as to obtain the protocol features corresponding to each protocol data to be classified.

[0194] The similarity acquisition module 403 is used to obtain the similarity between each protocol data to be classified based on each protocol feature, and based on the similarity, obtain the number of similar protocol data corresponding to each protocol data to be classified; similar protocol data refers to other protocol data to be classified whose similarity with the current protocol data to be classified is greater than or equal to a preset similarity threshold; the current protocol data to be classified is any one of the protocol data to be classified, and the other protocol data to be classified refers to the protocol data to be classified other than the current protocol data to be classified.

[0195] The data classification module 404 is configured to cluster the protocol data to be classified based on the number of similar protocol data corresponding to each protocol data to be classified and a preset protocol number threshold, and obtain a classification result for each protocol data to be classified according to the clustering result.

[0196] In one embodiment, the data classification module 404 further includes a distribution type acquisition submodule, a first clustering submodule, a second clustering submodule, and a third clustering submodule, wherein:

[0197] The distribution type acquisition submodule is used to obtain the density distribution type corresponding to each protocol data to be classified based on the number of similar protocol data corresponding to each protocol data to be classified and a preset protocol number threshold; the density distribution type includes: core protocol type, boundary protocol type and noise protocol type;

[0198] The first clustering submodule is used to randomly obtain a target protocol data from each protocol data to be classified, and use the target protocol data as the cluster center of the cluster set corresponding to the target protocol data; wherein the target protocol data is the protocol data to be classified that has not been clustered and the corresponding density distribution type is the core protocol type;

[0199] The second clustering submodule is used to add similar protocol data corresponding to the target protocol data into the cluster set, and return to the step of randomly obtaining a target protocol data from each protocol data to be classified until the target protocol data is no longer included in the protocol data to be classified;

[0200] The third clustering submodule is used to add the protocol data to be classified whose density distribution type is the noise protocol type into the noise clustering set.

[0201] In one embodiment, the clustering second sub-module is also used to obtain current similar protocol data; the current similar protocol data is any one of the similar protocol data corresponding to the target protocol data; when the density distribution type of the current similar protocol data is a boundary protocol type, the current similar protocol data is added to the clustering set; when the density distribution type of the current similar protocol data is a core protocol type, the current similar protocol data is added to the clustering set, and the similar protocol data corresponding to the current similar protocol data is added as the similar protocol data corresponding to the target protocol data.

[0202] In an exemplary embodiment, similar protocol data corresponding to the target protocol data is added to the cluster set, and the step of randomly obtaining a target protocol data from each protocol data to be classified is returned to execute until the target protocol data is no longer contained in the protocol data to be classified. The second submodule of clustering is also used to obtain the remaining protocol data to be classified; the density distribution type of the remaining protocol data to be classified is the boundary protocol type; the similarity between the remaining protocol data to be classified and the target protocol data as the cluster center of each cluster set is obtained; and the remaining protocol data to be classified is added to the cluster set corresponding to the target protocol data with the highest similarity.

[0203] In one embodiment, the protocol quantity threshold includes a first protocol quantity threshold and a second protocol quantity threshold, and the first protocol quantity threshold is greater than the second protocol quantity threshold; the distribution type acquisition submodule is further used to determine that the density distribution type of the current protocol data to be classified is a core protocol type when the number of similar protocol data corresponding to the current protocol data to be classified is greater than or equal to the first protocol quantity threshold; when the number of similar protocol data corresponding to the current protocol data to be classified is less than the first protocol quantity threshold and greater than the second protocol quantity threshold, determine that the density distribution type of the current protocol data to be classified is a boundary protocol type; when the number of similar protocol data corresponding to the current protocol data to be classified is less than or equal to the second protocol quantity threshold, determine that the density distribution type of the current protocol data to be classified is a noise protocol type.

[0204] In an exemplary embodiment, the trained convolutional autoencoder includes a convolution layer and a pooling layer; the feature extraction module 402 is also used to input each protocol data to be classified into the convolution layer for feature extraction to obtain feature maps corresponding to the protocol data to be classified; each feature map is input into the pooling layer for feature dimensionality reduction to obtain protocol features corresponding to each protocol data to be classified.

[0205] In one embodiment, the power monitoring system protocol classification device also includes a training module for obtaining multiple sample protocol data; inputting each sample protocol data into the convolutional autoencoder to be trained for feature extraction to obtain sample protocol features corresponding to each sample protocol data; inputting each sample protocol feature into the decoder corresponding to the convolutional autoencoder to be trained for data reconstruction to obtain reconstructed protocol data corresponding to each sample protocol data; training the parameters of the convolutional autoencoder to be trained based on the sample protocol data and the reconstructed protocol data to obtain a trained convolutional autoencoder.

[0206] Each module in the above-mentioned power monitoring system protocol classification device can be implemented in whole or in part through software, hardware, or a combination thereof. Each of the above-mentioned modules can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the corresponding operations of each of the above modules.

[0207] In an exemplary embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in FIG. Figure 5 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store the network traffic of the current monitoring system, the protocol data to be classified, the protocol characteristics, the similarity, the similar protocol data, the number of similar protocol data, and the clustering results. The input / output interface of the computer device is used to exchange information between the processor and the external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a protocol classification method for a power monitoring system.

[0208] Those skilled in the art will understand that Figure 5 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0209] In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the power monitoring system protocol classification method of the above embodiment when executing the computer program.

[0210] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the power monitoring system protocol classification method of the above embodiment is implemented.

[0211] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the power monitoring system protocol classification method of the above embodiment is implemented.

[0212] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0213] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), quantum computing-based data processing logic devices, artificial intelligence (AI) processors, and the like.

[0214] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0215] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A method for classifying protocols of power monitoring systems, characterized in that: The method comprises: Collecting multiple protocol data to be classified from the network traffic of the power monitoring system; Inputting each of the protocol data to be classified into the trained convolutional autoencoder for feature extraction to obtain protocol features corresponding to each of the protocol data to be classified; Based on each of the protocol features, a degree of similarity between each of the protocol data to be classified is obtained, and based on the degree of similarity, a number of similar protocol data corresponding to each of the protocol data to be classified is obtained; the similar protocol data is other protocol data to be classified whose degree of similarity with the current protocol data to be classified is greater than or equal to a preset similarity threshold; the current protocol data to be classified is any one of the protocol data to be classified, and the other protocol data to be classified is the protocol data to be classified in each of the protocol data to be classified except the current protocol data to be classified; Based on the number of similar protocol data corresponding to each of the protocol data to be classified and a preset protocol number threshold, the protocol data to be classified are clustered, and a classification result of each of the protocol data to be classified is obtained according to the clustering result.

2. The method according to claim 1, characterized in that Clustering the protocol data to be classified based on the number of similar protocol data corresponding to each protocol data to be classified and a preset protocol number threshold includes: Based on the number of similar protocol data corresponding to each of the protocol data to be classified and a preset protocol number threshold, obtaining a density distribution type corresponding to each of the protocol data to be classified; the density distribution type includes: a core protocol type, a boundary protocol type, and a noise protocol type; Randomly obtaining a target protocol data from each of the protocol data to be classified, and using the target protocol data as the cluster center of the cluster set corresponding to the target protocol data; wherein the target protocol data is the protocol data to be classified that has not been clustered and whose corresponding density distribution type is the core protocol type; Adding similar protocol data corresponding to the target protocol data into the cluster set, and returning to the step of randomly acquiring a target protocol data from each of the protocol data to be classified until the protocol data to be classified does not contain the target protocol data; The protocol data to be classified whose density distribution type is the noise protocol type is added into the noise cluster set.

3. The method according to claim 2, characterized in that Adding similar protocol data corresponding to the target protocol data into the cluster set includes: Acquire current similar protocol data; the current similar protocol data is any one of the similar protocol data corresponding to the target protocol data; In a case where the density distribution type of the current similar protocol data is the boundary protocol type, adding the current similar protocol data into the cluster set; When the density distribution type of the current similar protocol data is the core protocol type, the current similar protocol data is added to the cluster set, and the similar protocol data corresponding to the current similar protocol data is added as the similar protocol data corresponding to the target protocol data.

4. The method according to claim 3, characterized in that The step of adding similar protocol data corresponding to the target protocol data to the cluster set and returning to the step of randomly obtaining a target protocol data from each of the protocol data to be classified until the target protocol data is no longer included in the protocol data to be classified further includes: Obtaining remaining protocol data to be classified; wherein the density distribution type of the remaining protocol data to be classified is the boundary protocol type; Obtaining the similarity between the remaining protocol data to be classified and the target protocol data serving as the cluster center of each cluster set; The remaining protocol data to be classified are added to the cluster set corresponding to the target protocol data with the highest similarity.

5. The method according to claim 2, characterized in that The protocol quantity threshold includes a first protocol quantity threshold and a second protocol quantity threshold, and the first protocol quantity threshold is greater than the second protocol quantity threshold; obtaining the density distribution type corresponding to each of the protocol data to be classified based on the quantity of similar protocol data corresponding to each of the protocol data to be classified and the preset protocol quantity threshold includes: When the number of similar protocol data corresponding to the current protocol data to be classified is greater than or equal to the first protocol number threshold, determining that the density distribution type of the current protocol data to be classified is the core protocol type; When the number of similar protocol data corresponding to the current protocol data to be classified is less than the first protocol number threshold and greater than the second protocol number threshold, determining that the density distribution type of the current protocol data to be classified is the boundary protocol type; When the number of similar protocol data corresponding to the current protocol data to be classified is less than or equal to the second protocol number threshold, the density distribution type of the current protocol data to be classified is determined to be the noise protocol type.

6. The method according to claim 1, characterized in that The trained convolutional autoencoder includes a convolutional layer and a pooling layer; The step of inputting each of the protocol data to be classified into a trained convolutional autoencoder for feature extraction to obtain protocol features corresponding to each of the protocol data to be classified comprises: Input each of the protocol data to be classified into the convolution layer for feature extraction to obtain feature maps corresponding to the protocol data to be classified; Each of the feature maps is input into the pooling layer to perform feature dimensionality reduction, so as to obtain the protocol features corresponding to each of the protocol data to be classified.

7. The method according to claim 6, characterized in that The convolutional autoencoder is trained by the following steps: Obtain multiple sample protocol data; Inputting each of the sample protocol data into the convolutional autoencoder to be trained for feature extraction to obtain sample protocol features corresponding to each of the sample protocol data; Inputting each of the sample protocol features into the decoder corresponding to the convolutional autoencoder to be trained to reconstruct data, thereby obtaining reconstructed protocol data corresponding to each of the sample protocol data; The parameters of the convolutional autoencoder to be trained are trained according to the sample protocol data and the reconstructed protocol data to obtain a trained convolutional autoencoder.

8. A protocol classification device for a power monitoring system, characterized in that: The device comprises: A data acquisition module is used to collect a plurality of protocol data to be classified from the network traffic of the current monitoring system; A feature extraction module is used to input each of the protocol data to be classified into a trained convolutional autoencoder for feature extraction to obtain protocol features corresponding to each of the protocol data to be classified; a similarity acquisition module, configured to acquire, based on each of the protocol features, a degree of similarity between each of the protocol data to be classified, and based on the degree of similarity, acquire a quantity of similar protocol data corresponding to each of the protocol data to be classified; the similar protocol data being other protocol data to be classified whose degree of similarity to the current protocol data to be classified is greater than or equal to a preset similarity threshold; the current protocol data to be classified being any one of the protocol data to be classified, and the other protocol data to be classified being protocol data to be classified other than the current protocol data to be classified in the protocol data to be classified; The data classification module is used to cluster the protocol data to be classified based on the number of similar protocol data corresponding to each protocol data to be classified and a preset protocol number threshold, and obtain a classification result of each protocol data to be classified according to the clustering result.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.