Power internet of things communication protocol data cleaning and feature extraction method

Through the data cleaning and feature extraction method of the power Internet of Things communication protocol, the data quality problem in the protocol identification and classification of the power Internet of Things terminal equipment is solved, and the accuracy of protocol identification and classification is improved.

CN120676025APending Publication Date: 2025-09-19STATE GRID ELECTRIC POWER RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510967708.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

In the existing technology, the protocols of power Internet of Things terminal devices are different, which makes it difficult for terminal devices to automatically register and perceive their operating status. In addition, data noise, outliers, duplicate data, and missing data affect the accuracy of protocol recognition and classification algorithms.

Method used

A method for cleaning and extracting features from a power Internet of Things communication protocol data is provided. The method includes structured processing, feature field generation, normalization processing, and feature similarity model establishment. Data is acquired through bus packet capture, and invalid characters are removed using encoding module decoding and regular expressions. Feature fields and normalized data sets are generated, and a feature similarity model is established to extract valid features.

Benefits of technology

The impact of factors such as data noise, outliers, and duplicate data on the protocol identification and classification algorithm of power Internet of Things terminal equipment is reduced, and the accuracy of protocol identification and classification is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120676025A_ABST
    Figure CN120676025A_ABST
Patent Text Reader

Abstract

The invention discloses an electric power Internet of Things communication protocol data cleaning and feature extraction method, which relates to the technical field of electric power Internet of Things, and comprises the following steps: generating exchange data based on a communication protocol of edge side equipment and intelligent electric power Internet of Things terminal equipment, carrying out structured processing on the exchange data, generating a feature field, and carrying out feature extraction on the feature field; generating an original data set based on the feature field; extracting feature data in the original data set based on a communication protocol, generating a key feature field based on the feature data, and performing normalization processing on the key feature data set to generate a normalized data set; establishing a feature similarity model based on a communication protocol, extracting effective features in the normalized data set by using the feature similarity model, and generating a training data set based on the effective features; a data basis is provided for side deployment of a power Internet of Things terminal protocol identification and classification scheme based on a data driving method, and conditions are provided for power Internet of Things terminal protocol identification and classification feature extraction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electric power Internet of Things, and in particular to a method for cleaning and extracting features of electric power Internet of Things communication protocol data. Background Art

[0002] With the rapid development of mobile Internet and artificial intelligence, the types of power terminal devices are becoming increasingly diverse, and consumers' expectations for the quality and efficiency of power services are also increasing. To meet these needs and enhance power users' awareness and participation in smart grids, the Power Internet of Things has emerged. The Power Internet of Things can flexibly and efficiently connect power users with various enterprises and devices to achieve data sharing. It provides users with a unified platform for access and information exchange, serving multiple participants such as users, power grids, power companies, and suppliers.

[0003] Power IoT technology, with its advantages such as flexible perception, real-time communication, and intelligent control, has gradually become an indispensable component of modern power grid construction. It not only enhances power users' awareness and participation in smart grids, but also provides strong technical support for grid management and operation and maintenance. Among them, smart power IoT terminal devices adopt the "cloud-pipe-edge-end" smart IoT architecture, accessing the distribution automation system, electricity consumption information collection system, and IoT management platform, providing great convenience for grid management staff. However, with the increase in smart power IoT terminal devices, the protocols between devices vary (including but not limited to 698 protocol, 645 protocol, Modbus protocol, etc.), and some device protocols are even unknown, resulting in an increasing workload for grid management and operation and maintenance personnel. In order to strengthen the operation and maintenance management of power IoT terminal devices and meet the application needs of grid management and operation and maintenance users, it is necessary to automatically identify the protocol category of connected smart devices;

[0004] The serious lack of holographic perception, connection, open sharing and other capabilities has become a bottleneck for the development of the power Internet of Things. At the perception terminal level, due to the lack of terminal identification-related systems, standards and basic technologies, it is difficult to achieve automatic registration and operation status perception of terminals. Secondly, due to the diverse types of perception terminals and the lack of a unified terminal identification mechanism and standard information interaction protocol, manual configuration of access terminal types, transmission protocols and information data categories is required, which cannot achieve automatic registration of terminal devices and low terminal access efficiency. Therefore, it is urgent to build a power Internet of Things terminal protocol identification and classification model to realize the automatic identification and management of existing power Internet of Things terminals by edge devices;

[0005] However, the existing traffic message data of power Internet of Things terminals has problems such as data complexity, irregularity, and diversity. The input of pcap package data of different specifications seriously affects the accuracy of the power Internet of Things terminal protocol identification and classification model. The present invention aims to propose a power Internet of Things communication protocol data cleaning and feature extraction method to reduce the impact of data quality factors such as data noise, outliers, duplicate data, and missing data on the accuracy of the power Internet of Things terminal device protocol identification and classification algorithm. Summary of the Invention

[0006] In order to overcome the above-mentioned technical problems, the purpose of the present invention is to provide a method for cleaning and extracting power Internet of Things communication protocol data to solve the problem in the prior art that data quality factors such as data noise, outliers, duplicate data, and missing data reduce the accuracy of the power Internet of Things terminal device protocol identification and classification algorithm.

[0007] The purpose of the present invention can be achieved through the following technical solutions:

[0008] Specifically, a method for cleaning and extracting features of power Internet of Things communication protocol data is provided, comprising the following steps:

[0009] S1: Generate exchange data based on the communication protocol between the edge device and the smart power IoT terminal device, perform structured processing on the exchange data, generate feature fields, and generate the original data set based on the feature fields;

[0010] S2: Extract feature data from the original data set based on the communication protocol, generate key feature fields based on the feature data, normalize the key feature data set, and generate a normalized data set;

[0011] S3: Establish a feature similarity model based on the communication protocol, use the feature similarity model to extract effective features in the normalized dataset, and generate a training dataset based on the effective features.

[0012] As a further solution of the present invention: S1 comprises the following steps:

[0013] S11: Define the ports to be traversed based on the communication protocol between the edge device and the smart power IoT terminal device, set serial port parameters based on the ports, and generate exchange data based on the serial port parameters;

[0014] S12: traverse and open the port. If the port is opened successfully, read N exchange data of the port, where N is a positive integer;

[0015] S13: If the port fails to open, the output capture exception;

[0016] S14: decoding the read N exchange data of the port using the encoding module to generate decoded data;

[0017] S15: using a regular expression to remove invalid characters in the decoded data, retaining hexadecimal characters, and saving the data in a pcap data packet format, and generating a feature field based on the data in the pcap data packet;

[0018] S16: Generate an original data set based on the feature fields and write the original data set into an Excel file.

[0019] As a further solution of the present invention: the encoding module includes utf-8 encoding, latin1 encoding, gbk encoding or iso-8859-1 encoding.

[0020] As a further solution of the present invention: S2 comprises the following steps:

[0021] S21: Segmenting the original data set based on the communication protocol to form a number of feature data, and merging every two unit data in the number of feature data to form a key feature field;

[0022] S22: converting the hexadecimal in each key feature field into decimal to form a decimal feature field;

[0023] S23: normalize the decimal feature field to generate a normalized data set;

[0024] S24: Output the normalized dataset.

[0025] As a further solution of the present invention: S23 includes the following steps:

[0026] S231: Divide each feature bit in the decimal feature field by 255 to obtain a processing feature field;

[0027] S232: Fill the processing feature field. If the length of the processing feature field is less than 196, fill it with 0. If the length of the processing feature field exceeds 196, delete the redundant field.

[0028] As a further solution of the present invention: S3 comprises the following steps:

[0029] S31: Establish feature similarity model based on communication protocol;

[0030] S32: Determine the communication protocol direction and terminal based on the normalized data set and generate basic communication features;

[0031] S33: identifying similarities of basic communication features using a feature similarity model based on the basic communication features, and generating similarity features;

[0032] S34: Determine the data length of the similarity feature and select a valid feature based on the data length;

[0033] S35: Generate a training dataset based on the valid features.

[0034] As a further solution of the present invention: the similarity of the basic communication characteristics includes the similarity of traffic data between the receiving and sending terminals of the same terminal or protocol, the similarity of basic data fields in the traffic data received or sent by the same terminal or protocol, or the similarity of basic data fields generated by the same protocol.

[0035] As a further solution of the present invention: the data length includes the length of a single data packet, the average length of a data packet, the difference between the length of a single data packet and the average length of a data packet, the difference between the length of a single data packet and the longest data packet, or the difference between the length of a single data packet and the shortest data packet.

[0036] As a further solution of the present invention: the edge side device and the smart power Internet of Things terminal device obtain the data set by bus packet capture.

[0037] As a further solution of the present invention: the bus packet capture is an industrial bus protocol, and the industrial bus protocol includes Modbus, CAN, RS485 or Profinet.

[0038] Beneficial effects of the present invention:

[0039] In the present invention, a data basis is provided for the side-deployed power Internet of Things terminal protocol identification and classification scheme based on a data-driven method, conditions are provided for the power Internet of Things terminal protocol identification and classification feature extraction, and the influence of data quality factors such as data noise, outliers, duplicate data, and missing data on the accuracy of the power Internet of Things terminal device protocol identification and classification algorithm is reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] The present invention will be further described below with reference to the accompanying drawings.

[0041] Figure 1 This is a flow chart of a method for cleaning and extracting features of a power Internet of Things communication protocol data in the present invention;

[0042] Figure 2 This is a data receiving flow chart of a method for cleaning and extracting features of a power Internet of Things communication protocol in the present invention;

[0043] Figure 3 This is a schematic diagram of a hexadecimal original message data set of a method for cleaning and extracting features of a power Internet of Things communication protocol in the present invention;

[0044] Figure 4 This is a flow chart of a data preprocessing solution for a PCAP data packet format of a method for cleaning and extracting features of a power Internet of Things communication protocol in the present invention;

[0045] Figure 5 This is a data preprocessing flow chart of Example 1 of a method for cleaning and extracting features of a power Internet of Things communication protocol according to the present invention;

[0046] Figure 6 This is a data preprocessing flow chart of Example 2 in a method for cleaning and extracting features of a power Internet of Things communication protocol of the present invention. DETAILED DESCRIPTION

[0047] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0048] Example 1

[0049] like Figure 1-Figure 5 As shown, the present invention discloses a method for cleaning and extracting features of power Internet of Things communication protocol data, comprising the following steps:

[0050] The first step is to generate exchange data based on the communication protocol between the edge device and the smart power IoT terminal device. The exchange data is structured to generate feature fields, and then the original data set is generated based on the feature fields. The edge device and the smart power IoT terminal device obtain the data set by bus packet capture. The bus packet capture is based on the industrial bus protocol, which includes Modbus, CAN, RS485, or Profinet. It should be noted that:

[0051] First, define the ports to be traversed based on the communication protocol between the edge device and the smart power IoT terminal device. Then, set the serial port parameters based on the ports and generate exchange data based on the serial port parameters. Note that when the industrial bus protocol is Modbus, the communication protocol is also Modbus. When the industrial bus protocol is CAN / RS485 / Profinet, the communication protocol is also CAN / RS485 / Profinet.

[0052] Then traverse and open the port. If the port is opened successfully, read N exchange data of the port, where N is a positive integer. For example, when N=5, read 5 exchange data of the port. The size of N can be arbitrarily set by those skilled in the art.

[0053] If the port fails to open, the output captures the exception;

[0054] Decode the read port N exchange data using an encoding module to generate decoded data. The encoding module includes but is not limited to utf-8 encoding, latin1 encoding, gbk encoding or iso-8859-1 encoding;

[0055] It's important to note that UTF-8 is a variable-length encoding module for encoding Unicode characters into byte sequences. Its design goal is to be compatible with ASCII and provide unified encoding capabilities across languages. It's an implementation of the Unicode standard, known as "Unicode Transformation Format, 8-bit," and uses one to four bytes to encode each Unicode character. Its core principle is to convert Unicode code points into a set of byte sequences in a specific format that doesn't conflict with ASCII, thus ensuring backward compatibility. The key to UTF-8 encoding lies in its highly recognizable and decode-deterministic byte structure. Each UTF-8-encoded byte sequence begins with a high-order 1 that clearly indicates the number of bytes comprising the character. For example, single-byte characters begin with 0, double-byte characters begin with 110, and three-byte characters begin with 1110. All subsequent bytes other than the first byte uniformly begin with 10. This design allows data to be intercepted at any point in a byte stream, and the byte prefix can be analyzed to determine whether the current byte is the start of a character. This allows for synchronization and positioning, effectively avoiding garbled characters or synchronization errors.

[0056] The Latin1 encoding maps each character to a fixed 8-bit integer value, that is, a value between 0 and 255. The first half of the encoding (0-127) is reserved for standard ASCII characters, including English letters (uppercase and lowercase), numbers, punctuation marks, control characters, etc. For example, the Latin1 encoding value of the letter "A" is 0x41, which is 65 in decimal, and the space character corresponds to 0x20, which is 32 in decimal. This part is exactly the same as the ASCII encoding, thus ensuring that Latin1 is fully compatible with ASCII. In the range 0x80 to 0x9F, Latin1 reserves a set of control characters. These characters were used in early computer systems to represent non-displayable information in text streams, such as page feeds and device control. However, this range is rarely used in modern applications. The range 0xA0 to 0xFF is the core part of Latin1, covering a variety of diacritical letters and symbols unique to Western European languages. For example, the character "é" is encoded in Latin1 as 0xE9, which is 233 in decimal; the character Corresponds to 0xF1; non-breaking space is 0xA0;

[0057] GBK encoding is based on character set mapping, which maps Unicode characters to a set of byte combinations designed specifically for Chinese. Its core mechanism is table lookup encoding: that is, each Chinese character or symbol has a unique corresponding code point in the GBK encoding table. This code point consists of two bytes. During the encoding process, the system first determines whether the current character is a character in the ASCII range. If so, the corresponding value 0x00 to 0x7F is directly output in single-byte form; if the character is Chinese or other special symbols, the system will look up the double-byte encoding value corresponding to the character from the GBK encoding table, and then output the two bytes sequentially. During the decoding process of GBK encoding, the receiver reads a byte. If the byte is between 0x00 and 0x7F, it is directly parsed as an ASCII character; if it is between 0x81 and 0xFE, it indicates that the character is a double-byte encoded Chinese character. The system will continue to read the next byte (should be between 0x40 and 0xFE), and then splice the two bytes to form a code value, and look up the corresponding character in the GBK encoding table.

[0058] ASCII code assigns a unique integer number to commonly used English letters, numbers, and symbols, then converts this number into a 7-bit binary representation and transmits it as an electrical signal or byte. Of these 128 codes, the first 32 (0 to 31) are control characters used to control device behavior. These characters themselves do not have visible symbolic forms.

[0059] ISO-8859-1 encoding is a single-byte character encoding scheme that keeps the first 128 characters of ASCII unchanged and uses the remaining 128 bytes of encoding space (i.e. 0x80 to 0xFF) to increase support for Latin letter variants, common symbols, and some control characters.

[0060] Use regular expressions to remove invalid characters from the decoded data, retain hexadecimal characters, and save it in pcap packet format. Generate feature fields based on the data in the pcap packet. It should be noted that the method of using regular expressions to remove invalid characters from the decoded data is to use regular expressions to match and filter the string pattern, and then replace the matched invalid parts with empty spaces, thereby obtaining cleaned data containing only the required valid characters. In different application scenarios, the definition of invalid characters may vary. For example, control characters may be considered invalid in log cleaning, characters outside the ASCII range may be considered invalid in device protocol data processing, and non-target language characters may be excluded in multilingual text processing. The specific selection is based on the adaptability of the technology in the field. Then, the exclusion character set of the regular expression, such as [^...], is used to match all characters not within the range specified by the square brackets. When the regular expression performs a matching operation on the string, it checks each character from the beginning of the string. Any characters that match outside the valid range are identified as deletion targets. These deletion targets are then replaced with empty strings, that is, directly deleted, thereby completing the cleaning of the decoded data.

[0061] Generate an original data set based on the feature fields and write the original data set into an Excel file. It should be noted that after extracting the feature fields, these fields need to be combined row by row to form an original data set. Each row of data represents a communication or a data record, and the columns store different feature fields, such as timestamp, function code, register address, data length, data value, checksum, etc. After forming such an original data set, it can be written into an Excel file.

[0062] The second step is to extract feature data from the original dataset based on the communication protocol, generate key feature fields based on the feature data, and normalize the key feature dataset to generate a normalized dataset. It should be noted that:

[0063] The original data set is segmented based on the communication protocol to form several feature data. Every two units of data in the several feature data are merged to form a key feature field. The hexadecimal in each key feature field is converted to decimal to form a decimal feature field. The decimal feature field is normalized to generate a normalized data set. It should be noted that each feature bit in the decimal feature field is divided by 255 to obtain a processing feature field. The processing feature field is padded. If the length of the processing feature field is less than 196, it is padded with 0. If the length of the processing feature field exceeds 196, the redundant fields are deleted and the normalized data set is output.

[0064] The third step is to establish a feature similarity model based on the communication protocol, use the feature similarity model to extract effective features from the normalized dataset, and generate a training dataset based on the effective features. Specifically:

[0065] A feature similarity model is established based on the communication protocol. It should be noted that the feature similarity model performs feature engineering analysis based on prior knowledge. The prior knowledge includes but is not limited to the similarity of traffic data between the receiving and sending terminals of the same terminal or protocol, the similarity of basic data fields in the traffic data received or sent by the same terminal or protocol, and the similarity of basic data fields generated by the same protocol. On this basis, effective features closely related to data length are selected, including but not limited to the length of a single data packet, the average length of a data packet, the difference between the length of a single data packet and the average length of a data packet, the difference between the length of a single data packet and the longest data packet, or the difference between the length of a single data packet and the shortest data packet.

[0066] Determine the communication protocol direction and terminal based on the normalized data set and generate basic communication features;

[0067] Based on the basic communication features, a feature similarity model is used to identify the similarity of basic communication features and generate similarity features. It should be noted that the similarity of basic communication features includes the similarity of traffic data between the receiving and sending terminals of the same terminal or protocol, the similarity of basic data fields in the traffic data received or sent by the same terminal or protocol, or the similarity of basic data fields generated by the same protocol;

[0068] Determine the data length of similarity features and select valid features based on the data length. It should be noted that data length includes the length of a single data packet, the average length of data packets, the difference between the length of a single data packet and the average length of data packets, the difference between the length of a single data packet and the longest data packet, or the difference between the length of a single data packet and the shortest data packet.

[0069] The training data set is generated based on effective features. It should be noted that the training data set generated by effective features is used as the input of the machine learning model to reduce the data noise, outliers, duplicate data, and missing data of the training data, and improve the accuracy of the protocol identification and classification algorithm of the power Internet of Things terminal equipment.

[0070] Example 2

[0071] like Figure 1-Figure 5 As shown, the present invention discloses a method for cleaning and extracting features of power Internet of Things communication protocol data, taking the industrial bus protocol RS485 as an example, which specifically includes the following steps:

[0072] The first step is to generate exchange data based on the communication protocol between the edge device and the smart power IoT terminal device, perform structured processing on the exchange data, generate feature fields, and generate the original data set based on the feature fields. The edge device and the smart power IoT terminal device obtain the data set by bus packet capture. The bus packet capture is an industrial bus protocol, where the bus protocol is RS485 and the communication protocol is RS485. First, define the ports to be traversed based on RS485. The number of ports is 8. Bus packet capture is used to monitor 8 ports and set the serial port parameters. Consider setting the serial port parameters to baudrate = 9600: the baud rate is set to 9600, bytesize = 8: the data bits are set to 8 bits, and stopbits = 1: the stop bit is set to 1 bit to form exchange data.

[0073] Traverse and open all ports. If they are successfully opened, enter the data reading stage and read N=5 data. If it fails, output capture exception. After obtaining the data, use multiple encoding modules (including but not limited to 'utf-8', 'latin1', 'gbk', 'ascii', 'iso-8859-1', etc.) to decode the data to form decoded data. Then use regular expressions to remove invalid characters in the decoded data and only retain hexadecimal characters (numbers 0-9 and letters a-fA-F). UTF-8 encoding, latin1 encoding, gbk encoding or iso-8859-1 encoding Example 1 has been described in detail and will not be repeated again.

[0074] And save it in pcap data packet format, generate feature fields based on the data in the pcap data packet, generate original data sets based on the feature fields, and write the original data sets into the Excel file.

[0075] In the second step, since the received data cannot be directly predicted using the model, preprocessing operations are required, as shown in the following example: Figure 2 As shown, the following steps are included:

[0076] Step 1: First, the original data set must be segmented to form several feature data. Each two unit data in the feature data are combined to form a key feature field (two hexadecimal digits are equivalent to 1 byte, that is, 8 binary digits). Specifically, the original data set is segmented based on the communication protocol to form several feature data. Each two unit data in the feature data are combined to form a key feature field.

[0077] Step 2: Then convert the hexadecimal in each key feature field into a decimal number that the algorithm can learn, forming a decimal feature field;

[0078] Step 3: To improve the convergence speed of the model and reduce numerical instability, the decimal feature field needs to be normalized (each feature bit is divided by 255, i.e. FF) to obtain the processed feature field;

[0079] Step 4: Then fill the processing feature field. If the processing feature field length is less than 196, fill it with 0. If the processing feature field length exceeds 196, delete the extra fields. The specific implementation steps are as follows Figure 5 As shown;

[0080] Finally, obtain the pre-processed data sets of the communication protocols (including but not limited to Modbus, CAN, RS485 or Profinet, etc.) of different smart power IoT terminal devices;

[0081] The processed hexadecimal data is stored in the edge device results list and the port number is added. The format is F{port number}{cleaned hexadecimal data}. Finally, the result is saved in pcap data packet format to form a hexadecimal original data set. The result is saved to a file in Excel. The original data set format is as follows Figure 3 As shown;

[0082] The third step is to establish a feature similarity model based on the communication protocol, use the feature similarity model to extract effective features from the normalized dataset, and generate a training dataset based on the effective features. Specifically, in terms of feature selection, prior knowledge is summarized to conduct feature engineering analysis, including but not limited to: there are certain differences in the similarity of traffic data between the receiving and sending terminals of the same terminal or protocol;

[0083] The similarity of basic data fields in traffic data received or sent by the same terminal or protocol is much higher than that in traffic data from mixed terminals or protocols;

[0084] When information data with a specific function is generated by the same protocol, the similarity of the basic data fields of the information data is usually higher than the similarity between information data generated by different functions;

[0085] On this basis, select effective features closely related to length, including but not limited to single packet length, average packet length, difference between single packet length and average packet length, difference between single packet length and longest packet length, difference between single packet length and shortest packet length, etc., and calculate the above features for each sample of each protocol type based on the extracted pre-processed payload data;

[0086] First, there are certain differences in the traffic data between the receiving and sending terminals;

[0087] Secondly, the similarity of basic data fields in traffic data received or sent by the same terminal or protocol is much higher than that in traffic data from mixed terminals or protocols. The same protocol generates information data with specific functions, and the similarity of basic data fields is usually higher than that between information data generated by different functions.

[0088] At the same time, analysis found that the terminal communication protocols of different manufacturers also have different modes. Although they are all Modbus communication protocols, some manufacturers use Modbus-TCP communication protocol, while some manufacturers use Modbus / RTU communication protocol. The data frames of the two are composed of different formats.

[0089] Furthermore, the similarity of register addresses corresponding to function codes in traffic data received or sent by the same terminal or protocol is much higher than that of register addresses corresponding to function codes in traffic data from mixed terminals or protocols;

[0090] Based on the above analysis, we select effective features that are closely related to information data such as message byte length, distance between messages and address, such as single data packet length, average data packet length, difference between single data packet and average data packet length, difference between single data packet and longest data packet length, difference between single data packet and shortest data packet length, etc., as shown in Table 1;

[0091] Table 1 Protocol identification feature extraction

[0092]

[0093] where γ i is the length of the i-th data packet, γ avg is the average length of the data packet, γ max is the longest packet length, γ min is the shortest data packet length, n is the total number of samples, m is the sequence number of the mth sample, k is the total dimension of single sample features, j is the jth feature of a single sample, x m,j Refers to the j-dimensional feature data of the m-th sample, x i,j Refers to the j-th dimension feature data of the i-th sample, d i The payload data is formed by the mean distance between all samples in the sample set and the full-dimensional features of the i-th sample. Based on the extracted payload data, the above features are calculated for each sample of each protocol type, and the above features are merged with the effective features to generate a training dataset, which is used as the input data for model training.

[0094] Example 3

[0095] like Figure 1 、 Figure 3 、 Figure 4 、 Figure 5 and Figure 6 As shown, the present invention discloses a method for cleaning and extracting features of power Internet of Things communication protocol data, taking the industrial bus protocol CAN as an example, which specifically includes the following steps:

[0096] The first step is to generate exchange data based on the communication protocol between the edge device and the smart power Internet of Things terminal device, perform structured processing on the exchange data, generate feature fields, and generate the original data set based on the feature fields. The edge device and the smart power Internet of Things terminal device obtain the data set by bus packet capture, which is an industrial bus protocol. The bus protocol is CAN and the communication protocol is CAN. First, the ports to be traversed are defined based on CAN. The number of ports is 8. 8 ports are monitored through bus packet capture. The serial port parameters are set based on the port to form exchange data. Then, the port is traversed and opened. If the port is opened successfully, port N=5 data is read. If the port fails to be opened, the capture exception is output. Typically, the N data of the port read are decoded using an encoding module to form decoded data. The encoding module includes utf-8 encoding, latin1 encoding, gbk encoding or iso-8859-1 encoding. Utf-8 encoding, latin1 encoding, gbk encoding or iso-8859-1 encoding Example 1 has been described in detail and will not be repeated here. A regular expression is used to clear invalid characters in the decoded data, and the processed hexadecimal data is stored in the edge device results list and the port number is added. The format is F{port number}{cleaned hexadecimal data}, and finally the result is saved in pcap data packet format to form a hexadecimal original data set;

[0097] The second step is to segment the payload original application data set, treat every two bits of data in the payload original application data set as a feature data, convert the hexadecimal in each feature data into decimal to form decimal feature data, and normalize the decimal feature data. Specifically, each feature bit in the decimal feature field is divided by 255 to obtain a processing feature field, and the processing feature field is padded. If the length of the processing feature field is less than 196, it is padded with 0. If the length of the processing feature field exceeds 196, the redundant fields are deleted and the normalized data set is output;

[0098] The third step is to establish a feature similarity model based on the communication protocol, use the feature similarity model to extract valid features in the normalized data set, generate a training data set based on the valid features, establish a feature similarity model based on the communication protocol, determine the communication protocol direction and terminal based on the normalized data set, generate basic communication features, use the feature similarity model based on the basic communication features to identify the similarity of the basic communication features, generate similarity features, determine the data length of the similarity features, select valid features based on the data length, and generate a training data set based on the valid features. This is the same as Example 1 and Example 2 and will not be repeated here.

[0099] The above is a detailed description of an embodiment of the present invention. However, the content described is only a preferred embodiment of the present invention and should not be considered to limit the scope of the present invention. All equivalent changes and improvements made within the scope of the present invention should still fall within the scope of the patent coverage of the present invention.

Claims

1. A method for cleaning and extracting features of power Internet of Things communication protocol data, characterized in that: The following steps are involved: S1: Generate exchange data based on the communication protocol between the edge device and the smart power IoT terminal device, perform structured processing on the exchange data, generate feature fields, and generate the original data set based on the feature fields; S2: Extract feature data from the original data set based on the communication protocol, generate key feature fields based on the feature data, normalize the key feature data set, and generate a normalized data set; S3: Establish a feature similarity model based on the communication protocol, use the feature similarity model to extract effective features in the normalized dataset, and generate a training dataset based on the effective features.

2. A method for cleaning and extracting features of power Internet of Things communication protocol data according to claim 1, characterized in that: Said S1 comprises the following steps: S11: Define the ports to be traversed based on the communication protocol between the edge device and the smart power IoT terminal device, set serial port parameters based on the ports, and generate exchange data based on the serial port parameters; S12: traverse and open the port. If the port is opened successfully, read N exchange data of the port, where N is a positive integer; S13: If the port fails to open, the output capture exception; S14: decoding the read N exchange data of the port using the encoding module to generate decoded data; S15: using a regular expression to remove invalid characters in the decoded data, retaining hexadecimal characters, and saving the data in a pcap data packet format, and generating a feature field based on the data in the pcap data packet; S16: Generate an original data set based on the feature fields and write the original data set into an Excel file.

3. A method for cleaning and extracting features of power Internet of Things communication protocol data according to claim 2, characterized in that: The encoding module includes utf-8 encoding, latin1 encoding, gbk encoding or iso-8859-1 encoding.

4. The method for cleaning and extracting features of a power Internet of Things communication protocol data according to claim 1, characterized in that: The S2 comprises the following steps: S21: Segmenting the original data set based on the communication protocol to form a number of feature data, and merging every two unit data in the number of feature data to form a key feature field; S22: converting the hexadecimal in each key feature field into decimal to form a decimal feature field; S23: normalize the decimal feature field to generate a normalized data set; S24: Output the normalized dataset.

5. A method for cleaning and extracting features of power Internet of Things communication protocol data according to claim 4, characterized in that: The S23 includes the following steps: S231: Divide each feature bit in the decimal feature field by 255 to obtain a processing feature field; S232: Fill the processing feature field. If the length of the processing feature field is less than 196, fill it with 0. If the length of the processing feature field exceeds 196, delete the redundant field.

6. A method for cleaning and extracting features of power Internet of Things communication protocol data according to claim 1, characterized in that: The S3 includes the following steps: S31: Establish feature similarity model based on communication protocol; S32: Determine the communication protocol direction and terminal based on the normalized data set and generate basic communication features; S33: identifying similarities of basic communication features using a feature similarity model based on the basic communication features, and generating similarity features; S34: Determine the data length of the similarity feature and select a valid feature based on the data length; S35: Generate a training dataset based on the valid features.

7. A method for cleaning and extracting features of power Internet of Things communication protocol data according to claim 6, characterized in that: The similarity of the basic communication characteristics includes the similarity of traffic data between the receiving and sending terminals of the same terminal or protocol, the similarity of basic data fields in the traffic data received or sent by the same terminal or protocol, or the similarity of basic data fields generated by the same protocol.

8. The method for cleaning and extracting features of a power Internet of Things communication protocol data according to claim 6, characterized in that: The data length includes the length of a single data packet, the average length of a data packet, the difference between the length of a single data packet and the average length of a data packet, the difference between the length of a single data packet and the longest data packet, or the difference between the length of a single data packet and the shortest data packet.

9. The method for cleaning and extracting features of a power Internet of Things communication protocol data according to claim 1, characterized in that: The edge side device and the smart power Internet of Things terminal device obtain the data set by bus packet capture.

10. A method for cleaning and extracting features of power Internet of Things communication protocol data according to claim 9, characterized in that: The bus packet capture is an industrial bus protocol, which includes Modbus, CAN, RS485 or Profinet.