Protocol analysis method based on proprietary industrial network
Through customized data acquisition and multi-method fusion analysis, binary and text-based protocols are targetedly analyzed, combined with deep learning technology, data heterogeneity and security problems in industrial network protocol analysis are solved, efficient and secure protocol analysis and data processing are achieved, and industrial intelligence is promoted.
Patent Information
- Application Number
- CN202510623515.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-08-08
AI Technical Summary
The existing industrial network protocol analysis technology has problems such as data heterogeneity and insufficient noise processing, difficulty in analyzing protocol structure and encryption mechanism, lack of targeted parsing strategies, vulnerabilities in security protection systems, lack of systematic processing processes, and difficult to break through real-time constraints, resulting in reduced credibility of the analysis results, delayed response of the control system, and threatened facility security.
We adopt customized data acquisition, multi-method fusion analysis, targeted protocol analysis, data processing and display optimization, algorithm optimization and deep learning assistance methods, through Wireshark software packet capture, layered research, state organization construction, machine learning and big data collaborative processing, and analyze binary and text-based protocols separately to establish a secure data transmission and storage mechanism.
It realizes accurate collection and analysis of industrial network protocols, improves the accuracy and speed of analysis, ensures data security, adapts to the real-time analysis needs of complex industrial networks, and promotes the development of industrial intelligence.
Smart Images

Figure CN120455566A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of industrial control protocol analysis, and in particular relates to a protocol analysis method based on a proprietary industrial network. Background Art
[0002] The existing industrial network protocol parsing technology has the following defects: insufficient data heterogeneity and noise processing, difficulty in parsing protocol structure and encryption mechanism, lack of targeted parsing strategy, loopholes in security protection system, lack of systematic processing flow, and difficulty in breaking through real-time constraints. Specifically:
[0003] The data formats generated by industrial equipment vary significantly. For example, in the nuclear energy sector, sensor output is a mix of binary and text formats, and contains significant noise and erroneous data. Traditional preprocessing methods are unable to effectively eliminate this interference, resulting in reduced confidence in analysis results and a lack of reliable support for industrial control.
[0004] Industrial proprietary protocols often utilize complex nested data structures and specialized encryption methods (such as the dynamic key mechanism used in nuclear facility monitoring systems). Existing general-purpose parsing technologies are unable to effectively identify the multi-layered protocol architecture and encryption rules, resulting in increased parsing errors. For example, parsing delays in nuclear reactor monitoring data can lead to control system responsiveness failures, threatening facility safety.
[0005] Existing methods use a unified parsing framework for binary protocols (such as proprietary binary encoding) and text protocols (such as ASCII encoding), without establishing adaptation mechanisms for data type differences. Furthermore, fixed template matching strategies struggle to identify semantic rules specific to industrial protocols (such as nuclear energy parameter encoding specifications), resulting in incomplete extraction of key information.
[0006] Encryption strength during data transmission is insufficient, especially for sensitive information such as nuclear facility operating parameters, which lacks quantum-resistant encryption protection. Access control mechanisms are imperfect and unable to effectively defend against new cyberattacks (such as APT attacks based on protocol vulnerabilities), posing a risk of data leakage.
[0007] In the processing flow, in the acquisition stage, the data structures of traditional devices and new intelligent devices are poorly compatible, and private binary codes are difficult to parse; in the preprocessing stage: fixed cleaning rules cannot adapt to the variable-length structure and special verification mechanism of industrial protocols, resulting in the loss of valid data; in the parsing stage: protocol reverse engineering is inefficient, and rule updates lag behind protocol version iterations; in the classification and storage stage: general semantic models have difficulty identifying contextual associations of industrial data, and the classification results are not accurate enough.
[0008] Scenarios such as nuclear energy control require millisecond-level analysis and response, but existing technologies are limited by the efficiency of complex protocol parsing algorithms, which can easily lead to chain risks such as control instruction delays. Summary of the Invention
[0009] To overcome the shortcomings of the aforementioned prior art, the present invention aims to provide a protocol parsing method based on proprietary industrial networks that accurately collects valid data while minimizing noise and erroneous data. This method employs a comprehensive approach to interpreting various types of protocols, while also establishing a comprehensive data security system that encrypts and controls access to sensitive information during data collection, data transmission, and data storage. This not only improves interpretation accuracy and speed, maximizing the value of data, but also ensures industrial data security.
[0010] In order to achieve the above object, the technical solution adopted by the present invention is:
[0011] A protocol parsing method based on a proprietary industrial network includes the following steps:
[0012] Step 1: Customized data collection: Customize scenario-based design solutions to capture specific data packets and obtain data;
[0013] Step 2: Multi-method fusion analysis: By integrating three methods, namely, packet layer analysis, protocol analysis and reproduction, and large-scale data comparison, the captured data is analyzed in an orderly and comprehensive manner;
[0014] Step 3: Targeted protocol analysis: Use corresponding algorithms to analyze binary and text protocols respectively to ensure the correctness and accuracy of the analysis results. During the process, repeat step 2 to ensure the accuracy of step 3.
[0015] Step 4: Data processing and display optimization: Write a program to assist in the analysis process. At the same time, use big data to collaboratively process the suspected but unknown fields during the data processing process; write the corresponding UI interface; and combine the existing results from step 3 to jointly process the parsed part with the part to be parsed, so as to explore the meaning of the unparsed fields.
[0016] Step 5: Algorithm optimization and deep learning assistance: The captured data is divided into a data set and a training set, and the training set is used to verify the correctness of the data set analysis results. During the verification process, step 1 is repeated, and the environment is changed multiple times to verify the conclusion.
[0017] The capture of a specific data packet in step 1 is specifically as follows:
[0018] Deploy Wireshark software on the virtual controller and virtual host, then start Wireshark to capture data packets. After capturing packets, save the file and copy it.
[0019] The scenario-based design solution is divided into stand-alone device operation scenario and multi-device collaborative operation scenario. For stand-alone device operation scenario, by setting multiple trigger conditions for changing communication data, such as device offline and online, sending specific instructions, and powering on and off, it is possible to accurately capture data packets containing specific operations.
[0020] The specific operations are:
[0021] For the first collection, the collection point is on the virtual host;
[0022] A certain configuration of the physical controller is offline and online;
[0023] All configurations of physical controllers are offline and online;
[0024] A certain configuration of the virtual controller goes offline and online;
[0025] All configurations of virtual controllers are offline and online;
[0026] Modify the inter-station communication variable output sent from the virtual controller to the physical controller;
[0027] Modify the physical controller from the second layer;
[0028] Turn off the virtual controller;
[0029] Unplug the network cable of the physical controller;
[0030] Restart the second layer.
[0031] For multi-machine collaborative operation scenarios, multiple actual collection points are set up to capture multiple data packets between different devices in the same operation, thereby collecting a large amount of comparative data and comprehensively monitoring the overall operation of the communication system.
[0032] The specific operations are:
[0033] During the process of establishing a link between a virtual controller and a virtual host, the collection points are on the virtual host and virtual controller, and the collection is done simultaneously.
[0034] The virtual controller can be powered off and on directly;
[0035] Control the virtual controller offline and online.
[0036] The specific steps of the three methods in step 2 are:
[0037] In the comprehensive application of data packet layering research, Wireshark software is used to layer the captured data packets and distinguish other layer data from link layer data; the analysis of industrial control protocols is mainly based on link layer data;
[0038] Through the layered view function of Wireshark software, you can clearly see the packet structure, data characteristics, and whether the underlying data has special text meaning, and understand the rules and characteristics of the protocol format;
[0039] In terms of protocol parsing and reproduction, we follow the complete process of data preprocessing, format extraction, and state machine establishment to analyze the captured packet data and obtain preliminary field division and semantic inference;
[0040] On this basis, a large amount of data of different operations is compared based on semantic inference.
[0041] The data preprocessing is to filter the captured massive data, and process it using the MAC addresses and data packet lengths of the communicating parties as filtering conditions to filter out useless information or information that does not need to be processed temporarily.
[0042] The format extraction includes three steps: field division, structure recognition, and semantic annotation;
[0043] The field division uses a sequence alignment algorithm to compare and fill in the data, obtain the similarity between the data to preliminarily infer the meaning of certain fields, and divide the data into dynamic and static fields at the same time;
[0044] Structural recognition processes the message content based on field division and preliminarily obtains the protocol format;
[0045] Semantic annotation is to infer the meaning of each field after division, and at the same time, a large amount of data is processed to verify the correctness of the conclusion.
[0046] The state machine inference and construction is based on semantic annotation, and the fields related to the protocol state are found from the message sequence annotated with semantics, so as to construct the state machine.
[0047] By establishing the possible potential correspondence between fields and semantics, using machine learning methods, continuously deepening the big data model, and establishing a comparison table between fields and semantics, it is possible to accurately grasp irregular or less regular data by looking up the table.
[0048] The potential correspondence between fields and semantics for the following premises is as follows:
[0049] (1): The corresponding relationship between the second-layer data and the first-layer fields has been found;
[0050] (2): The semantics of the second-layer data has been parsed;
[0051] (3): No specific construction method was found for the first layer of data;
[0052] Establish a corresponding table between the second-layer data and the field in the first-layer data, and complete the analysis of the field by looking up the table;
[0053] Layer 1 data: binary data used by the virtual controller and the virtual host for link layer communication;
[0054] Layer 2 data: text data uploaded by data stream in virtual host;
[0055] Model training: The second-layer data is uploaded from the first-layer data, and there is a corresponding relationship between the second-layer data and the first-layer data.
[0056] The step 3 is specifically as follows:
[0057] For binary protocols, the data preprocessing and semantic recognition described above are used for parsing, which is the focus of parsing industrial control protocols.
[0058] For text-based protocols, taking advantage of their readability, one approach is to check the ASCII code displayed by the Wireshark software for hexadecimal data to see if there are specific fields to display as text; the other approach is to check the text in the trace stream; for the same system, the trace stream in the second layer usually comes from the binary protocol in the first layer. Through preliminary analysis of the protocol and the readability of the second layer data, the correspondence between the first and second layers can be found. Through this relationship, a large amount of readable data from the second layer can be obtained, thereby promoting the analysis of the first layer data.
[0059] In the later stages of development, binary and textual protocols were parsed simultaneously, with each verifying the other. For binary protocols, the key was to repeatedly perform step 2, continuously adjusting the filtering conditions during parsing to conduct in-depth analysis of different fields. Simultaneously, the sequence alignment algorithm was continuously adjusted and optimized based on actual results to accelerate subsequent semantic recognition. This enabled more accurate identification of fields related to protocol status, allowing the construction of a state machine as the foundation for further development of textual protocols.
[0060] The step 4 is specifically as follows:
[0061] For large amounts of data, algorithms are used for programming, thereby shortening the parsing cycle; for fields whose semantics have been guessed but the specific algorithms have not been parsed, a large model is used to establish a comparison table, making table lookup possible; for how to construct data through table lookup, a database is used to save the required data, and it is dynamically organized and used; for situations where the parsing results are not obvious or are relatively confusing, a specific UI page is designed, which nests the parsing algorithm toolbox and the complete parsing process to achieve an overall display of the parsing results.
[0062] The specific UI page is specifically:
[0063] The timestamp, millisecond timestamp, quality bit, and command value in the second-layer data are converted into first-layer data according to different conversion methods. The location of the data packet and the corresponding first-layer data are found in the corresponding first-layer pcapng file to analyze the relationship between the second-layer data and the first-layer data and the conditions that the data can meet before being uploaded to the second layer.
[0064] When a piece of extracted data needs to be parsed, the Data payload of the data packet is used as input. The software analyzes the input data, determines its type, outputs the data in segments, and performs in-depth analysis.
[0065] The step 5 is specifically as follows:
[0066] The captured data is divided into data sets and training sets. The hypothesis is obtained by analyzing the data set, and then verified using the training set. After the verification is passed, the capture environment needs to be changed and multiple verifications need to be performed to ensure that the conclusion is not accidental.
[0067] During this process, we built a deep learning model to shorten the verification cycle, and also organized the parsing process into an algorithm. We hope that the model can learn the algorithm and thus acquire the ability to automatically parse.
[0068] During the training process, a variety of optimization algorithms and techniques were used, such as adjusting the learning rate and increasing the diversity of training data, to improve the performance of the model;
[0069] With the assistance of deep learning technology, the meaning of unknown fields can be parsed more accurately, making protocol parsing more complete and accurate, and meeting the needs of complex industrial network protocol parsing.
[0070] Beneficial effects of the present invention:
[0071] 1. Comprehensive and in-depth analysis: The multi-method integrated analysis method analyzes the same protocol from multiple aspects. Compared with using a single analysis method alone, it can establish a more complete and accurate mapping relationship between the protocol algorithm and the protocol format. This method not only improves the accuracy of the analysis, but also increases the depth of the analysis, excavating the potential value in the protocol. In practical applications, it is conducive to a deeper study of the working principles of industrial network protocols and provides a foundation for realizing industrial production optimization and control.
[0072] 2. Advantages of unknown field parsing: The application of deep learning technology can effectively enhance the ability to parse unknown fields. In more complex industrial network protocols, the corresponding unknown field meanings can be parsed more accurately, making the protocol parsing effect better, more complete and accurate, so that the present invention can better meet the needs of coping with complex industrial network protocol parsing tasks, allowing the present invention to better complete the protocol parsing work for different scenarios, effectively improving the technical strength of the present invention.
[0073] 3. Promoting Industrial Intelligence: Industrial network intelligence, which is beneficial to the development of intelligent industrial networks, can drive the development of intelligent industry by providing secure and efficient data processing capabilities. This can enable efficient management and control of industrial production processes at the enterprise level, improving production efficiency and reducing production costs, possessing significant advantages and application value. Industrial network protocol analysis can achieve relatively good results, enabling better services to a certain extent. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] Figure 1 It is the overall flow chart of the analysis process of the present invention.
[0075] Figure 2 This is a flowchart for customized data collection.
[0076] Figure 3 Flowchart for multi-method fusion analysis.
[0077] Figure 4 This is a flowchart for analyzing targeted protocols.
[0078] Figure 5 A flowchart to assist algorithm optimization and deep learning.
[0079] Figure 6 This is a multi-layer perception structure diagram.
[0080] Figure 7 This is part of the Python code for the four-layer MLP model.
[0081] Figure 8 To select the loss function and optimizer.
[0082] Figure 9 Set the number of training rounds (200 rounds).
[0083] Figure 10 Predict output for the model. DETAILED DESCRIPTION
[0084] The present invention will be further described in detail below with reference to the accompanying drawings.
[0085] like Figure 1As shown, a protocol parsing method based on a proprietary industrial network includes the following steps:
[0086] Step 1: Customized data collection: Customize scenario-specific design plans to capture specific data packets and obtain data; this will facilitate training of specific models in step 5.
[0087] Step 2: Multi-method fusion analysis: By integrating three methods: packet layering research, protocol analysis and reproduction, and large-scale data comparison, the captured data is analyzed in an orderly and comprehensive manner;
[0088] This multi-method analysis method helps to deeply explore the similarities and differences between the data, facilitating the classification of the protocols in step 3.
[0089] Step 3: Targeted protocol analysis: Use corresponding algorithms to analyze binary and text protocols respectively to ensure the correctness and accuracy of the analysis results. This operation is the basis for step 4 and is combined with step 4 to conduct in-depth data analysis. In addition, step 2 is repeated during the process to ensure the accuracy of step 3.
[0090] Step 4: Data processing and display optimization: Considering the huge amount of data, a program is written to assist in the analysis work. At the same time, big data is used to collaboratively process the suspected but unknown fields during the data processing process. In order to facilitate the display and subsequent use by enterprises, a corresponding UI interface is written. At the same time, combined with the existing results in step 3, the parsed part and the part to be parsed are processed together to explore the meaning of the unparsed fields.
[0091] Step 5: Algorithm optimization and deep learning assistance: The captured data is divided into a data set and a training set, and the training set is used to verify the correctness of the data set analysis results. During the verification process, step 1 is repeated, and the environment is changed multiple times to verify the conclusion.
[0092] In step 1, the wireshark software is deployed on the virtual host and the virtual controller, and then the wireshark software is started to capture data packets, and then related operations are performed. After the packet capture is completed, the file is saved and copied.
[0093] The scenario-based design scheme is divided into single-machine operation scenario and multi-machine collaborative operation scenario.
[0094] The step 1 is specifically as follows:
[0095] For stand-alone device operation scenarios, by setting multiple trigger conditions for changing communication data, such as device offline and online, sending specific commands, and powering on and off, accurate capture of data packets containing specific operations can be achieved.
[0096] The specific operations are:
[0097] For the first collection, the collection point is on the virtual host;
[0098] A certain configuration of the physical controller is offline and online;
[0099] All configurations of physical controllers are offline and online;
[0100] A certain configuration of the virtual controller goes offline and online;
[0101] All configurations of virtual controllers are offline and online;
[0102] Modify the inter-station communication variable output sent from the virtual controller to the physical controller;
[0103] Modify the physical controller from the second layer;
[0104] Turn off the virtual controller;
[0105] Unplug the network cable of the physical controller;
[0106] Restart the second layer.
[0107] For multi-machine collaborative operation scenarios, by setting up multiple actual collection points and capturing multiple data packets between different devices performing the same operation, a large amount of comparative data can be collected to comprehensively monitor the overall operation of the communication system.
[0108] The specific operations are:
[0109] During the process of establishing a link between a virtual controller and a virtual host, the collection points are on the virtual controller and the virtual host, and the collection is done simultaneously;
[0110] The virtual controller can be powered off and on directly;
[0111] Control the virtual controller offline and online.
[0112] The step 2 is specifically as follows:
[0113] In the comprehensive application of data packet layering research, Wireshark software is used to layer the captured data packets and distinguish other layer data from link layer data; the analysis of industrial control protocols is mainly based on link layer data;
[0114] Through the layered view function of the software, you can clearly see the data packet structure, data characteristics and whether the underlying data has special text meaning, and grasp the rules and characteristics of the protocol format;
[0115] In terms of protocol parsing and reproduction, we analyzed the data from the complete process of data preprocessing, format extraction, and state machine establishment in a simulated paper experimental environment, obtained preliminary field division and semantic inference, and on this basis, innovated on semantic inference and compared large amounts of data from different operations.
[0116] The data preprocessing is to filter the captured massive data, and process it using the MAC addresses of the communicating parties, the length of the data packet, etc. as filtering conditions to filter out useless information or information that does not need to be processed temporarily.
[0117] The format extraction includes three steps: field division, structure recognition, and semantic annotation;
[0118] The field division uses a sequence alignment algorithm to compare and fill in the data, obtain the similarity between the data to preliminarily infer the meaning of certain fields, and divide the data into dynamic and static fields at the same time;
[0119] Structural recognition processes the message content based on field division and preliminarily obtains the protocol format;
[0120] Semantic annotation is to infer the meaning of each field after division, and at the same time, a large amount of data is processed to verify the correctness of the conclusion.
[0121] The state machine inference and construction is based on semantic annotation, and the fields related to the protocol state are found from the message sequence annotated with semantics, so as to construct the state machine.
[0122] By adopting the method of machine learning, we continuously deepen the big data model by establishing the possible potential correspondence between fields and semantics, and establish a comparison table between fields and semantics, so that irregular or less regular data can be accurately grasped by looking up the table.
[0123] The possible potential correspondence between fields and semantics is as follows:
[0124] Layer 1 data: binary data used by the virtual controller and the virtual host for link layer communication;
[0125] Layer 2 data: text data uploaded by data stream in virtual host;
[0126] Model training: The second-layer data is uploaded from the first-layer data, and there is a corresponding relationship between the second-layer data and the first-layer data;
[0127] Based on the following premises: 1. The corresponding relationship between the second-layer data and the first-layer fields has been found, 2. The semantics of the second-layer data has been parsed, and 3. The specific construction method of the first-layer data has not been found;
[0128] A corresponding table between the second-layer data and the field in the first-layer data is established, and the analysis of the field can be completed by looking up the table.
[0129] The step 3 is specifically as follows:
[0130] According to the different types of protocols, they are divided into binary protocols and text protocols;
[0131] For binary protocols, parsing is performed based on the data preprocessing and semantic recognition described above. This part serves as the focus of parsing industrial control protocols.
[0132] For text-based protocols, taking advantage of their readability, one approach is to check the ASCII code displayed by the Wireshark software for hexadecimal data to see if there are specific fields to display as text; the other approach is to check the text in the trace stream; for the same system, the trace stream in the second layer usually comes from the binary protocol in the first layer. Through preliminary analysis of the protocol and the readability of the second layer data, the correspondence between the first and second layers can be found. Through this relationship, a large amount of readable data from the second layer can be obtained, thereby promoting the analysis of the first layer data.
[0133] The step 4 is specifically as follows:
[0134] For large amounts of data, algorithms are used for programming, thereby shortening the parsing cycle; for fields whose semantics have been guessed but the specific algorithms have not been parsed, a large model is used to establish a comparison table, making table lookup possible; for how to construct data through table lookup, a database is used to save the required data, and it is dynamically organized and used; for situations where the parsing results are not obvious or are relatively confusing, a specific UI page is designed, which nests the parsing algorithm toolbox and the complete parsing process to achieve an overall display of the parsing results.
[0135] The specific UI page is specifically:
[0136] The timestamps, millisecond timestamps, quality bits, and instruction values in the second-layer data are converted into first-layer data according to different conversion methods. The location of the data packet containing the data and the corresponding first-layer data are found in the corresponding first-layer pcapng file to analyze the relationship between the second-layer data and the first-layer data and the conditions that the data can meet before being uploaded to the second layer.
[0137] Specific UI page software usage instructions:
[0138] Find Data Input
[0139] First, click "Select File" to select the .pcapng file path you want to search. For the sake of convenience, the input interface here already has the initial file path in the input box.
[0140] If no .pcapng file is found or no valid .pcapng file is selected, a pop-up window will appear.
[0141] Select the Layer 2 trace stream data you want to search for. Enter the timestamp, millisecond timestamp, quality bit, and command value into the corresponding input boxes. To the right of the four input boxes are the corresponding check boxes. To use them as search criteria, select the corresponding check boxes.
[0142] In most cases, you need to check all four data points and there is a "Select All" check box in the lower right corner of the interface; in addition, when searching, the weights of the four data points are different, so the corresponding weights of the data points are provided on the right side of the check box, and the value range is set from 1 to 5, which can be changed according to your needs.
[0143] After confirming the above file path, four second-layer data and corresponding weights, click the "Data Conversion Search" button to search the data. The search results will be displayed in the new pop-up interface.
[0144] Search result export
[0145] To facilitate further analysis and processing of the search results, click the "Print" button on the input interface to export the data as a Word file.
[0146] Parsing interface: When a piece of extracted data needs to be parsed, the Data payload of the data packet is used as input. The software analyzes the input data, determines its type, outputs the data in segments and performs in-depth parsing, which can reduce manual segmentation, data conversion, etc., saving time.
[0147] Instructions for using the parsing interface:
[0148] Input data format requirements
[0149] The input data is the Data payload data, so you need to pay attention to the data format.
[0150] If you copied the data portion from the Wireshark capture software, you need to add the "81" field to the data, as shown below:
[0151] The data obtained at this time is:
[0152] 2501001c002e52000d49303030303030000c2977cb8e0d4930303030303000006cc000560863408700c3029175;
[0153] Add the "81" field as input to the Enter Data entry box:
[0154] 812501001c002e52000d49303030303030000c2977cb8e0d4930303030303000006cc000560863408700c3029175;
[0155] If the input data is incorrect, there will be an output prompt in the output box.
[0156] If the payload data is obtained by overlaying the .payload in Python code, the above processing is not required.
[0157] Parsing Operation
[0158] Click the Parse button to trigger the data parsing operation. The data entered into the input box can be modified at will. If the data is changed, there is no need to rerun the program. Click the Parse button again to enter the parsing of the new data. You can perform data parsing as many times as you want on this interface.
[0159] Data Output
[0160] The output box provides an in-depth analysis of the data in the input box. Different methods are used to parse the data based on the data type. The parsed data is printed to the output box in alternate lines. A scroll bar is attached to the right side of the output box for easy viewing of the data.
[0161] The step 5 is specifically as follows:
[0162] The captured data is divided into data sets and training sets. The hypothesis is obtained by analyzing the data set, and then verified using the training set. After the verification is passed, the capture environment needs to be changed and multiple verifications need to be performed to ensure that the conclusion is not accidental.
[0163] During this process, we built a deep learning model to shorten the verification cycle, and also organized the parsing process into an algorithm, hoping that the model could learn the algorithm and thus acquire the ability to automatically parse.
[0164] During the training process, a variety of optimization algorithms and techniques were used, such as adjusting the learning rate and increasing the diversity of training data, to improve the performance of the model;
[0165] With the assistance of deep learning technology, the meaning of unknown fields can be parsed more accurately, making protocol parsing more complete and accurate, and meeting the needs of complex industrial network protocol parsing.
[0166] like Figure 2 As shown, customized data collection:
[0167] Closely focusing on the actual needs of complex industrial systems and analysis processes, we carefully design data collection conditions and build a multi-machine interactive data collection architecture. For single-machine device operation scenarios, by setting various trigger conditions that change communication data, such as device offline and online, sending specific instructions, and powering on and off, we can accurately capture data packets containing specific operations. For multi-machine collaborative operation scenarios, by setting multiple actual collection points, we can capture multiple data packets between different devices performing the same operation, collect a large amount of comparative data, and comprehensively monitor the overall operation of the communication system. Through this customized data collection method, the collected data can accurately reflect the operating status of the equipment, providing a rich and reliable information foundation for subsequent protocol analysis. It also helps the system to advance the progress of the analysis process as a whole, realizing the cycle of designing data solutions based on the analysis progress and collecting data to promote the analysis progress.
[0168] like Figure 3 As shown, the multi-method fusion analysis combines three approaches: packet layering, protocol parsing and replication, and large-scale data comparison. Leveraging the powerful Wireshark software, captured packets are layered to distinguish data from other layers and link layer data. The analysis of industrial control protocols focuses on link layer data. The software's layered view function also clearly demonstrates packet structure, data characteristics, and whether the underlying data possesses specific textual meaning, allowing for a precise understanding of the patterns and characteristics of the protocol format. Regarding protocol parsing and replication, the team thoroughly studied protocol parsing algorithms from existing papers, analyzing data from data preprocessing, semantic inference, and state machine construction. Simulating the experimental environment of the papers, the team analyzed data to obtain preliminary field segmentation and semantic inference. Building on this, they innovated in semantic inference, comparing large amounts of data from different operations rather than relying solely on clustering algorithms for semantic inference. This approach uncovered underlying patterns within the data from a macro perspective. Furthermore, machine learning methods were employed to establish potential correspondences between fields and semantics, continuously deepening the big data model and establishing a comparison table between fields and semantics. This allows for precise identification of data with irregular or weak patterns through table lookup.
[0169] like Figure 4As shown, targeted protocol analysis: based on the different types of protocols, they are divided into binary protocols and text protocols. For binary protocols, analysis is performed based on the data preprocessing and semantic recognition described above. This part serves as the focus of the analysis of industrial control protocols. For text protocols, taking advantage of their readability, one is to view the ASCII code displayed by the Wireshark software for hexadecimal data to see if there are specific fields to display as text type; the other is to view the text in the trace stream. For the same system, the trace stream in the second layer usually comes from the binary protocol in the first layer. The correspondence between the first and second layers can be found through preliminary analysis of the protocol and the readability of the second layer data. Through this relationship, a large amount of readable data from the second layer can be obtained, thereby promoting the analysis of the first layer data.
[0170] Data processing and display optimization: For large amounts of data, algorithms are used for programming, thereby shortening the parsing cycle; for fields whose semantics have been guessed but the specific algorithms have not been parsed, a large model is used to establish a comparison table, making table lookup possible; for how to construct data through table lookup, the required data is saved in the database, and dynamically organized and used; for situations where the parsing results are unclear or relatively confusing, a specific UI page is designed, which embeds the parsing algorithm toolbox and the complete parsing process to realize the overall display of parsing results.
[0171] like Figure 5 As shown, algorithm optimization and deep learning assistance: The captured data is divided into a dataset and a training set. The dataset is parsed to generate a hypothesis for the analysis, which is then verified using the training set. After verification, the capture environment is modified and multiple verifications are performed to ensure the conclusion is not accidental. During this process, a deep learning model is constructed to shorten the verification cycle. The parsing process is also algorithmically organized, hoping that the model can learn from this algorithm and acquire the ability to automatically parse. During training, a variety of optimization algorithms and techniques are employed, such as adjusting the learning rate and increasing the diversity of training data, to improve model performance. With the assistance of deep learning technology, the meaning of unknown fields can be more accurately parsed, making protocol parsing more complete and accurate, meeting the needs of parsing complex industrial network protocols.
[0172] Example:
[0173] Model training: Establish the corresponding relationship between the known data in the second layer and the corresponding fields in the first layer:
[0174] (1) Dataset preparation
[0175] Extract the second layer data;
[0176] Match the second-layer data with the first-layer data;
[0177] Extract the second-layer known data and the first-layer corresponding fields as data for model training.
[0178] A total of more than 5000 rows of data sets were prepared.
[0179] (2) Model training
[0180] The first and second layer data are independent of time, so the MLP model is used;
[0181] A multilayer perceptron (MLP) is also called an artificial neural network (ANN). In addition to the input and output layers, it can have multiple hidden layers. The simplest MLP contains only one hidden layer, that is, a three-layer structure.
[0182] from Figure 6 As can be seen, each node in the three given layers (input, intermediate, and output layers) of the multilayer perceptron is connected to each node in the adjacent layer (fully connected), so one of the most important components of an MLP is the Dense Layer (fully connected layer, linear layer, dense layer, here called the fully connected layer).
[0183] The fully connected layer has a learnable parameter w (n: the dimension of the input feature, m: the length of the output vector), and a parameter b (bias, length m). So in this layer, we will calculate the following formula on the input data x and get the output y. (y is also a vector of length m)
[0184] y=w·x+b
[0185] The three-layer MLP structure is both simple and capable of fitting arbitrary curves. Therefore, it is often used in deep learning as a small component, a "black box model that can fit arbitrary relationships." Specifically, when the relationship between x and y is unclear, it is simply assumed that y = MLP(x), and then trained using data.
[0186] The model only recognizes numbers, and takes the first layer of data as model input and the second layer of data as model input;
[0187] Define the MLP model and adjust hyperparameters
[0188] After multiple trainings and attempts, the final defined model has a high accuracy rate and can reduce the loss to 0.
[0189] A four-layer MLP model (two hidden layers) is used, such as Figure 7 shown.
[0190] The selection of loss function and optimizer is as follows Figure 8, and dynamically update the learning rate, setting the learning rate to multiply by 0.1 every 50 epochs.
[0191] The ratio of training set to test set is 8:2, training is done for 200 rounds, output the results and save the model. Figure 9 shown.
[0192] (3) Model prediction output
[0193] like Figure 10 As shown in the figure, the saved model is used to predict the output of the four-dimensional input of the new file with a high accuracy.
Claims
1. A protocol parsing method based on a proprietary industrial network, characterized in that: The following steps are included: Step 1: Customize scenario-based design solutions to capture specific data packets and obtain data; Step 2: Conduct an orderly and comprehensive analysis of the captured data by integrating three methods: packet layer analysis, protocol analysis and reproduction, and large-scale data comparison; Step 3: Use the corresponding algorithms to parse the binary protocol and text protocol respectively to ensure the correctness and accuracy of the parsing results; during the progress, repeat step 2 to ensure the accuracy of step 3; Step 4: Write a program to assist in the analysis process. At the same time, use big data to collaboratively process the suspected but unknown fields during the data processing process; write the corresponding UI interface; and combine the existing results in step 3 to jointly process the parsed part with the part to be parsed. In order to mine the meaning of unparsed fields; Step 5: Divide the captured data into a data set and a training set, and use the training set to verify the correctness of the data set analysis results. During the verification process, repeat step 1 and change the environment multiple times to verify the conclusion.
2. The protocol parsing method based on a proprietary industrial network according to claim 1, characterized in that: The capture of a specific data packet in step 1 is specifically as follows: Deploy Wireshark software on the virtual controller and virtual host, then start Wireshark to capture data packets, and then perform related operations. After the packet capture is completed, save the file and copy it; The scenario-based design scheme is divided into single-machine operation scenario and multi-machine collaborative operation scenario.
3. The protocol parsing method based on a proprietary industrial network according to claim 2, characterized in that: For stand-alone device operation scenarios, by setting multiple trigger conditions for changing communication data, such as device offline and online, sending specific commands, and powering on and off, accurate capture of data packets containing specific operations can be achieved. The specific operations are: For the first collection, the collection point is on the virtual host; A certain configuration of the physical controller is offline and online; All configurations of physical controllers are offline and online; A certain configuration of the virtual controller goes offline and online; All configurations of virtual controllers are offline and online; Modify the inter-station communication variable output sent from the virtual controller to the physical controller; Modify the physical controller from the second layer; Turn off the virtual controller; Unplug the network cable of the physical controller; Restart the second layer.
4. The protocol parsing method based on a proprietary industrial network according to claim 3, characterized in that: For multi-machine collaborative operation scenarios, by setting up multiple actual collection points and capturing multiple data packets between different devices in the same operation, a large amount of comparative data can be collected to comprehensively monitor the overall operation of the communication system. The specific operations are as follows: During the process of establishing a link between a virtual controller and a virtual host, the collection points are on the virtual controller and the virtual host, and the collection is done simultaneously; The virtual controller can be powered off and on directly; Control the virtual controller offline and online.
5. The protocol parsing method based on a proprietary industrial network according to claim 1, characterized in that: The specific steps of the three methods in step 2 are: In the comprehensive application of data packet layering research, Wireshark software is used to layer the captured data packets and distinguish other layer data from link layer data; the analysis of industrial control protocols is mainly based on link layer data; Through the layered view function of Wireshark software, you can clearly see the packet structure, data characteristics, and whether the underlying data has special text meaning, and understand the rules and characteristics of the protocol format; In terms of protocol parsing and reproduction, we analyzed the data from the complete process of data preprocessing, format extraction, and state machine establishment in a simulated paper experimental environment to obtain preliminary field division and semantic inference; On this basis, a large amount of data of different operations is compared based on semantic inference.
6. The protocol analysis method based on a proprietary industrial network according to claim 5, characterized in that: The data preprocessing is to filter the captured massive data, using the MAC addresses and data packet lengths of the communicating parties as filtering conditions to filter out useless or temporarily unprocessed information; The format extraction includes three steps: field division, structure recognition, and semantic annotation; The field division uses a sequence alignment algorithm to compare and fill in the data, obtain the similarity between the data to preliminarily infer the meaning of certain fields, and at the same time divide the data into dynamic and static fields; The structure recognition processes the message content based on field division to preliminarily obtain the protocol format; The semantic annotation is to infer the meaning of each field after division, and at the same time, a large amount of data is processed to verify the correctness of the conclusion; The state machine inference and construction is based on semantic annotation, and the fields related to the protocol state are found from the message sequence annotated with semantics, so as to construct the state machine.
7. The protocol analysis method based on a proprietary industrial network according to claim 6, characterized in that: By establishing the potential correspondence between fields and semantics, using machine learning methods to continuously deepen the big data model, and establishing a comparison table between fields and semantics, it is possible to accurately grasp irregular or less regular data by looking up the table; The potential correspondence between fields and semantics is specific to the following premises: (1): The corresponding relationship between the second-layer data and the first-layer fields has been found; (2): The semantics of the second-layer data has been parsed; (3): No specific construction method was found for the first layer of data; Establish a corresponding table between the second-layer data and the field in the first-layer data, and complete the analysis of the field by looking up the table; Layer 1 data: binary data used by the virtual controller and the virtual host for link layer communication; Layer 2 data: text data uploaded by data stream in virtual host; Model training: The second-layer data is uploaded from the first-layer data, and there is a corresponding relationship between the second-layer data and the first-layer data.
8. The protocol analysis method based on a proprietary industrial network according to claim 5, characterized in that: The step 3 is specifically as follows: For binary protocols, the data preprocessing and semantic recognition described above are used for parsing, which is the focus of parsing industrial control protocols. For text-based protocols, taking advantage of their readability, one approach is to check the ASCII code displayed by the Wireshark software for hexadecimal data to see if there are specific fields to display as text; the other approach is to check the text in the trace stream; for the same system, the trace stream in the second layer usually comes from the binary protocol in the first layer. Through preliminary analysis of the protocol and the readability of the second layer data, the correspondence between the first and second layers can be found. Through this relationship, a large amount of readable data from the second layer can be obtained, thereby promoting the analysis of the first layer data.
9. The protocol parsing method based on a proprietary industrial network according to claim 8, characterized in that: The step 4 is specifically as follows: For large amounts of data, algorithms are used for programming, thus shortening the analysis cycle; For fields whose semantics have been guessed but whose specific algorithms have not been parsed, a large model is used to create a comparison table, making table lookup possible; How to look up tables to construct data, use databases to save required data, and dynamically organize and use it; In order to solve the problem that the analysis results are not obvious and are relatively confusing, a specific UI page is designed, which nests the toolbox of the analysis algorithm and the complete analysis process to realize the overall display of the analysis results.
10. A protocol parsing method based on a proprietary industrial network according to claim 9, characterized in that: The specific UI page is specifically: The timestamp, millisecond timestamp, quality bit, and command value in the second-layer data are converted into first-layer data according to different conversion methods. The location of the data packet and the corresponding first-layer data are found in the corresponding first-layer pcapng file to analyze the relationship between the second-layer data and the first-layer data and the conditions that the data can meet to be uploaded to the second layer. When a piece of extracted data needs to be parsed, the data payload of the data packet is used as input. The software analyzes the input data, determines its type, outputs the data in segments, and performs in-depth parsing. The step 5 is specifically as follows: The captured data is divided into data sets and training sets. The hypothesis is obtained by analyzing the data set, and then verified using the training set. After the verification is passed, the capture environment needs to be changed and multiple verifications need to be performed to ensure that the conclusion is not accidental. During this process, a deep learning model is built to shorten the verification cycle, and the parsing process is also algorithmically organized.