A method for reverse engineering communication protocols based on publicly available digital power grid documents
By using a reverse engineering method based on the communication protocol of publicly available digital power grid documents, and by marking and transforming text blocks using a network model to generate and optimize tag sequences, the problem of low efficiency in communication protocol reverse engineering is solved, and more efficient and accurate finite state machine determination is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-02
- Publication Date
- 2026-03-10
AI Technical Summary
The problem of low reverse engineering efficiency in existing communication protocols mainly relies on manual analysis and interpretation of protocol specification documents, resulting in low efficiency and limitations due to the level of artificial intelligence.
By using a reverse engineering method based on publicly available digital power grid documentation for communication protocols, text blocks are marked and transformed using first and second network models to generate a tag sequence. The tag sequence is then optimized to ultimately determine the finite state machine of the communication protocol.
It improves the efficiency and accuracy of communication protocol reverse engineering, enabling faster and more accurate determination of finite state machines, thus solving the problem of low efficiency in traditional methods.
Smart Images

Figure CN119697286B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of information security, and more specifically, to a method for reverse engineering a communication protocol based on publicly available digital power grid documentation. Background Technology
[0002] Reverse engineering communication protocols is a critical task, particularly in the field of digital power grid communications, where it involves parsing and understanding the structure and characteristics of digital power grid protocols. Traditional methods primarily rely on manual analysis and interpretation of protocol specification documents; however, this approach is inefficient, time-consuming, and limited by the level of artificial intelligence.
[0003] This indicates that the related technologies suffer from low reverse engineering efficiency of deterministic communication protocols.
[0004] There is currently no effective solution to the aforementioned problems in the relevant technologies. Summary of the Invention
[0005] This invention provides a communication protocol reverse engineering method based on publicly available digital power grid documentation, which at least solves the problem of low efficiency in communication protocol reverse engineering in related technologies.
[0006] According to an embodiment of the present invention, a method for reverse engineering a communication protocol based on a publicly available digital power grid document is provided, comprising: determining a target document including a communication protocol; inputting each text block included in the target document into a first network model to label each text block to obtain a first tag sequence, wherein the first tag sequence includes each text block and tag information of the text block; converting each text block into a text vector through a second network model and labeling the text vector to obtain a second tag sequence, wherein the second tag sequence includes each text vector and tag information of the text vector; optimizing the first tag sequence and the second tag sequence to obtain a target tag sequence; and determining a finite state machine of the communication protocol based on the target tag sequence to complete the reverse engineering of the communication protocol.
[0007] According to another embodiment of the present invention, a reverse engineering device for a communication protocol based on publicly available digital power grid documents is provided, comprising: a first determining module, configured to determine a target document including a communication protocol; a first marking module, configured to input each text block included in the target document into a first network model to mark each text block to obtain a first tag sequence, wherein the first tag sequence includes each text block and tag information of the text block; a second marking module, configured to convert each text block into a text vector through a second network model and mark the text vector to obtain a second tag sequence, wherein the second tag sequence includes each text vector and tag information of the text vector; an optimization module, configured to optimize the first tag sequence and the second tag sequence to obtain a target tag sequence; and a second determining module, configured to determine a finite state machine of the communication protocol based on the target tag sequence to complete the reverse engineering of the communication protocol.
[0008] According to yet another embodiment of the present invention, a computer-readable storage medium is also provided, wherein a computer program is stored therein, wherein the computer program is configured to perform the steps in any of the above method embodiments when executed.
[0009] According to yet another embodiment of the present invention, an electronic device is also provided, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.
[0010] According to yet another embodiment of the present invention, a computer program product is also provided, including a computer program that, when executed by a processor, implements the steps of the methods described in various embodiments of the present application.
[0011] This invention identifies a target document containing a communication protocol; inputs each text block in the target document into a first network model to label each text block, obtaining a first tag sequence, wherein the first tag sequence includes each text block and its tag information; converts each text block into a text vector using a second network model, and labels the text vectors, obtaining a second tag sequence, wherein the second tag sequence includes each text vector and its tag information; optimizes the first and second tag sequences to obtain a target tag sequence; and determines a finite state machine for the communication protocol based on the target tag sequence. Since each text block in the target document can be input into the first network model to label the text block and obtain the first tag sequence, and the text block can be converted into a text vector using the second network model, and the text vector can be labeled to obtain the second tag sequence, and the first and second tag sequences can be optimized to obtain the target tag sequence, the finite state machine for the communication protocol can be determined based on the target tag sequence to complete the reverse engineering of the communication protocol. The first label sequence is obtained through a first network model, and the second label sequence is obtained through a second network model, which improves the speed of determining the label sequence. Optimizing the first and second label sequences and training a finite state machine based on the optimized target labels, and then using the optimized label sequence to determine the finite state machine, helps improve the accuracy of the finite state machine determination. Therefore, this addresses the problem of low reverse engineering efficiency in related technologies, achieving the effect of improving both the efficiency and accuracy of communication protocol reverse engineering. Attached Figure Description
[0012] Figure 1 This is a hardware structure block diagram of a mobile terminal based on a communication protocol reverse engineering method from a publicly available digital power grid document, according to an embodiment of the present invention.
[0013] Figure 2 This is a flowchart of a communication protocol reverse engineering method based on publicly available digital power grid documentation, according to an embodiment of the present invention.
[0014] Figure 3 This is a flowchart of a communication protocol reverse engineering method based on publicly available digital power grid documents according to a specific embodiment of the present invention;
[0015] Figure 4 This is a structural block diagram of a communication protocol reverse engineering device based on publicly available digital power grid documents according to an embodiment of the present invention. Detailed Implementation
[0016] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings and examples.
[0017] It should be noted that the terms "first," "second," etc., in the specification, claims, and drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0018] The methods and embodiments provided in this application can be executed on a mobile terminal, computer terminal, or similar computing device. Taking running on a mobile terminal as an example, Figure 1 This is a hardware structure block diagram of a mobile terminal based on a communication protocol reverse engineering method from a publicly available digital power grid document, according to an embodiment of the present invention. Figure 1 As shown, a mobile terminal may include one or more ( Figure 1 Only one is shown in the diagram. A processor 102 (which may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.) and a memory 104 for storing data are also shown. The mobile terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the mobile terminal described above. For example, the mobile terminal may also include components that are more... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.
[0019] The memory 104 can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to the communication protocol reverse engineering method based on publicly available digital power grid documents in this embodiment of the invention. The processor 102 executes various functional applications and data processing by running the computer programs stored in the memory 104, thereby implementing the aforementioned method. The memory 104 may include high-speed random access memory and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to a mobile terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0020] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the mobile terminal's communication provider. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.
[0021] This embodiment provides a method for reverse engineering communication protocols based on publicly available digital power grid documents. Figure 2 This is a flowchart of a communication protocol reverse engineering method based on publicly available digital power grid documentation according to an embodiment of the present invention, such as... Figure 2 As shown, the process includes the following steps:
[0022] Step S202: Determine the target document, including the communication protocol;
[0023] Step S204: Input each text block included in the target document into the first network model to label each text block and obtain a first label sequence, wherein the first label sequence includes each text block and the label information of the text block;
[0024] Step S206: Convert each text block into a text vector using a second network model, and label the text vectors to obtain a second label sequence, wherein the second label sequence includes each text vector and the label information of the text vector;
[0025] Step S208: Optimize the first tag sequence and the second tag sequence to obtain the target tag sequence;
[0026] Step S210: Determine the finite state machine of the communication protocol based on the target tag sequence, thus completing the reverse engineering of the communication protocol.
[0027] In the above embodiments, the communication protocol reverse engineering method can be applied to the field of digital power grids, the field of network communication, and other fields where standard protocols exist. When the communication protocol reverse engineering method is applied to different fields, the type of communication protocol can be different. For example, when applied to the field of digital power grids, the communication protocol can be a digital power grid communication protocol, and the target document is a document that includes the digital power grid communication protocol. When applied to the field of network communication, the communication protocol can be a network communication protocol, and the target document is a document that includes the network communication protocol. In this invention, the application of a finite state machine determination method to the field of digital power grids, the communication protocol being a digital power grid communication protocol, and the target document being a document that includes the digital power grid communication protocol will be used as an example for explanation.
[0028] In the above embodiments, the target document can be publicly available documentation in the field of digital power grids, including standards, protocols, research papers, technical forums, blogs, etc. The target document may include specifications and descriptions related to various digital power grid communication protocols. Target documents related to digital power grids can be collected in real time from various public sources in the field, such as standards organization websites, academic paper databases, and technical forums, ensuring access to the latest and most comprehensive digital power grid documentation resources, providing a sufficient material foundation for subsequent analysis and processing.
[0029] Following the above embodiments, after obtaining the target document, it can be segmented, such as dividing text paragraphs into multiple text blocks. Each text block is then input into a sequence-to-sequence model, which may include a first network model and a second network model. An initial sequence-to-sequence model can be trained first to parse digital power grid specification documents and label key information, such as states, events, and variables. Both the first and second network models can employ a BIO (Beginning, Internal, External) tagging system to accurately identify entities in the text and assign them corresponding labels. The first network model can be a convolutional neural network conditional random field (CRF), and the second network model can be a neural CRF model based on BERT embeddings. The initial models of the first and second network models can be trained using a large number of digital power grid documents to adapt to the language characteristics of the digital power grid domain. Training the first network model, i.e., the convolutional neural network conditional random field, can process a set of extracted features on each text unit to learn the sequence-to-sequence mapping. It models predictions as a probabilistic graphical model, specifically considering the order dependencies in the predictions. The second network model, a neural CRF model based on BERT embeddings, is trained by enhancing the BERT encoder with GRU and CRF layers. The BERT encoder is used to create chunk-level representations from word sequences, and then the GRU computes the embedded text unit sequence to obtain the representation. The neural CRF model based on BERT embeddings can leverage collected digital power grid document information to transform domain-specific technical language into distributed word representations.
[0030] In the above embodiments, the output of the sequence-to-sequence model, namely the first label sequence and the second label sequence, can be post-processed and optimized, including noise removal, error correction, and structure optimization, to obtain the target label sequence. For example, rules can be applied to correct some simple cases that the model failed to recognize, such as the extraction of states, events, and variables, to ensure the accuracy and completeness of the extracted information. Then, the various components of the finite state machine, including the set of states, the set of inputs and outputs, the initial state, and the set of state transitions, can be extracted from the output of the post-processed and optimized sequence-to-sequence model, namely the target label sequence.
[0031] In the above embodiments, the determined finite state machine can be applied to the reverse engineering of communication protocols. When performing reverse engineering of communication protocols, finite state machines can be used to analyze and model the behavior of the protocols. By observing the data interactions and state transitions of the protocols, they can be abstracted into a finite state machine model, thereby helping to understand the working principles and design logic of the protocols. Simultaneously, finite state machines can also be used to verify the results of reverse engineering. By simulating and testing the finite state machine model, the accuracy and completeness of the analysis results can be verified. Therefore, the reverse engineering of communication protocols and finite state machines are closely related. Finite state machines provide a formalized modeling and analysis tool for reverse engineering, helping to understand and crack the working principles of communication protocols.
[0032] This invention identifies a target document containing a communication protocol; inputs each text block in the target document into a first network model to label each text block, obtaining a first tag sequence, wherein the first tag sequence includes each text block and its tag information; converts each text block into a text vector using a second network model, and labels the text vectors, obtaining a second tag sequence, wherein the second tag sequence includes each text vector and its tag information; optimizes the first and second tag sequences to obtain a target tag sequence; and determines a finite state machine for the communication protocol based on the target tag sequence to complete the reverse engineering of the communication protocol. Since each text block in the target document can be input into the first network model to label the text block and obtain the first tag sequence, and the text block can also be converted into a text vector using the second network model, and the text vector can be labeled to obtain the second tag sequence, and the first and second tag sequences can be optimized to obtain the target tag sequence, the finite state machine for the communication protocol can be determined based on the target tag sequence. The first label sequence is obtained through a first network model, and the second label sequence is obtained through a second network model, which improves the speed of determining the label sequence. Optimizing the first and second label sequences and training a finite state machine based on the optimized target labels, and then using the optimized label sequence to determine the finite state machine, helps improve the accuracy of the finite state machine determination. Therefore, this addresses the problem of low reverse engineering efficiency in related technologies, achieving the effect of improving both the efficiency and accuracy of communication protocol reverse engineering.
[0033] Optionally, the entity performing the above steps may be a background processor, or other devices with similar processing capabilities, or a machine that integrates at least a data processing device. The data processing device may include, but is not limited to, terminals such as computers and mobile phones.
[0034] In an exemplary embodiment, converting each text block into a text vector using a second network model and labeling the text vectors to obtain a second label sequence includes: converting the text block into the text vector using a bidirectional encoder representation from transformers (BERT) model included in the second network model; labeling the text vector using a gated recurrent unit (GRU) and a conditional random field (CRF) layer included in the second network model to obtain label information for the text vector; and determining the sequence consisting of the text vector and the label information of the text vector as the second label sequence. In this embodiment, the second network model may include a bidirectional encoder representation from transformers (BERT) model, a gated recurrent unit (GRU), and a conditional random field (CRF) layer. The second network model combines BERT, GRU, and CRF layers to better capture the contextual information of the text sequence and to label it using the dependencies between sequences. The text unit is converted into a representation, i.e., a text vector, using a BERT encoder, and then labeled using the GRU and CRF layers. By jointly training the parameters of the BERT encoder, GRU layer, and CRF layer, the model can be effectively optimized, thereby achieving accurate labeling and information extraction of digital power grid protocol documents.
[0035] In the above embodiment, the second network model can receive text blocks as input and output a sequence of tags corresponding to the grammar. To tag the text, BIO (beginning, inner, outer) tags can be used. Text blocks correspond to paragraphs in an RFC document. Paragraphs are first divided into smaller units (e.g., single words, chunks, or phrases). Each unit is then mapped to a specific tag. A convolutional neural network conditional random field processes a set of extracted features on each chunk. Conditional random fields model predictions as probabilistic graphical models; chained conditional random fields specifically consider order dependencies in predictions. Let y be the tag sequence and x be the input sequence of text units. The goal is to maximize the conditional probability:
[0036] Where f is in the feature vector x t The scoring function is learned using the parameter vector θ. p(y,x) represents the joint probability of a given input sequence x and label sequence y, i.e., the probability that both the input and label sequences occur simultaneously; p(x,y) represents the joint probability of a given label sequence y and input sequence x, i.e., the probability that both the label and input sequences occur simultaneously. y' represents all possible label sequences, used to sum over all possible label sequences; ∑ y' p(y', x) represents the sum of the joint probabilities of all possible label sequences given the input sequence x; y tLet represent the t-th label in the label sequence, and y t-1 Let represent the (t-1)th label in the label sequence. They are used to represent the label at the current position and the label at the previous position, respectively. To understand θ, the negative log-likelihood log P(y, x) is minimized. The partition function is calculated using a forward-backward algorithm: Z(x) = ∑ y′ p(y′,x). Z(x) represents the sum of the joint probabilities of all possible label sequences given the input sequence x, where y′ represents all possible label sequences. In a Conditional Random Field (CRF), Z(x) is used to normalize the conditional probability p(y|x), ensuring that the sum of the probabilities of all possible label sequences is 1.
[0037] The BERT-embedded Neural CRF model (NEURALCRF) combines gated recurrent units (GRUs) and CRF layers to better capture the contextual information of text sequences and leverage dependencies between sequences for labeling. A BERT encoder transforms text units into representations, which are then labeled using GRU and CRF layers. By jointly training the parameters of the BERT encoder, GRU layers, and CRF layers, the model can be effectively optimized, enabling accurate labeling and information extraction of digital power grid protocol documents.
[0038] A CRF layer can be added to these representations by replacing the function f of the conditional probability: Where, x t h represents the input of this text unit. t The text unit is represented by the model computation, and P is the learned parameter matrix representing the transition between labels. Similar to the case of RNNCRF, the negative log-likelihood log P(y, x) is minimized to jointly learn the parameters of the BERT encoder, GRU layer, and transition matrix P.
[0039] In an exemplary embodiment, before converting the text block into the text vector using a bidirectional encoder-representation converter model included in the second network model, the method further includes: obtaining a target grammar for the structure of a document describing the communication protocol; obtaining a training document including the communication protocol; and performing unsupervised learning on an initial bidirectional encoder-representation converter model based on the target grammar and the training document to obtain the bidirectional encoder-representation converter model. In this embodiment, a general target grammar can be defined in the digital power grid protocol document to describe the structure and characteristics of the network protocol, covering related concepts such as states, events, and variables. The target grammar may include state definition tags: used to annotate the names of states, events, and variables associated with each protocol. These tags identify and define various states in the text, including network connection states, protocol execution states, etc. Through these tags, the various states described in the document can be clearly identified, providing a basis for labeling and identification in subsequent steps. Event definition tags: used to reference states and events in the text, linking to labeled states and events to clarify the terminology used in the text. These tags associate events involved in the text with defined states and variables, helping readers understand the various operations and processes described in the document. State Machine Labels: A set of two labels used to represent the logic of a state machine. These include transition and variable labels. Word embedding models can be pre-trained: technical language embedding models can be built using unsupervised learning algorithms. These labels describe the two parts of the state machine, including transitions between states and the use of variables. Transition Labels: Used to describe the transitional relationships between states in the state machine. In digital power grid protocol documents, transition labels can identify the transition from one state to another, describing the changes in state during protocol execution. Transition labels clearly define the conditions and triggering events for state transitions. Variable Labels: Used to identify and describe the variables or parameters involved in the digital power grid protocol. These variables may include various parameters used in the protocol, configuration options, state information, etc. Variable labels explicitly define the various variables involved in the protocol and guide their use and management during protocol execution.
[0040] In the above embodiments, the BERT model can be subjected to unsupervised learning based on the training documents and the target grammar. By utilizing the collected digital power grid document information, a technical language embedding model, such as the BERT model, is constructed. The goal of this model is to transform the proprietary technical language of the digital power grid domain into distributed word representations. This approach captures the semantic information between words, leading to a better understanding of the text content. Through the BERT model, words in the digital power grid domain are represented as distributed vectors, capturing their semantic relationships and meanings. This transforms proprietary terminology in the digital power grid domain into computer-understandable vector representations, providing a foundation for subsequent model training and semantic analysis.
[0041] In an exemplary embodiment, optimizing the first label sequence and the second label sequence to obtain a target label sequence includes: determining state text included in the first label sequence and the second label sequence, wherein the state text includes state information and verbs or directional prepositions; determining event text included in the first label sequence and the second label sequence, wherein the event text includes action verbs; determining first text included in the first label sequence and the second label sequence, wherein the first text is not marked with a span and includes state text or event text; determining second text included in the first label sequence and the second label sequence, wherein the second text is not marked with a span; and optimizing the state text, the event text, the first text, and the second text to obtain the target label sequence. In this embodiment, a set of rules can be used to correct some simple cases that the model fails to recognize. These rules are applied to the classification output to correct relevant cases by flipping the labels. These rules include: Transition span marking: finding text units that mention a state. If the unit mentions a state and contains a transition verb or directional preposition, the unit is marked as a transition span. Action span marking: finding text units that mention an event. If the cell mentions an event and contains an action verb (such as send, receive), mark the cell as an action span. Unmarked span marker: Mark any remaining unmarked spans and mention the state or event as the trigger. Variable marker: Mark any remaining unmarked spans as variable names.
[0042] In the above embodiments, the state text can refer to the state, that is, the state of the protocol, such as LISTEN for listening and RESPONSE for responding. Unlabeled spans can be understood as span labels that have not been recognized by the second network model. The target name can depend on the specific protocol, such as sliding window variables or congestion control variables in the TCP protocol, and can identify all remaining unlabeled content as variables.
[0043] In an exemplary embodiment, optimizing the first tag sequence and the second tag sequence to obtain a target tag sequence includes: marking the state text as a transition span; marking the event text as an action span; marking the first text as an unmarked span; marking the second text as a target name; and determining the sequence including the state text and the transition span, the event text and the action span, the first text and the unmarked span, the second text and the target name as the target tag sequence. In this embodiment, transition span marking involves: first, identifying text units that mention a state. If a unit mentions a state and contains a transition verb or a directional preposition, the unit is marked as a transition span. This helps identify portions of the text that describe state transitions. For example, when the text describes "from CLOSED to OPEN state," "from CLOSED to OPEN" is identified as a transition span to indicate the state transition from CLOSED to OPEN state. Action span marking involves: second, identifying text units that mention an event. If a unit describes an event and contains an action verb, the unit is marked as an action span. This helps identify the parts of the text that describe the action of an event. For example, when the text describes "sending a data packet," "sending a data packet" is marked as an action span to indicate the event that performs the sending operation. Unmarked span marker: Marks any remaining unmarked spans and mentions the state or event as the trigger. Variable marker: Marks any remaining unmarked spans as variable names, i.e., target names.
[0044] The introduction of these tags enables the model to more accurately identify state transitions and events in the text, thereby improving the overall tagging quality and parsing performance. By using these rules, some situations that the model failed to accurately identify can be corrected, further improving the performance and effectiveness of the inverse method.
[0045] In an exemplary embodiment, determining the finite state machine of the communication protocol based on the target label sequence includes: determining the set of states, the set of input / output events, and the set of transition relationships included in the target label sequence; and determining the set of states, the set of input / output events, and the set of transition relationships as the finite state machine. In this embodiment, the various components of the finite state machine can be extracted from the output of the post-processed optimized sequence to the sequence model. This may include a set of states, a set of input / output events, an initial state, and a set of state transitions. By analyzing the labeled sequence of the model output, the structure of the finite state machine is identified, providing important support for subsequent digital power grid system development and application. These transition relationships describe the transitions between states and the associated events. The set of states can be extracted from the output, where some states may be labeled as initial states. All states can be extracted from the model output, and some of these states can be identified as initial states. These states represent the states of the system at different stages, while the initial state identifies the initial state of the system. The sets of input and output events are extracted based on the information in the model output. The model output is analyzed to identify the sets of input and output events. Input events represent external inputs received by the system, while output events represent external outputs generated by the system. By extracting these event sets, we can understand the system's behavior and function. We analyze the transition information in the model output to extract the set of transition relationships between states. These transition relationships describe the system's transitions between different states and the associated events. By analyzing this information, we can construct the transition diagram of the finite state machine, thus fully describing the system's behavior and state transitions.
[0046] A protocol state machine, also known as a finite state machine, can be represented as P =<S,I,O,s0,T> A finite set of states S, containing a finite number of states, representing the various states the system can be in. A finite set of inputs I, containing a finite number of inputs, representing the external inputs received by the system. A finite set of outputs O, containing a finite number of outputs, representing the external outputs generated by the system. The output set does not overlap with the input set I. An initial state s0 ∈ S represents the initial state of the system. A finite set of transitions. It contains a finite number of transition relationships between system states. Each transition relationship consists of three parts: the starting state, the transition label (which can be empty, timeout, or input / output), and the target state.
[0047] In one exemplary embodiment, determining a target document including a communication protocol includes: acquiring a first initial document including the communication protocol; performing data cleaning and format standardization operations on the first initial document to obtain a second initial document; and preprocessing the second initial document using natural language processing techniques to obtain the target document. In this embodiment, firstly, a wide range of publicly available documents in the field of digital power grids are collected, including standards, protocols, research papers, etc. These documents may involve specifications and descriptions of various digital power grid communication protocols. Then, the collected documents are preprocessed, including removing formatting errors, standardizing text format, and segmenting text paragraphs, to facilitate subsequent processing and analysis.
[0048] In the above embodiments, digital power grid-related documents can be collected in real time from various publicly available sources in the field of digital power grids, such as standards organization websites, academic paper databases, and technical forums, to obtain a first initial document, which is then stored as structured data. This ensures access to the latest and most comprehensive digital power grid document resources, providing a sufficient material foundation for subsequent analysis and processing. The collected documents, i.e., the first initial document, undergo text cleaning and format standardization processing, including removing HTML tags, converting to a unified encoding format, and eliminating non-text content. Noise and interference in the document are eliminated to obtain a second initial document, ensuring that subsequent processing can be based on clean and standardized text data. The standardized text data, i.e., the second initial document, undergoes preprocessing, including natural language processing techniques such as word segmentation, stemming, and part-of-speech tagging. The original text data is transformed into a computer-processable form, extracting key information and features to prepare for model building and training in subsequent steps.
[0049] The following describes a method for reverse engineering communication protocols based on publicly available digital power grid documentation, using specific implementation details. This method involves reverse engineering the DNP3 protocol, a commonly used communication protocol for remote monitoring and control, which is widely applied in power systems.
[0050] Figure 3 This is a flowchart of a communication protocol reverse engineering method based on publicly available digital power grid documentation, according to a specific embodiment of the present invention. Figure 3 As shown, the process includes:
[0051] Step S302: Receive documents, clean the text, and standardize the format.
[0052] The collection of DNP3 protocol-related documents covered its official standard documents, technical specifications, RFC documents, and related research papers. These documents contain key information about the DNP3 protocol, including its communication mechanism, data format, and message structure. For example, the official standard documents include content such as data frame structure, function codes, and control fields. A typical DNP3 data frame includes the following tags:
[0053] Control Field: Contains control information for DNP3 data frames, used to indicate frame type, master-slave communication mode, etc.
[0054] Address Field: Used to indicate the source and destination addresses of a data frame in order to identify the target device for communication.
[0055] Function Code: Indicates the type of operation performed on the data frame, such as reading data or writing data.
[0056] Data Field: Contains the actual data content, such as information collected by sensors and control commands.
[0057] Step S304, Text Preprocessing. In the preprocessing stage, firstly, HTML tags and non-text content are removed from the document, and then the document's encoding format is standardized to UTF-8 to ensure the accuracy and consistency of subsequent processing.
[0058] Step S306, Pre-trained word embedding model: Use unsupervised learning algorithms to construct a technical language embedding model, represent words in the digital power grid field as distributed vectors, and capture the semantic relationships and meanings between them.
[0059] By collecting information from DNP3 protocol documentation, a technical language embedding model, namely the BERT model, was constructed. This model converts proprietary terms and concepts in the DNP3 protocol into distributed word representations. This embedding model can transform protocol terms into vector representations, thereby capturing the semantic relationships and meanings between terms. For example, terms such as "DNP3 data frame," "master station," and "slave station" are converted into continuous vector representations for computer understanding and processing.
[0060] Step S308, Training a Convolutional Neural Network Conditional Random Field Model (RNNCRF): Using a convolutional neural network and CRF layers to process a set of extracted features on each text unit, thereby learning a sequence-to-sequence mapping.
[0061] Step S310, train a GRU Conditional Random Field model based on BERT embeddings (GRUCRF): the BERT encoder is enhanced with GRU and CRF layers to create chunk-level representations from word sequences.
[0062] When training the sequence-to-sequence model, a preprocessed DNP3 specification document was used, and key information within it was annotated. The model is designed to identify key fields and operation types in data frames, such as control fields and function codes. The model can accurately label these fields and operation types as corresponding events or states. For example, in data frame parsing, the model can accurately identify information such as communication modes and address types in control fields and label them as corresponding events. Simultaneously, the model can also identify function codes, such as read data and write data operations, and label them as corresponding states.
[0063] Step S312, Rule Correction: Use a set of rules to correct some simple cases that the model failed to recognize, including transitions and variable labeling.
[0064] After the model output, post-processing and optimization are performed to improve the accuracy and completeness of the extracted information. This includes detecting and correcting labeling errors caused by unclear or ambiguous document descriptions, removing noise from the model output, and optimizing the structure of the output results. For example, when parsing data frames, control field labeling errors or missing data may occur; the post-processing can correct these errors. Simultaneously, some data frames may be parsed incompletely or incorrectly; these erroneous data can be removed to ensure the accuracy and completeness of the output results.
[0065] Step S314, Finite State Machine Extraction: Extract the state set: Extract the state set from the output, where some states may be labeled as the initial state. Extract the input and output event set: Extract the input and output event set based on the information in the model output. Analyze the transition relationships: Analyze the transition information in the model output to extract the finite state machine's transition relationship set, describing the transitions between states and the associated events.
[0066] Extracting the DNP3 protocol finite state machine from the post-processed optimized sequence to the sequence model output. This includes a set of states, an input / output set, an initial state, and a set of state transitions. For example, during data frame parsing, the state machine can describe the state transition sequence of the master station sending a command to the slave station and waiting for a response. The state set includes the master station state and the slave station state, the input / output set includes the data frame and the response frame, and the state transition set describes the communication process between the master and slave stations.
[0067] Step S316, finite state machine result of digital power grid dedicated protocol.
[0068] In the aforementioned embodiments, technologies such as digital power grid-specific XML markup syntax, BERT model, convolutional neural network conditional random field (RNNCRF), and neural CRF model based on BERT embedding (GRUCRF) are utilized to improve the reverse engineering process of digital power grid communication protocols.
[0069] First, the dedicated XML markup syntax for digital power grids provides the foundation for parsing and understanding digital power grid protocol documents. This syntax defines the general structure and characteristics of the protocol, including related concepts such as states, events, and variables. By using this dedicated markup syntax, key information in the protocol document can be extracted more accurately and efficiently, providing a reliable foundation for subsequent reverse engineering.
[0070] Secondly, the BERT model is a natural language processing model capable of converting text into semantically rich vector representations. The BERT model is used to learn the technical language embeddings in the digital power grid field, thereby better understanding and analyzing textual information in protocol documents. By pre-training the BERT model, complex semantic and contextual information in documents can be captured, improving the accuracy and efficiency of reverse engineering.
[0071] Third, Convolutional Neural Network Conditional Random Field (RNNCRF) is a probabilistic graphical model used for sequence labeling tasks. RNNCRF is used to label text blocks with specific grammatical tags and to identify states, events, and other important elements in the text. By using RNNCRF, the relationships between text blocks can be established, and information in the protocol document can be transformed into actionable sequence tags, providing an important preprocessing step for subsequent finite state machine extraction.
[0072] Furthermore, the BERT-embedded neural CRF model (GRUCRF) combines the semantic representation capabilities of the BERT model with the sequence labeling capabilities of the CRF model. GRUCRF is used to label sequences, thereby identifying and extracting the finite state machine structure in protocol documents. By combining the semantic understanding of the BERT model with the sequence labeling capabilities of the CRF model, GRUCRF can more accurately infer the transition relationships between states, achieving precise extraction and modeling of the protocol structure.
[0073] The present invention can achieve the following beneficial effects:
[0074] Improved accuracy: By converting publicly available documents in the field of digital power grids into intermediate representations and applying a series of rules for post-processing and optimization, the accuracy of information extraction can be improved. Training with a sequence-to-sequence model, combined with techniques such as semantic role labelers, effectively captures key information in the documents, thereby reducing errors and inaccuracies.
[0075] Enhanced Information Comprehensiveness: The finite state machine extraction step can completely extract the set of states, the set of input / output events, and the set of transition relationships between states from the optimized intermediate representation. This extraction ensures that the extracted finite state machine has complete information, comprehensively describing the state changes and event transitions of the digital power grid system, and providing sufficient reference for system development and maintenance.
[0076] Efficiency Optimization: A series of optimization techniques were employed, including post-processing of intermediate representations, optimized model training, and further processing of the extracted results, ensuring the high efficiency of the entire information extraction process. These optimization measures accelerate the information extraction process and reduce the time costs of system development and maintenance.
[0077] Enhanced flexibility: Thanks to its modular design, each step is independent, resulting in high flexibility. Users can selectively apply each step according to their actual needs, or make customized adjustments based on specific domain requirements, thus better adapting to different application scenarios and requirements.
[0078] Enhanced versatility: The technical methods and processing procedures employed in this invention possess strong versatility, applicable not only to the field of digital power grids but also extending to text information extraction tasks in other fields. By changing the model's input data and adjusting relevant parameters, this invention can be easily applied to text information extraction in other domains, providing a universal solution for information analysis and application in related fields.
[0079] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0080] This embodiment also provides a communication protocol reverse engineering device based on publicly available digital power grid documentation. This device is used to implement the above embodiments and preferred embodiments, and details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0081] Figure 4 This is a structural block diagram of a communication protocol reverse engineering device based on publicly available digital power grid documentation, according to an embodiment of the present invention. Figure 4 As shown, the device includes:
[0082] The first determining module 402 is used to determine the target document, including the communication protocol;
[0083] The first tagging module 404 is used to input each text block included in the target document into the first network model to tag each text block and obtain a first tag sequence, wherein the first tag sequence includes each text block and the tag information of the text block;
[0084] The second labeling module 406 is used to convert each text block into a text vector through a second network model and label the text vectors to obtain a second label sequence, wherein the second label sequence includes each text vector and the label information of the text vector;
[0085] Optimization module 408 is used to optimize the first tag sequence and the second tag sequence to obtain the target tag sequence;
[0086] The second determining module 410 is used to determine the finite state machine of the communication protocol based on the target tag sequence, thus completing the reverse engineering of the communication protocol.
[0087] In an exemplary embodiment, the second labeling module 406 can convert each text block into a text vector through a second network model and label the text vector to obtain a second label sequence in the following manner: the text block is converted into the text vector through a bidirectional encoder-representation converter model included in the second network model; the text vector is labeled through a gated recurrent unit and a conditional random field layer included in the second network model to obtain the label information of the text vector; and the sequence consisting of the text vector and the label information of the text vector is determined as the second label sequence.
[0088] In one exemplary embodiment, the apparatus may further be used to: obtain a target grammar for describing the structure of a document describing the communication protocol before converting the text block into the text vector through a bidirectional encoder-representation converter model included in the second network model; obtain a training document including the communication protocol; and perform unsupervised learning on an initial bidirectional encoder-representation converter model based on the target grammar and the training document to obtain the bidirectional encoder-representation converter model.
[0089] In an exemplary embodiment, the optimization module 408 can optimize the first tag sequence and the second tag sequence to obtain a target tag sequence by: determining the state text included in the first tag sequence and the second tag sequence, wherein the state text includes state information and verbs or directional prepositions; determining the event text included in the first tag sequence and the second tag sequence, wherein the event text includes action verbs; determining the first text included in the first tag sequence and the second tag sequence, wherein the first text is not marked with a span and includes state text or event text; determining the second text included in the first tag sequence and the second tag sequence, wherein the second text is not marked with a span; and optimizing the state text, the event text, the first text, and the second text to obtain the target tag sequence.
[0090] In an exemplary embodiment, the optimization module 408 can optimize the first tag sequence and the second tag sequence to obtain a target tag sequence by: marking the state text as a transition span; marking the event text as an action span; marking the first text as an unmarked span; marking the second text as a target name; and determining the sequence including the state text and the transition span, the event text and the action span, the first text and the unmarked span, the second text and the target name as the target tag sequence.
[0091] In an exemplary embodiment, the second determining module 410 may determine the finite state machine of the communication protocol based on the target tag sequence by: determining the set of states, the set of input / output events, and the set of transition relationships included in the target tag sequence; and determining the set of states, the set of input / output events, and the set of transition relationships as the finite state machine.
[0092] In an exemplary embodiment, the first determining module 402 may determine the target document including the communication protocol by: obtaining a first initial document including the communication protocol; performing data cleaning and format standardization operations on the first initial document to obtain a second initial document; and preprocessing the second initial document using natural language processing technology to obtain the target document.
[0093] It should be noted that the above modules can be implemented by software or hardware. For the latter, they can be implemented in the following ways, but are not limited to: all the above modules are located in the same processor; or, the above modules are located in different processors in any combination.
[0094] Embodiments of the present invention also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to perform the steps in any of the above method embodiments when executed.
[0095] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.
[0096] Embodiments of the present invention also provide an electronic device including a memory and a processor, the memory storing a computer program and the processor being configured to run the computer program to perform the steps in any of the above method embodiments.
[0097] In one exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor and the input / output device is connected to the processor.
[0098] Embodiments of the present invention also provide a computer program product, including a computer program that, when executed by a processor, implements the steps of the methods in various embodiments of the present application.
[0099] Specific examples in this embodiment can be found in the examples described in the above embodiments and exemplary implementations, and will not be repeated here.
[0100] It is obvious to those skilled in the art that the modules or steps of the present invention described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. They can be implemented using computer-executable program code, and thus can be stored in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those described herein, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.
[0101] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A reverse method of a communication protocol based on a digital grid public document, characterized in that, The method comprises the following steps: determining a target document comprising a communication protocol; inputting each text block included in the target document into a first network model to label each text block to obtain a first label sequence, wherein the first label sequence comprises each text block and label information of the text block; converting each text block into a text vector through a second network model and labeling the text vector to obtain a second label sequence, wherein the second label sequence comprises each text vector and label information of the text vector; optimizing the first label sequence and the second label sequence to obtain a target label sequence; determining a finite state machine of the communication protocol based on the target label sequence to complete reverse engineering of the communication protocol; optimizing the first label sequence and the second label sequence to obtain a target label sequence comprises: marking state text as a transition span, wherein the state text is text included in the first label sequence and the second label sequence, and the state text comprises state information and a verb or a directional preposition; marking event text as an action span, wherein the event text is text included in the first label sequence and the second label sequence, and the event text comprises an action verb; marking first text as an unmarked span, wherein the first text is text included in the first label sequence and the second label sequence, and the first text is not span marked, and the first text comprises state text or event text; marking second text as a target name, wherein the second text is text included in the first label sequence and the second label sequence, and the second text is not span marked; determining a sequence comprising the state text and the transition span, the event text and the action span, the first text and the unmarked span, and the second text and the target name as the target label sequence; the second network model is a neural CRF model based on BERT embedding, and the second network model comprises a BERT encoder, a gated recurrent unit, and a conditional random field layer.
2. The method of claim 1, wherein, converting each text block into a text vector through a second network model and labeling the text vector to obtain a second label sequence comprises: converting the text block into the text vector through a bidirectional encoder representation transformer model included in the second network model; labeling the text vector through a gated recurrent unit and a conditional random field layer included in the second network model to obtain label information of the text vector; determining a sequence formed by the text vector and the label information of the text vector as the second label sequence.
3. The method of claim 2, wherein, Before converting the text block into the text vector through the bidirectional encoder representation transformer model included in the second network model, the method further comprises: obtaining a target grammar for describing the structure of the document of the communication protocol; obtaining a training document comprising the communication protocol; unsupervised learning is performed on the initial bidirectional encoder representation converter model based on the target syntax and the training document, to obtain the bidirectional encoder representation converter model.
4. The method of claim 1, wherein, optimizing the first label sequence and the second label sequence to obtain a target label sequence includes: determining the state text included in the first label sequence and the second label sequence; determining the event text included in the first label sequence and the second label sequence; determining the first text included in the first label sequence and the second label sequence; determining the second text included in the first label sequence and the second label sequence; optimizing the state text, the event text, the first text, and the second text to obtain the target label sequence.
5. The method of claim 1, wherein, determining a finite state machine of the communication protocol based on the target label sequence includes: determining a state set, an input / output event set, and a transition relationship set included in the target label sequence; determining the state set, the input / output event set, and the transition relationship set as the finite state machine.
6. The method of claim 1, wherein, determining a target document including a communication protocol includes: obtaining a first initial document including the communication protocol; performing data cleaning and format standardization operations on the first initial document to obtain a second initial document; preprocessing the second initial document using natural language processing technology to obtain the target document.
7. A communication protocol reverse device based on digital grid public documents, characterized by, includes: a first determining module configured to determine a target document including a communication protocol; a first marking module configured to input each text block included in the target document into a first network model to mark each text block to obtain a first label sequence, wherein the first label sequence includes each text block and label information of the text block; a second marking module configured to convert each text block into a text vector by a second network model and mark the text vector to obtain a second label sequence, wherein the second label sequence includes each text vector and label information of the text vector; an optimization module configured to optimize the first label sequence and the second label sequence to obtain a target label sequence; a second determining module configured to determine a finite state machine of the communication protocol based on the target label sequence to complete reverse engineering of the communication protocol. The optimization module optimizes the first label sequence and the second label sequence to obtain a target label sequence by: marking state text as a transition span, wherein the state text is text included in the first label sequence and the second label sequence, and the state text includes state information and a verb or a directional preposition; marking event text as an action span, wherein the event text is text included in the first label sequence and the second label sequence, and the event text includes an action verb; marking first text as an unmarked span, wherein the first text is text included in the first label sequence and the second label sequence, the first text is not span marked, and the first text includes state text or event text; marking second text as a target name, wherein the second text is text included in the first label sequence and the second label sequence, and the second text is not span marked; and determining a sequence including the state text and the transition span, the event text and the action span, the first text and the unmarked span, and the second text and the target name as the target label sequence. The second network model is a neural CRF model based on BERT embedding, and the second network model includes a BERT encoder, a gated recurrent unit, and a conditional random field layer.
8. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, wherein the computer program is configured to execute the method in any one of claims 1 to 6 when running. 9.An electronic device comprising a memory and a processor, the electronic device characterized by, The memory stores a computer program, and the processor is configured to execute the computer program to execute the method in any one of claims 1 to 6.
10. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method in any one of claims 1 to 6.
Citation Information
Patent Citations
Method and device for identifying unknown entities
CN111222336A
Unknown protocol reverse system based on network traffic
CN111314279A
Method and device for tagging network protocol document and extracting finite-state machine
CN116451684A