Network protocol reverse analysis method based on deep learning and graph neural network
The three-tier framework with Needleman-Wunsch, CRF, and graph neural networks improves protocol parsing efficiency and accuracy by detecting fields, clustering formats, and resolving nested structures, addressing data dependency and semantic limitations in existing methods.
Patent Information
- Application Number
- CN202510815256.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-06-18
AI Technical Summary
The prior art has problems such as strong data dependence, insufficient adaptability and limited semantic analysis in reverse analysis of network protocols, making it difficult to efficiently handle complex protocols.
Using a three-level progressive architecture, combined with the improved Needleman-Wunsch algorithm, knowledge-enhanced CRF and graph neural network, protocol format segmentation and semantic inference are realized through sliding window embedding, bidirectional LSTM encoding and graph attention network.
It significantly improves the degree of automation and accuracy of protocol parsing, can handle complex nested structures, supports the parsing of arbitrary deep nested protocols, and improves the efficiency and consistency of protocol parsing.
Smart Images

Figure CN120321322A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of network protocol reverse parsing, and particularly relates to a network protocol reverse parsing method based on deep learning and graph neural network. Background Art
[0002] With the rapid development of modern communication technology, network protocol reverse engineering (PRE) has become an important tool in network security research, used to reveal the format, semantics, and behavior of unknown protocols, especially in the absence of protocol specifications. Existing research on protocol reverse engineering mainly focuses on the parsing of text protocols and binary protocols. Among them, binary protocols are considered more challenging because of their compact format, lack of obvious field delimiters, and difficult-to-infer semantics.
[0003] Traditional protocol format segmentation methods mainly include alignment-based and probability-based techniques. Alignment-based methods extract field keywords through sequence alignment techniques (such as multiple sequence alignment algorithms), but are vulnerable to noise and have high computational complexity. Probability-based methods use the frequency distribution of fields in messages and locate fields through techniques such as frequent itemset mining and information entropy. However, these methods are inefficient in processing large-scale message calculations and are difficult to extract the complete protocol format.
[0004] Semantic inference is another important task in protocol reverse parsing. Existing methods mainly include type matching and heuristic mining. Type matching infers semantics by identifying common data types of protocol fields (such as integers, floating-point numbers, timestamps), but is limited to common reusable field types. Heuristic mining discovers fields such as length and checksum by calculating the relationships between fields, but requires specific algorithm design and takes a long time.
[0005] In recent years, deep learning models have gradually been applied to protocol reverse parsing. For example, using multi-scale feature extraction and knowledge-driven traffic simulation techniques, binary protocols are segmented and semantically inferred through deep learning models, which show high precision and recall rates in the format segmentation and semantic inference of some unknown protocols. However, this method has a high dependence on high-quality training data and has certain limitations in dealing with complex semantic association fields (such as length, offset, and checksum).
[0006] Generally speaking, the existing technologies still have problems such as strong data dependence, insufficient adaptability, and limited depth of semantic parsing in protocol reverse parsing, and there is an urgent need for more efficient and general technical solutions to meet the diverse protocol reverse requirements. Summary of the Invention
[0007] To address the above deficiencies, the present invention provides a network protocol reverse parsing method based on deep learning and graph neural networks, which proposes a three-level progressive parsing architecture. By means of an improved Needleman-Wunsch algorithm, knowledge-enhanced CRF, and a graph neural network recursive parsing mechanism, the automation degree and accuracy of complex protocol parsing are significantly improved.
[0008] A network protocol reverse parsing method based on deep learning and graph neural networks includes:
[0009] Step 1, basic field detection: Byte-level feature extraction is performed on the binary data stream using sliding window embedding, bidirectional LSTM encoding, and knowledge-enhanced CRF decoding, and a structured field annotation sequence is output.
[0010] Step 2, protocol format clustering based on the sequence: Calculate multi-dimensional similarity through an improved Needleman-Wunsch algorithm, and combine with the dynamic density clustering algorithm optimized by LSH to achieve automatic classification of unknown protocols, and output protocol clusters.
[0011] Step 3, composite structure parsing based on the protocol clusters: Construct a protocol syntax tree based on a graph neural network, perform multi-round message passing through the graph attention network GAT attention mechanism, identify nested structures and recursively parse them, and generate a multi-level protocol syntax tree to represent the protocol structure.
[0012] Preferably, in step 1.1, perform sliding window processing on the input binary data stream, splice the embedding vectors of the 2 bytes before and after each central byte, generate a 640-dimensional combined feature vector, and construct local context; in step 1.2, extract global context features through a bidirectional LSTM network and output a 512-dimensional hidden state sequence; in step 1.3, use a knowledge-enhanced CRF decoder, adjust the label transition weight through a dynamic priority injection mechanism, and output a quadruple sequence of field start offset, length, type, and confidence.
[0013] Preferably, the bidirectional LSTM network adopts a bidirectional encoding mechanism, and the forward LSTM generates a hidden state sequence along the data stream direction , and the backward LSTM processes it in reverse to generate , and a comprehensive feature representation containing global context information is obtained by splicing the hidden states in both directions .
[0014] Preferably, the CRF decoder defines a state transition matrix , where K is the number of label categories; introduce a dynamic priority injection mechanism to improve the calculation of the transition score , where is a predefined domain knowledge weight matrix, is a learnable scaling factor; BIO annotation system is adopted for label prediction, and the conditional probability calculation is expressed as: , wherein, is a label-related weight vector; finally, the optimal label sequence is obtained through Viterbi algorithm decoding, and the structured field description tuple (start offset, field length, type identifier, confidence) is output after merging consecutive similar labels.
[0015] Preferably, step 2 specifically includes: step 2.1, converting the field annotation sequence into a mixed representation of normalized position and type; step 2.2, performing multi-dimensional alignment based on the improved Needleman-Wunsch algorithm, and calculating the protocol similarity by combining the type matching score function and the non-linear length penalty mechanism; step 2.3, adopting the dynamic density clustering algorithm optimized by local sensitive hashing LSH to achieve fine-grained division of the protocol cluster through adaptive adjustment of the neighborhood radius.
[0016] Preferably, the construction of the multi-feature fusion similarity function in step 2.2 specifically includes: step 2.21, using the improved Needleman-Wunsch algorithm to perform global alignment on two annotation sequences , and its dynamic programming recurrence formula is: , where the type matching score function is defined as: , wherein, the gap penalty coefficient .
[0017] Step 2.22, after alignment, calculate the comprehensive similarity through the following formula: , wherein, is the position weight, is the length penalty factor, is the type matching score function, K is the number of matching field pairs obtained after field alignment, and respectively represent the field lengths of the th pair of matching fields in and , and respectively represent the field types of the th pair of matching fields in the sequence and .
[0018] Preferably, in step 2.3, the adaptive neighborhood adjustment mechanism based on Locality-Sensitive Hashing (LSH) includes: the initial neighborhood radius takes the i-th percentile of the pairwise sample similarities in the dataset and dynamically shrinks according to the cluster density during the iteration: , where is the shrinkage rate, is the number of samples in the identified dense cluster, is the total number of samples; through the dynamic adjustment strategy, fine-grained partitioning of the dense protocol is performed while avoiding over-segmentation of the sparse protocol clusters.
[0019] Preferably, step 3 includes: Step 3.1, converting the field annotation sequence output in step 2 into an initial graph structure , where the node set V includes a quadruple: offset, length, type, confidence, and the edge set contains three types of connection relationships: physical connection edges of adjacent fields, logical association edges with a type transition probability exceeding a threshold , and nested edges generated by a predefined nested structure pattern; Step 3.2, performing multi-round message passing to update node features through a Graph Attention Network (GAT). The feature update formula for the i-th node in the -th layer is: , In the formula, represents the feature, the superscript represents the -th layer, the subscript represents the -th node, represents the set of neighbor nodes connected to the -th node, represents the learnable weight matrix of the -th layer of the graph neural network; The attention weight is calculated through a bidirectional attention mechanism: ; Step 3.3, recursively calling the basic detection module to parse nested fields and integrating multi-level syntax trees through a graph fusion operator.
[0020] Preferably, in step 3, the recursive parsing process includes: Step 3.31, when the node type confidence is greater than the set threshold, it is determined as a composite structure node, and the value part byte stream is extracted; Step 3.32, recursively calling the basic detection module to parse and generate a subgraph ; Step 3.33, integrating the sub - graph into the parent - graph structure through a graph fusion operator: , where the projection function converts the sub - graph node offset into local coordinates relative to the parent node; Step 3.34, the recursive process continues until all leaf nodes meet the basic type conditions , generating a protocol syntax tree that supports arbitrarily deep nesting.
[0021] The present invention also discloses a network protocol reverse parsing system based on deep learning and graph neural networks, including:
[0022] A basic field detection module, which uses a sliding - window embedding unit, a bidirectional LSTM encoding unit, and a knowledge - enhanced CRF decoding unit to form a processing pipeline for byte - level feature extraction of the input binary data stream, and outputs a structured field annotation sequence including start offset, length, type, and confidence;
[0023] A protocol format clustering module, which includes a multiple - sequence alignment unit and a dynamic density clustering unit. The multiple - sequence alignment unit calculates the multi - dimensional similarity of the annotation sequences using an improved Needleman - Wunsch algorithm, and the dynamic density clustering unit realizes the automatic classification of protocol clusters through the DBSCAN algorithm optimized by LSH;
[0024] A composite structure parsing module, which includes a protocol syntax tree construction unit, a graph attention network unit, and a recursive parsing unit. The protocol syntax tree construction unit converts the protocol cluster into an initial graph structure, the graph attention network unit updates the node features through multiple rounds of message passing, and the recursive parsing unit performs recursive detection and sub - graph fusion on the nested structure.
[0025] The beneficial effects of the present invention are:
[0026] The protocol parsing system / method disclosed by the present invention, on the basis of traditional protocol parsing technologies, significantly improves the automation degree and parsing accuracy of protocol analysis by introducing deep learning and graph neural network technologies.
[0027] (1) The present invention adopts a three - level progressive architecture, covering three core modules: basic field detection, protocol format clustering, and composite structure parsing, forming an efficient and accurate protocol parsing closed - loop. This architecture can share knowledge among multiple analysis levels and optimize the protocol parsing process through repeated feedback, improving the accuracy and consistency of the analysis results.
[0028] (2) The present invention adopts a protocol format clustering method that combines an improved Needleman-Wunsch algorithm and an adaptive density clustering algorithm. Through the comprehensive comparison of cross-field features, the system can automatically discover the format of unknown protocols and effectively identify version differences, thereby providing prior information for protocol parsing and significantly improving the automation and accuracy of parsing.
[0029] (3) The composite structure parsing process of the present invention realizes the recursive parsing of complex nested structures such as TLV and LV through a graph neural network (GNN) architecture combined with a graph attention network (GAT) mechanism. This innovative solution allows the system to automatically reconstruct the syntax tree of the protocol when facing nested structures and continuously adjust the parsing strategy when the nesting level is uncertain, supporting the parsing of arbitrarily deep nested protocols.
[0030] (4) In order to improve the accuracy of field type prediction, the present invention uses a knowledge-enhanced conditional random field (CRF) for label sequence decoding. By introducing a domain knowledge injection mechanism based on the traditional CRF model, the system can dynamically adjust the field type priority in a complex protocol environment, further improving the accuracy of complex data stream parsing.
[0031] (5) Through the cross-layer feature fusion mechanism of the present invention, the outputs at different layers can be transmitted and optimized with each other, ensuring that the system can extract information layer by layer from the underlying byte features to the high-level protocol structure. At each stage of protocol parsing, the higher-level semantic information extracted in the previous stage can be utilized, thereby maintaining the accuracy and consistency of parsing. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments.
[0033] Figure 1 It is a three-level progressive framework diagram of an embodiment of the present invention;
[0034] Figure 2 It is a schematic diagram of the basic field detection process of an embodiment of the present invention;
[0035] Figure 3 It is a schematic diagram of the protocol format clustering process of an embodiment of the present invention;
[0036] Figure 4 It is a schematic diagram of the composite structure parsing process of an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0037] Embodiments of the present invention provide a method for making the objectives, technical solutions and advantages of the present invention clearer. The following further details the present invention in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. Embodiment 1
[0038] As Figures 1 to 4 shown, this embodiment discloses a network protocol reverse parsing method based on deep learning and graph neural network, including the following steps:
[0039] Step 1, basic field detection: Byte-level feature extraction is performed on the binary data stream using sliding window embedding, bidirectional LSTM encoding, and knowledge-enhanced CRF decoding, and a structured field annotation sequence is output.
[0040] As Figure 2 shown, the basic field detection module proposed by the present invention first performs byte-level segmentation processing on the input binary data stream to generate an ordered byte sequence , where each byte value . By constructing a trainable embedding matrix with dimensions of 256×128 , each byte value is mapped to a 128-dimensional vector representation . To capture the data type feature patterns across bytes, a five-byte sliding window mechanism is used to construct local context: for each central byte , the embedding vectors of its two adjacent bytes before and after are concatenated to form a 640-dimensional combined feature , and zero padding is used at the boundary positions to keep the window size constant.
[0041] In the context feature extraction stage, a bidirectional long short-term memory network (BiLSTM) is used for sequence modeling. The forward LSTM generates a hidden state sequence along the data stream direction , representing the subsequence from the 1st element to the tth element, and the backward LSTM processes it in reverse to generate , representing the subsequence from the tth element to the nth element. By concatenating the bidirectional hidden states, a comprehensive feature representation containing global context information is obtained. By using this bidirectional encoding mechanism, the type features of the previous fields and the structural information of the subsequent fields can be utilized simultaneously. For example, when identifying floating-point data that appears after the timestamp field, the forward LSTM retains the timestamp feature, and the backward LSTM captures the subsequent floating-point pattern.
[0042] In the label sequence decoding stage, the conditional random field (CRF) layer models the legal transition rules of the label sequence through the state transition matrix. Define the state transition matrix (where K is the number of label categories), and a dynamic priority injection mechanism is introduced to improve the calculation of the transition score. Among them, is the improved state transition matrix, that is, the updated state transition matrix; is the index value of the matrix, is the predefined domain knowledge weight matrix, is the learnable scaling coefficient. When the four-byte data satisfies both integer and floating-point characteristics at the same time, this mechanism automatically adjusts the type priority according to the protocol semantics. The label prediction adopts the BIO annotation system, and the conditional probability is calculated as follows. Among them, is the label-related weight vector.
[0043] , In the formula, is the transition score from the label at the previous time step of the specific label sequence to the label at the current time step; is the transition score from the label at the previous time step in all possible label sequences to the label at the current time step; is the observation score of the input feature at the current time step and the label in the specific label sequence at the current time step; is the observation score of the input feature at the current time step and the label in all possible label sequences at the current time step. is the label-related weight vector; finally, the optimal label sequence is decoded through the Viterbi algorithm, and the continuous similar labels are merged and then the structured field description tuple is output. The subscript represents the currently predicted label sequence, corresponding to the actually observed protocol field, represents any label sequence in the candidate label sequence space, corresponding to the set of unobserved potential output states, is the total time step of the sequence, that is, the length of the byte sequence to be annotated, is the current time step, and The optimal label sequence is finally decoded through the Viterbi algorithm, and after merging consecutive similar labels, a structured field description tuple (starting offset, field length, type identifier, confidence) is output, such as (0x0004, 4, "Float32", 0.92). Through the synergistic effect of the context window mechanism, two-stream feature fusion, and knowledge-enhanced CRF, the parsing accuracy of complex protocol data is significantly improved in this step.
[0044] Step 2: Cluster the protocol formats based on the sequence: Calculate the multi-dimensional similarity through the improved Needleman-Wunsch algorithm, and combine it with the dynamically density clustering algorithm optimized by LSH to achieve the automatic classification of unknown protocols and output protocol clusters.
[0045] As Figure 3 shown, first convert the input quadruple sequence (starting offset, field length, type identifier, confidence) into a normalized representation , where the unknown type field (UNK) is regarded as a wildcard in the type dimension but still retains its length feature, forming an intermediate representation that takes into account both structural features and type compatibility. For example, convert the quadruple (0x0004, 4, "Float32", 0.92) into ("Float32", 0.25), where 0.25 represents the normalized proportion of the field starting position in the total protocol length.
[0046] The core of the clustering process lies in defining a multi-feature fusion similarity function between protocol formats. Given two labeled sequences and , use the improved Needleman-Wunsch algorithm for global sequence alignment, and its dynamic programming recurrence formula is: , represents the maximum similarity score when globally aligning the first elements of two protocol labeled sequences and of elements. represents the index of sequence , indicating that the current consideration is the th field in sequence , similarly; : represents the index of sequence , indicating that the current consideration is the th field in sequence , similarly.
[0047] Among them, the type matching score function is defined as: , represents the type identifier of the th field in the sequence. For example, if a field is recognized as "Float32", then is "Float32". represents the type identifier of the th field in the sequence. The gap penalty coefficient .
[0048] This method extends the field to include a case with a wildcard UNK to adapt to the results of the above BiLSTM-CRF, and adapts to the custom type cases of non-basic types. It is more suitable for fuzzy classification in the protocol syntax scenario compared to the traditional NW algorithm. The input is no longer directly the byte sequence of the message, but a quadruple sequence (starting offset, field length, type identifier, confidence) obtained through a deep learning algorithm. The scoring mechanism is reconstructed based on the improved NW algorithm and the structure awareness ability is increased.
[0049] After the alignment is completed, the comprehensive similarity calculation introduces a non-linear length penalty and a position weighting mechanism: ,
[0050] In the formula, is the position weight, is the length penalty factor, is the type matching score function, K is the number of matching field pairs obtained after the field alignment, and respectively represent the th and field lengths of the matching fields in and respectively represent the th and field types of the matching fields in the sequence In this embodiment, the position weight makes the matching of protocol header fields have a higher weight, and the length penalty factor
[0051] The comprehensive similarity calculation formula disclosed in this application realizes difference perception through an exponential decay function: it accurately reflects the impact of tiny differences in field lengths on the protocol structure, is sensitive to changes in the lengths of key fields, and tolerates differences in irrelevant fields. For example, the difference penalty (0.24) between 1 byte and 2 bytes is significantly higher than that between 100 bytes and 101 bytes (0.005), which conforms to the domain characteristics of protocol analysis; and it has excellent anti-interference ability: it balances the offset caused by field detection errors through structure penalties, and can still maintain the robustness of the overall similarity evaluation when the field boundaries are not accurately recognized.
[0052] The clustering algorithm adopts an improved density clustering framework and proposes an adaptive neighborhood adjustment mechanism based on Locality-Sensitive Hashing (LSH). A composite hash key is designed for the mixed features (type, length, position) of protocol fields, and a mathematical relationship is established between the LSH collision probability and the domain radius through a similarity percentile threshold to achieve dynamic adjustment. The initial neighborhood radius takes the 15th percentile of the pairwise sample similarities in the dataset and dynamically shrinks according to the cluster density during the iteration: ,
[0053] where is the new neighborhood radius, which defines the range for examining the "neighbors" around a point. When becomes smaller, the clustering becomes more fine-grained. is the old (current) neighborhood radius, which is dynamically adjusted according to the density of the cluster during the iteration. The initial neighborhood radius is the 15th percentile of the pairwise sample similarities in the dataset. is the shrinkage rate, is the number of samples in the identified dense cluster, is the total number of samples. This dynamic adjustment strategy enables finer-grained partitioning in dense regions (such as the HTTP protocol cluster containing multiple sub-versions), while avoiding over-segmentation of sparse regions (such as independent protocols like SSH).
[0054] Step 3, perform composite structure parsing based on the protocol cluster: construct a protocol syntax tree based on a graph neural network, perform multi-round message passing through the graph attention network GAT attention mechanism to identify nested structures and recursively parse them to generate a multi-level protocol syntax tree; the parsing results are fed back to Steps 1 and 2 to form a cross-layer closed-loop optimization. The syntax tree can represent the structure of the protocol, including the positions, types, lengths, and nested relationships of fields. Generating a syntax tree provides a structured representation of the protocol to support applications such as security detection and vulnerability mining. Step 3 specifically includes:
[0055] Step 3.1, convert the field annotation sequence output in Step 2 into an initial graph structure where the node set V includes a quadruple: offset, length, type, confidence, and the edge set contains three types of connection relationships: physical connection edges of adjacent fields, logical association edges with a type transition probability exceeding a threshold and nested edges generated by predefined nested structure patterns.
[0056] Step 3.2, perform multi-round message passing to update node features through the Graph Attention Network (GAT). The feature update formula for node i at the l-th layer is: , where represents the feature, the superscript represents the l-th layer, and the subscript represents the i-th node, represents the set of neighbor nodes connected to the j-th node, represents the learnable weight matrix of the l-th layer of the graph neural network, , is the feature vector representation of node i at the l-th layer, and after update, it becomes .
[0057] The attention weight is calculated through a bidirectional attention mechanism: , where is the transpose of a learnable attention vector, which performs a dot product with the concatenated result of the transformed node features to calculate the attention score. The weights of this vector are learned during training to determine which feature combinations contribute more to the attention score. represents the feature vector of node i after a linear transformation (completed by the weight matrix is a learnable weight matrix that projects the features of node i into a new feature space for better attention calculation. Here, is the feature vector of node i at the current GAT layer (the l-th layer). represents the feature vector of node j after the same linear transformation. Similar to Similarly, it is the result of a linear transformation of the features of neighboring nodes. When calculating the attention, the relationship between the central node and its neighboring nodes will be considered. Denote the feature vector of node after the same linear transformation The result. This term appears in the summation in the denominator, where represents all neighboring nodes of node The role of the denominator is to normalize the attention scores of all neighboring nodes and the central node to ensure that the sum of the attention weights is 1.
[0058] Step 3.3, recursively call the basic detection module to parse nested fields and integrate multi-level syntax trees through graph fusion operators. The recursive parsing process includes: when the node type confidence is greater than the set threshold, it is determined as a composite structure node, and the byte stream of its value part is extracted ; recursively call the basic detection module for to parse and generate a subgraph ; integrate the subgraph into the parent graph structure through the graph fusion operator: , represents the parent graph (or the current graph) structure in the -th round of iteration in the recursive parsing process. This graph contains the parsed protocol fields and the relationships between them. When a nested structure (such as a TLV field) is detected, its value part needs to be recursively parsed to generate a subgraph. is the existing graph structure into which this subgraph will be integrated. represents the updated graph structure in the -th round of iteration after being processed by the graph fusion operator. This is the result after successfully integrating the newly parsed subgraph (represented by ) into the parent graph , forming a more complete and multi-level protocol syntax tree. represents a graph neural network model with a parameter set of θ. The projection function converts the subgraph node offset to the local coordinates relative to the parent node; the recursive process continues until all leaf nodes meet the basic type conditions , generating a protocol syntax tree that supports arbitrary depth nesting. Embodiment 2
[0059] This embodiment discloses a network protocol reverse parsing system based on deep learning and graph neural networks. It adopts a three-level progressive architecture and consists of three core modules: basic field detection, protocol format clustering, and composite structure parsing, which form a collaborative analysis system. First, the system extracts byte-level features from the input binary data stream through the basic field detection module. It uses a three-stage process of sliding window embedding, bidirectional LSTM encoding, and knowledge-enhanced CRF decoding to output a structured annotation sequence (including start offset, length, type, and confidence). These annotation sequences are then passed to the protocol format clustering module, which calculates multi-dimensional similarity through an improved Needleman-Wunsch algorithm and combines it with a dynamically density clustering algorithm optimized by LSH to achieve automatic classification and version identification of unknown protocols. Finally, the composite structure parsing module constructs a protocol syntax tree based on graph neural networks, identifies nested structures such as TLV / LV through the GAT attention mechanism, and uses a recursive parsing engine to achieve multi-level structure reconstruction, supporting the generation of a complete syntax tree for arbitrarily deeply nested protocols.
[0060] The entire system works through a closed-loop optimization process of "detection → clustering → parsing → re-detection". The quadruple sequence (including the start offset, length, type, and confidence of the field) provided by the basic field detection module is supplied to the clustering and parsing modules. The clustering module uses a protocol template library to provide prior knowledge of the protocol structure for the composite parsing module, while the parsing module triggers a recursive call when a nested structure is detected. The closed-loop process ensures the accuracy and consistency of the entire protocol parsing process.
[0061] The system gradually extracts and fuses information through hierarchical feature abstraction, from the byte level to the field level, then to the protocol level and the composite structure level. The output of each layer provides higher-level semantic information for the next layer, ensuring the parsing accuracy of complex protocols. Through a cross-layer knowledge sharing mechanism, the system can achieve automatic parsing from the original bit stream to the final syntax tree, greatly improving the efficiency and accuracy of protocol parsing.
[0062] The present invention improves the accuracy of network protocol reverse parsing and reduces the false alarm rate. It has important application value in the fields of security detection, vulnerability discovery, and network defense. It can help security researchers more quickly and accurately identify and parse unknown protocols, improving the overall network security level. The innovation and practicality of this method provide new technical means for the network security field and have broad application prospects.
[0063] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements on some of the technical features. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A network protocol reverse parsing method based on deep learning and graph neural networks, characterized in that, Including: Step 1, basic field detection: Use sliding window embedding, bidirectional LSTM encoding, and knowledge-enhanced CRF decoding to extract byte-level features from the binary data stream, and output a structured field annotation sequence; Step 2, protocol format clustering based on the sequence: Calculate multi-dimensional similarity through an improved Needleman-Wunsch algorithm, and combine the dynamic density clustering algorithm optimized by LSH to achieve automatic classification of unknown protocols, and output protocol clusters; Step 3, composite structure parsing based on the protocol clusters: Construct a protocol syntax tree based on a graph neural network, perform multi-round message passing through the graph attention network GAT attention mechanism, identify nested structures and recursively parse them to generate a multi-level protocol syntax tree to represent the protocol structure.
2. The method according to claim 1, characterized in that, Step 1 specifically includes: Step 1.1, perform sliding window processing on the input binary data stream, splice the embedding vectors of the 2 bytes before and after each central byte to generate a 640-dimensional combined feature vector, and construct local context; Step 1.2, extract global context features through a bidirectional LSTM network and output a 512-dimensional hidden state sequence; Step 1.3, use a knowledge-enhanced CRF decoder to adjust the label transition weights through a dynamic priority injection mechanism, and output a quadruple sequence of field start offset, length, type, and confidence.
3. The method according to claim 2, wherein The bidirectional LSTM network adopts a bidirectional encoding mechanism. The forward LSTM generates a hidden state sequence along the data flow direction: , and the backward LSTM processes it in reverse to generate a sequence: , represents the subsequence from the 1st element to the t-th element, represents the subsequence from the t-th element to the n-th element. The comprehensive feature representation containing global context information is obtained by concatenating the bidirectional hidden states .
4. The method according to claim 2, wherein The CRF decoder defines a state transition matrix , where K is the number of label categories; a dynamic priority injection mechanism is introduced to improve the calculation of the transition score , where is the updated state transition matrix, is the predefined domain knowledge weight matrix, is the index value of the matrix, is the learnable scaling coefficient; BIO annotation system is adopted for label prediction, and the conditional probability calculation is expressed as: , where is the transition score from the label at the previous time step of a specific tag sequence to the label at the current time step; is the transition score from the label at the previous time step to the label at the current time step among all possible tag sequences ; is the observation score of the input feature at the current time step and the label of a specific tag sequence at the current time step; is the observation score of the input feature at the current time step and the label of all possible tag sequences at the current time step; is the total number of time steps of the sequence, i.e., the length of the byte sequence to be labeled, and is the current time step.
5. The method according to any one of claims 1-4, characterized in that Step 2 specifically includes: Step 2.1, convert the field annotation sequence into a mixed representation of normalized position and type; Step 2.2, perform multi-dimensional alignment based on the improved Needleman-Wunsch algorithm, and calculate the protocol similarity by combining the type matching score function and the non-linear length penalty mechanism; Step 2.3, use the dynamic density clustering algorithm optimized by local sensitive hashing LSH to achieve fine-grained division of protocol clusters through adaptive adjustment of the neighborhood radius.
6. The method according to claim 5, wherein The construction of the multi-feature fusion similarity function in Step 2.2 specifically includes: Step 2.21, use the improved Needleman-Wunsch algorithm to perform global alignment on the two labeled sequences The dynamic programming recurrence formula is as follows: , Wherein, represents the maximum similarity score when globally aligning the first elements of two protocol annotation sequences and the elements of ; are elements in the sequence , and the type matching score function M is defined as: , wherein represents the type identifier of the -th field in the sequence, represents the type identifier of the -th field in the sequence, and the gap penalty coefficient ; Step 2.22, after alignment is completed, calculate the comprehensive similarity by the following formula: , Wherein, is the position weight, is the length penalty factor, is the type matching score function, K is the number of matching field pairs obtained after field alignment, and respectively represent the th pair of matching fields in and The field lengths in, and respectively represent the th pair of matching fields in the sequence and The field types in.
7. The method according to claim 5, wherein In step 2.3, the adaptive neighborhood adjustment mechanism based on Locality-Sensitive Hashing (LSH) includes: the initial neighborhood radius takes the i-th percentile of the pairwise sample similarities in the dataset and dynamically shrinks according to the cluster density during the iteration: , where is the new neighborhood radius, is the current neighborhood radius, is the shrinkage rate, is the number of samples in the identified dense cluster, is the total number of samples; through the dynamic adjustment strategy, the dense protocol is finely divided while avoiding over-segmentation of the sparse protocol clusters.
8. The method according to claim 1, characterized in that, Step 3 includes: Step 3.1, convert the field annotation sequence output in Step 2 into an initial graph structure: , where the node set V includes quadruples: offset, length, type, confidence, and the edge set includes three types of connection relationships: physical connection edges of adjacent fields, logical association edges with a type transition probability exceeding the threshold and nested edges generated by predefined nested structure patterns; Step 3.2, perform multi-round message passing to update node features through the Graph Attention Network (GAT). The feature update formula for node i at the -th layer is as follows: , where represents the feature, the superscript represents the -th layer, the subscript represents the -th node, represents the set of neighbor nodes connected to the -th node, represents the learnable weight matrix of the -th layer of the graph neural network, is the feature vector representation of node at the -th layer, and after update it becomes ; the attention weight is calculated through a bidirectional attention mechanism: ; where is the transpose of a learnable attention vector; is the result of the feature vector of node after a linear transformation, is the result of the feature vector of node after the same linear transformation , is the result of the feature vector of node after the same linear transformation . Step 3.3, recursively call the basic detection module to parse nested fields, and integrate multi-level syntax trees through a graph fusion operator.
9. The method according to claim 8, wherein In Step 3, the recursive parsing process includes: Step 3.31, when the node type confidence is greater than the set threshold, it is determined as a composite structure node, and the value part byte stream is extracted ; Step 3.32, recursively call the basic detection module to perform parsing to generate a sub-graph ; Step 3.33, integrating the sub-graph into the parent graph structure through the graph fusion operator: , where is the parent graph structure in the -th iteration of the recursive parsing process, is the updated graph structure in the -th iteration of the recursive parsing process after being processed by the graph fusion operator, is a graph neural network model with its parameter set as θ, and the projection function converts the sub-graph node offset into local coordinates relative to the parent node; Step 3.34, the recursive process continues until all leaf nodes meet the base type conditions , generating a protocol syntax tree that supports arbitrarily deep nesting.
10. A network protocol reverse parsing system based on deep learning and graph neural network, characterized in that, Including: A basic field detection module, which uses a sliding window embedding unit, a bidirectional LSTM encoding unit, and a knowledge-enhanced CRF decoding unit to form a processing pipeline for extracting byte-level features from the input binary data stream and outputting a structured field annotation sequence containing start offset, length, type, and confidence; A protocol format clustering module, which includes a multi-sequence alignment unit and a dynamic density clustering unit. The multi-sequence alignment unit calculates the multi-dimensional similarity of the annotation sequence through an improved Needleman-Wunsch algorithm, and the dynamic density clustering unit realizes the automatic classification of protocol clusters through the DBSCAN algorithm optimized by LSH; A composite structure parsing module, which includes a protocol syntax tree construction unit, a graph attention network unit, and a recursive parsing unit. The protocol syntax tree construction unit converts the protocol cluster into an initial graph structure, the graph attention network unit updates the node features through multi-round message passing, and the recursive parsing unit performs recursive detection and subgraph fusion on nested structures.
Citation Information
Patent Citations
Network protocol reverse analysis method based on message field separator identification
CN107707540A
Industrial control protocol reverse analysis method based on semantic pre-mining
CN111585832A
Efficient industrial control protocol analysis method based on deep learning
CN114553983A
Unknown industrial control protocol reverse analysis method and device
CN117354207A
Reverse analysis method for communication protocol of electric power internet of things
CN118612124A
Cited By
Encryption protocol structure restoration method and system based on graph neural network
CN120750678A
Graph neural network-based encrypted protocol structure restoration method and system
CN120750678B
Network data classification method based on artificial intelligence
CN120951035A
Intelligent networked automobile password security detection system
CN121056149A
Smart power grid cloud edge collaborative heterogeneous data security access method
CN121151097A