Protocol Type Adaptive Clustering Method and System for Unknown Message Formats

Through the protocol type adaptive clustering method for unknown message formats, the problem of insufficient packet feature depth extraction and protocol type clustering performance in the prior art is solved, and efficient packet clustering and protocol type recognition are achieved.

CN119766913BActive Publication Date: 2025-06-13NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510253040.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-13
Estimated Expiration
2045-03-05

AI Technical Summary

Technical Problem

The existing protocol clustering methods have shortcomings in the deep extraction of packet feature and the protocol type clustering of unknown types, which limit the improvement of clustering performance.

Method used

A protocol-type adaptive clustering method for unknown packet formats is proposed. By vectoring packets, constructing message feature enhancement modules, and adopting contrast learning adaptive clustering, the optimal cluster number is generated and clustered.

Benefits of technology

The semantic features in the message are effectively extracted, the optimal clustering number is automatically determined, which improves the performance of protocol type clustering, and deals with the problem of poor clustering effect when the type and type of packets are unknown.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119766913B_ABST
    Figure CN119766913B_ABST
Patent Text Reader

Abstract

The present invention discloses a protocol type adaptive clustering method for unknown message formats, including: vectorizing the message byte sequence to obtain a two-dimensional vector corresponding to each segment of the message; constructing a message feature enhancement module, through message vector alignment technology and a hybrid expert self-attention mechanism, unifying message vectors of different dimensions and enhancing the category specialization ability of the model when extracting features; adaptive clustering based on contrast learning, generating an adjacency matrix by calculating the similarity of message vectors and the message format distance, and using the adjacency matrix for Laplacian filtering to generate a contrast view, and obtaining the message vectors for clustering through contrast learning; evaluating the clustering quality according to the message vectors to generate the optimal number of clusters, and clustering the message vectors in combination with the optimal number of clusters to obtain the clustering result. The present invention can effectively extract the semantic features of the message, and further solve the problem of poor message clustering effect when both the types and kinds of messages are unknown.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of software engineering and artificial intelligence, and particularly to a protocol type adaptive clustering method and system for unknown message formats. Background Art

[0002] The cloud-edge-terminal architecture has become the core mode of the power Internet of Things system. Its characteristics are that the terminal device resources are limited, while the edge device and the cloud server undertake the main data processing and computing tasks. This distributed architecture significantly improves the efficiency and scalability of the power Internet of Things. However, the traditional power Internet of Things uses a wide variety of communication protocols for terminal devices, which has caused a major obstacle to the migration and integration of terminal devices into new systems. Therefore, how to enable a new system to effectively identify and parse the messages of various protocols has become a key problem to be solved urgently.

[0003] Traditional protocol analysis methods highly rely on protocol documents or manually designed parsing rules. This not only requires in-depth analysis of protocol details but also faces the challenges of the rapid increase in the number of Internet of Things devices and protocol diversification, making manual parsing both impractical and inefficient. In this context, protocol reverse engineering (PRE) emerged. Its basis lies in the effective clustering of unknown protocol messages, and accurate protocol clustering results can greatly facilitate the subsequent parsing of messages.

[0004] In recent years, many research works on clustering of unknown protocol messages have made great progress. Among them, len-Kmeans uses the maxBIDE algorithm to mine the maximum frequent sequences in the protocol message set, and then calculates the fuzzy membership degree of each protocol message to each maximum frequent sequence according to the fuzzy set theory to construct a fuzzy membership degree vector. Finally, the k-means algorithm is used to divide different types of protocol messages into multiple clusters. NKW uses the PSS algorithm to perform rough clustering on the dataset to be measured based on the continuity of protocol format fields, and extracts the initial clustering centers of the K-means algorithm. Then, the message distance and an improved convergence function are used to perform iterative processing on the data to achieve further clustering of the messages. However, the message features extracted by these protocol clustering methods are relatively coarse-grained and cannot fully capture the semantic information in the messages. At the same time, these methods usually rely on a pre-determined number of clusters, and in real scenarios, this information is not always available.

[0005] To address the above challenges, researchers have proposed a network protocol recognition method based on convolutional neural networks, leveraging the powerful capabilities of neural networks to extract deep features from packets. However, this method faces limitations in obtaining labeled sample data when training the packet feature extraction module, and when performing clustering analysis, the number of types of unknown protocols still needs to be preset in advance. FEAC analyzes the general structure of protocol specifications and then characterizes unknown traffic as protocol specification fusion vectors (PSFVs). Subsequently, representation learning is used to refine the information of PSFVs to compress dimensions and reduce computational complexity. Finally, the refined PSFVs and the DBSCAN algorithm are combined to achieve protocol clustering of unknown traffic. However, the training process of the FEAC method is not end-to-end, resulting in inconsistent optimization between feature extraction and clustering tasks, which limits the improvement of the overall clustering performance.

[0006] In summary, although significant progress has been made in protocol clustering methods in recent years, existing solutions still have deficiencies in the deep extraction of packet features and the clustering of protocol types with unknown numbers of types, which limits the improvement of clustering performance. Therefore, the present invention proposes a protocol type adaptive clustering method and system for unknown packet formats to solve the above problems. Summary of the Invention

[0007] The object of the present invention is to propose a protocol type adaptive clustering method and system for unknown packet formats, which can effectively handle the problems of deep feature extraction and unknown number of types in the protocol clustering task.

[0008] To achieve the above technical objectives, the technical solutions adopted by the present invention are as follows:

[0009] In a first aspect, the present invention discloses a protocol type adaptive clustering method for unknown packet formats, and the method includes the following steps:

[0010] Vectorize the byte sequence of the packet to obtain a two-dimensional vector corresponding to each segment of the packet;

[0011] Construct a packet feature enhancement module, and through packet vector alignment technology and a mixture of experts self-attention mechanism, unify packet vectors of different dimensions and enhance the category specialization ability of the model when extracting features;

[0012] Adaptive clustering based on contrast learning, generate an adjacency matrix by calculating the similarity of packet vectors and the distance of packet formats, and use the adjacency matrix for Laplacian filtering to generate a contrast view, and obtain packet vectors for clustering through contrast learning;

[0013] Evaluate the clustering quality based on the packet vectors to generate the optimal number of clusters, and cluster the packet vectors in combination with the optimal number of clusters to obtain the clustering result.

[0014] Further, the process of vectorizing the message byte sequence to obtain the two-dimensional vector corresponding to each segment of the message includes the following steps:

[0015] For the message , where is the number of bytes of the message, in bytes as the unit, through the embedding function map each byte to a one-dimensional vector :

[0016] ;

[0017] In the formula represents the mapping relationship between the byte and the vector in the randomly initialized embedding matrix of size ; represents the size of the input space, represents the dimension size of the one-dimensional vector embedding;

[0018] Stack the obtained multiple one-dimensional vectors to obtain the two-dimensional vector corresponding to each segment of the message :

[0019] ;

[0020] In the formula represents the stacking operation, and splices multiple one-dimensional vectors into a two-dimensional vector by rows.

[0021] Further, the process of constructing a message feature enhancement module and unifying message vectors of different dimensions through message vector alignment technology and a mixture-of-experts self-attention mechanism includes the following steps:

[0022] Construct a first mixture-of-experts model, which consists of a first gating layer and independent multi-layer perceptron models; among them, after filling the two-dimensional vector corresponding to each segment of the message to a unified size, perform multiple divisions on the filled message vector through a preset multiple segmentation scheme; after each division, discard the part that does not contain valid information, and input the results obtained by each segmentation scheme into the corresponding multi-layer perceptron model for processing respectively:

[0023] ;

[0024] ;

[0025] Among them represents the operation function of the multi-layer perceptron model, represents the filling operation, represents the segmentation operation of the th segmentation scheme, denotes multiple segments processed according to the th segmentation scheme, and all segments in are two-dimensional vectors of dimension ; denotes the hidden layer size, denotes the dimension size of the one-dimensional vector embedding;

[0026] Each padded two-dimensional vector is calculated through the first gating layer, and outputs a probability distribution:

[0027] ;

[0028] wherein, denotes the gating layer, is a trainable weight matrix; Each element in corresponds to the weight of the processing result of a multi-layer perceptron model ; ;

[0029] Based on the scores of the gating layer, the processing results of the multi-layer perceptron model are weighted and combined, and the aligned message vector is calculated as:

[0030] ;

[0031] wherein, is the operation function of the first mixture-of-experts model.

[0032] Furthermore, the process of enhancing the class specialization ability of the model when extracting features includes the following steps:

[0033] For the aligned message vector , respectively pass through three different linear transformation layers , , to generate the query matrix , the key-value matrix and the value matrix :

[0034] ;

[0035] where , and are all of size , and perform matrix multiplication to obtain a feature weight matrix of size ; after activating the feature weight matrix, it is combined with the value matrix ​ Perform matrix multiplication to obtain the attention matrix:

[0036] ;

[0037] where represents the activation function, is a hyperparameter used to measure the attention weight, represents the transpose operation;

[0038] Construct the second mixture-of-experts model, which consists of a second gating layer and feed-forward layers. Each input attention matrix is calculated through the second gating layer, and the feed-forward layer that best matches the current input characteristics is selected. The gating mechanism calculates the activation scores of each feed-forward layer based on the input :

[0039] ;

[0040] where and are the learnable parameters of the gating layer, and the element in the activation score corresponds to the score of the th feed-forward layer;

[0041] Let the set of feed-forward layers be denoted as , and each feed-forward layer specializes the input attention matrix to generate intermediate outputs for specific categories:

[0042] ;

[0043] After the feed-forward layers generate intermediate outputs, the intermediate output set

[0044] is obtained. Then, the activation scores calculated through the second gating layer are used to weight and aggregate the intermediate outputs to generate the final feature representation:

[0045] where is a one-dimensional vector of size .

[0046] Furthermore, for the adaptive clustering based on contrastive learning, the process of generating the adjacency matrix by calculating the similarity of packet vectors and the distance of packet formats, and then generating the contrastive view through Laplacian filtering using the adjacency matrix, and obtaining the packet vectors for clustering through contrastive learning includes the following steps:

[0047] Calculate the similarity matrix of packet pairs through cosine similarity, and the formula is as follows:

[0048] ;

[0049] Among them, represents the inner product operation, represents the modulus operation, represents the th feature representation of the message, represents the th feature representation of the message, , after calculating the cosine similarity between each pair of message feature representations, a message similarity matrix with a shape of is formed, represents the number of messages;

[0050] By the adjustable threshold parameter to control the proportion of the similarity matrix mapped to 1, the message pair similarity matrix is mapped to an adjacency matrix ; specifically, compare the value in the message pair correlation matrix with the adjustable threshold parameter . If the correlation value is greater than or equal to the adjustable threshold parameter , set the element at the corresponding position to 1, indicating the existence of an edge connection; if the correlation value is less than the adjustable threshold parameter , the element at the corresponding position is 0, indicating no edge connection;

[0051] By calculating the minimum operation cost required for conversion between all message pairs, a distance matrix is constructed, and by the adjustable threshold parameter to control the proportion of the distance matrix mapped to 1, the distance matrix is mapped to an adjacency matrix ; specifically, compare the value in the distance matrix with the adjustable threshold parameter . If the distance value is greater than or equal to the adjustable threshold parameter , set the element at the corresponding position to 0, indicating no edge connection; if the distance value is less than the adjustable threshold parameter , the element at the corresponding position is 0, indicating the existence of an edge connection;

[0052] Stack the embedding vectors of all messages to form a feature matrix:

[0053] ;

[0054] In the formula represents the stacking operation;

[0055] Using the adjacency matrix and the adjacency matrix , and their corresponding degree matrices and , respectively, for the feature matrix Perform Laplacian filtering to generate two view feature matrices with different semantic association characteristics and :

[0056] ;

[0057] ;

[0058] where represents the number of filtering times, represents the identity matrix;

[0059] Perform contrastive learning between the two view feature matrices and , and then use linear combination to calculate the final feature matrix of the message:

[0060] .

[0061] Furthermore, the process of generating the optimal number of clusters by clustering quality assessment based on the message vector and clustering the message vector with the optimal number of clusters to obtain the clustering result includes the following steps:

[0062] Represent the current state as a combination of the feature matrix and the cluster state . The feature matrix is the state of the feature matrix at the moment after contrastive learning. The cluster state is the mean of the feature matrix ;

[0063] Output the quality score vector of different numbers of clusters by the quality assessment module according to ; Based on the quality score vector , select the number of clusters at time t through the greedy strategy:

[0064] ;

[0065] where represents selecting the item with the largest median value, represents the random probability, which gradually increases as the training progresses, ;

[0066] Combine the number of clusters , and use k-means clustering for the feature matrix to obtain the clustering result.

[0067] Further, through the quality assessment module, according to output the quality score vectors of different numbers of clusters The process includes the following steps:

[0068] Through the linear transformation layer and respectively perform linear transformations on the feature matrix and the cluster state to extract more representative features; normalize the transformed results and activate them through the activation function and connect the two activation results together:

[0069] ;

[0070] where represents the concatenation operation;

[0071] Adopt the output linear layer and functions to calculate the quality score vector :

[0072] .

[0073] Further, the method further includes:

[0074] Use the cluster state to perform quality assessment to obtain the number of clusters , perform k-means clustering on the feature matrix and calculate the overall clustering optimization loss function to update the parameters in the byte embedding function , the first mixture-of-experts model, the second mixture-of-experts model, and the contrastive learning module; after the parameters are updated, obtain the new feature matrix ; based on the number of clusters perform k-means clustering on the feature matrix to obtain the cluster centers; according to the obtained cluster centers and the feature matrix , calculate the clustering quality evaluation index , and optimize the quality assessment module through the clustering quality evaluation index :

[0075] ;

[0076] where and respectively represent the th and th cluster centers, represents belonging to the the message vectors of the cluster centers; calculate the optimized loss function of quality assessment :

[0077] .

[0078] Furthermore, the calculation process of the overall clustering optimization loss function includes the following steps:

[0079] Let represent the number of preset segmentation schemes, then represents the proportion of valid information of different segmentation schemes. Combining the weights generated by the gating layer , calculate the byte sequence information redundancy loss function :

[0080] ;

[0081] where represents the second norm;

[0082] Calculate the contrastive learning loss function as:

[0083] ;

[0084] where and respectively represent the vector representations of the th message in two different views and , is a hyperparameter for similarity scaling, represents the number of messages, and there are message vector representations in the two views, refers to the th vector representation among the remaining message vector representations except the vector representation ;

[0085] Combining the number of clusters at time , use k-means clustering on the feature matrix to obtain the cluster centers ; the clustering loss function

[0086] ;

[0087] ;

[0088] ;

[0089] where is the vector representation of the th message in the feature matrix , represents calculating the KL divergence, represents the clustering distribution, represents the sharpened clustering distribution;

[0090] For , , perform weighted averaging to obtain the overall clustering optimization loss function . Update the model according to the overall clustering optimization loss function :

[0091] ;

[0092] In the formula, , and are hyperparameters that control the contribution of each term to the loss.

[0093] In a second aspect, the present invention discloses a protocol type adaptive clustering system for unknown message formats, and the system includes a clustering number determination module and a clustering module;

[0094] The clustering module vectorizes the message byte sequence to obtain a two-dimensional vector corresponding to each segment of the message; constructs a message feature enhancement module through message vector alignment technology and a mixture of experts self-attention mechanism; performs adaptive clustering based on contrast learning, obtains message vectors for clustering through contrast learning, and performs k-means clustering on the message vectors according to the optimal clustering number generated by the clustering number module to obtain a clustering result; the clustering number module evaluates the clustering quality of the message vectors generated in the clustering module and generates an optimal clustering according to the quality evaluation result.

[0095] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0096] The protocol type adaptive clustering method and system for unknown message formats of the present invention effectively realizes the in-depth extraction of message features through message sequence vectorization, message feature enhancement, and contrast learning, supports protocol type adaptive clustering through a quality evaluation module, improves the performance of protocol type clustering, and further solves the problem of poor message clustering effect when both the message type and category are unknown. BRIEF DESCRIPTION OF THE DRAWINGS

[0097] Figure 1 is a flowchart of the protocol type adaptive clustering method for unknown message formats;

[0098] Figure 2 It is the principle framework diagram of the protocol type adaptive clustering method for unknown message formats;

[0099] Figure 3 It is the clustering visualization result diagram provided by the embodiment of the present invention. Specific embodiments

[0100] The following further describes the embodiments of the present invention in detail with reference to the accompanying drawings.

[0101] As Figure 1 shown, the present invention discloses a protocol type adaptive clustering method for unknown message formats, including message vectorization, message feature enhancement, and adaptive clustering based on contrast learning. The present invention first vectorizes the message byte sequence to obtain a high-dimensional and dense vector representation to serve the subsequent message feature enhancement module; secondly, through the message vector alignment technology and the mixture-of-experts self-attention mechanism, the message vectors of different dimensions are unified, and the category specialization ability of the model in feature extraction is enhanced; then, by calculating the message vector similarity and the message format distance to generate an adjacency matrix, and using these matrices for Laplacian filtering to generate a contrast view, and obtaining the message vectors for clustering through contrast learning; finally, the clustering quality is evaluated according to the message vectors to generate the optimal number of clusters, and the message vectors are clustered in combination with the optimal number of clusters to obtain the clustering result. The present invention effectively extracts the semantic features contained in the message, can automatically determine the optimal number of clusters, and realizes the performance improvement of the protocol type adaptive clustering task.

[0102] As Figure 2 shown is the structural diagram of the protocol type adaptive clustering system for unknown message formats of the present invention, including message sequence vectorization, message vector low-redundancy alignment module, mixture-of-experts self-attention mechanism, contrast learning module for constructing message similarity view and distance view, and adaptive clustering module.

[0103] S1, Message sequence vectorization.

[0104] Vectorize the message byte sequence to obtain a high-dimensional and dense vector. First, for the message , where is the number of bytes of the message, taking the byte as the unit, and mapping each byte to a one-dimensional vector through the embedding function , indicating the dimension size of the vector embedding:

[0105] ;

[0106] In the formula represents at size In the randomly initialized embedding matrix, the mapping relationship between bytes and vectors represents the size of the input space. During the model training process, the embedding matrix is continuously optimized to learn a better embedding strategy. Then, multiple one-dimensional vectors obtained are stacked to obtain the two-dimensional vector corresponding to each segment of the message. .

[0107] ;

[0108] where represents the stacking operation, that is, multiple one-dimensional vectors are concatenated by rows to form a two-dimensional vector.

[0109] S2, the message feature enhancement module.

[0110] Align the message vectors of different lengths, and then use the mixture-of-experts self-attention module to enhance the message features.

[0111] (1) Through the message vector alignment technology, the message vectors of different dimensions are unified. As Figure 2 shown, the message vector alignment technology includes multi-scheme segmentation and MLP mapping as well as segmentation scheme weight calculation. After padding the two-dimensional message vectors to a unified size, through a preset number of segmentation schemes, the message vectors are divided multiple times according to these schemes. After each division, the parts that do not contain valid information are discarded, and the results obtained by each segmentation scheme are respectively input into independent multi-layer perceptron (MLP) models for processing:

[0112] ;

[0113] ;

[0114] where represents the padding operation, represents the segmentation operation of the th segmentation scheme, represents the multiple segments after processing according to the th segmentation scheme, and all in all are two-dimensional vectors with a dimension of , represents the hidden layer size.

[0115] The segmentation scheme weight calculation is used to fuse the two-dimensional vectors of different segmentation schemes, and the mixture-of-experts model (MoE) is used to calculate the weights for fusion of different processing results. MoE consists of a gating layer and the above-mentioned multiple independent MLPs, and these MLPs are called "experts". Each padded two-dimensional vector will be calculated through the gating layer to output a probability distribution:

[0116] ;

[0117] Among them, represents the gating layer, is a trainable weight matrix, representing the number of experts. Each element in corresponds to the weight of the processing result of an expert. Based on the score of the gating layer, the processing results of the experts are weighted and combined, and the aligned vector is calculated as:

[0118] .

[0119] (2) Through the mixture-of-experts self-attention mechanism, enhance the category specialization ability of the model when extracting features.

[0120] Figure 2 The Self-Attention in for the message vector , respectively passes through three different linear transformation layers , to generate the query matrix , the key-value matrix and the value matrix :

[0121] ;

[0122] where , and are all of size . and perform matrix multiplication to obtain the feature weight matrix of size . After activation of the feature weight matrix, matrix multiplication is performed with the value matrix to obtain the attention matrix:

[0123] ;

[0124] where represents the activation function, is a hyperparameter used to measure the attention weight, represents the transpose operation.

[0125] Figure 2The Mixture of Experts (MoE) model consists of a gating layer and a series of FFNs, which are called "experts". The attention matrix of each input is calculated through the gating layer to select the expert module that best matches the current input characteristics. The gating mechanism calculates the activation scores of each expert based on the input :

[0126] ;

[0127] where and are the learnable parameters of the gating network. The elements in this activation score correspond to the scores of the -th expert. Let the set of experts be denoted as , represents the number of experts. Each expert specializes in the attention matrix of the input to generate intermediate outputs for specific categories:

[0128] ;

[0129] After the feedforward layers generate intermediate outputs, the intermediate output set

[0130] ;

[0131] where is a one-dimensional vector of size .

[0132] S3, Adaptive Clustering Based on Contrastive Learning.

[0133] Construct the similar view and distance view of the message, obtain the message vector for clustering through contrastive learning between views, evaluate the clustering quality according to the message vector to generate the optimal number of clusters, and cluster the message vectors in combination with the optimal number of clusters to obtain the clustering result.

[0134] (1) Contrastive learning. Generate the adjacency matrix by calculating the message vector similarity and message format distance, and use these matrices for Laplacian filtering to generate the contrastive view. Obtain the message vector for clustering through contrastive learning between views.

[0135] After each message is processed by the message feature enhancement module, a one-dimensional vector will be obtained. The one-dimensional vectors of multiple messages form the message feature matrix. Calculate the similarity between message pairs in the message feature matrix through cosine similarity to form the similarity matrix. The formula is as follows:

[0136] ;

[0137] wherein represents the inner product operation, represents the modulus operation, . By an adjustable threshold parameter to control the proportion of the similarity matrix mapped to 1. Compare the values in the correlation matrix with this threshold. If the correlation value is greater than or equal to the threshold, set the element at the corresponding position to 1; if the correlation value is less than the threshold, the element at this position is 0.

[0138] Measure the similarity between bytes by the overlap degree of the possible functions between the bytes of the message. For two bytes and , their function sets are respectively and . By calculating the intersection and union of these two sets, the Jaccard similarity coefficient can be used to represent their similarity:

[0139] ;

[0140] where represents the difference set, represents the intersection, represents the number of elements in the set. On this basis, represent their difference by the Jaccard distance:

[0141] ;

[0142] After obtaining the difference between bytes, the minimum operation cost required to convert one byte sequence to another can be further calculated by the method of dynamic programming. The formula is as follows:

[0143] ;

[0144] where represents the minimum cost of converting the first bytes of the first byte sequence to the first bytes of the second, represents the cost required to insert the th byte in the second byte sequence after the th byte in the first byte sequence, represents the cost required to delete the th byte in the second sequence. By calculating the minimum operation cost required for conversion between all message pairs, a distance matrix is constructed. By an adjustable threshold parameter Control the proportion of the distance matrix mapped to 1. Compare the values in the distance matrix with this threshold. If the distance value is greater than or equal to the threshold, set the element at the corresponding position to 0; if the distance value is less than the threshold, the element at this position is 1.

[0145] Stack the embedding vectors of all packets to form a feature matrix:

[0146] ;

[0147] Use two different adjacency matrices and , as well as their corresponding degree matrices and , respectively perform Laplacian filtering on the feature matrix to generate two contrast views with different semantic association characteristics:

[0148] ;

[0149] ;

[0150] where represents the number of filtering times, represents the identity matrix.

[0151] After the above Laplacian filtering process, we obtain two different view feature matrices and . We calculate the final feature matrix of the packet through the linear combination of the two views. The formula is:

[0152] .

[0153] (2) Adaptive clustering. First, calculate the mean of the feature matrix to obtain the cluster state . The current state is the combination of the feature matrix and the cluster state . Then, through the quality evaluation module, according to output the quality estimates of different numbers of clusters.

[0154] ;

[0155] ;

[0156] where represents the concatenation operation, and are the input linear layers, is the output linear layer, is the activation function.

[0157] In , represents the clustering quality estimated when the number of clusters is . Based on the quality score vector , the number of clusters is selected through a greedy strategy, and its formula is:

[0158] ;

[0159] where represents the item with the largest value in the selected vector, represents the probability. At the beginning of training, a small random probability can promote the model to explore more possible numbers of clusters. As training progresses, will gradually increase, making the model more inclined to select the optimal number of clusters generated by the quality evaluation module.

[0160] Combined with the number of clusters , the feature matrix is used for k-means clustering to obtain the clustering result.

[0161] (3) Joint training. The overall optimization of the model is achieved through joint training, including two parts: clustering optimization and quality evaluation optimization.

[0162] Clustering optimization includes three parts: the byte sequence information redundancy loss function, the contrast learning loss function, and the clustering loss function during message vector alignment. The three loss functions are combined to obtain the joint loss, and the model is updated according to the joint loss.

[0163] After the splitting operation in message vector alignment, the ratio of the effective information in the split segments is calculated to obtain . Combined with the weights generated by the MoE gating layer, the byte sequence information redundancy loss function is as follows:

[0164] ;

[0165] where represents the second norm. The contrast learning loss function is:

[0166] ;

[0167] where and respectively represent the vectors of the th message in two different views and , is a hyperparameter for similarity scaling, indicating the number of packets. There are packet vector representations in the two views, referring to the remaining packet vector representations except for the vector representation and the

[0168] k - means clustering is performed on the feature matrix by combining the number of clusters generated by the quality assessment module at time t. After clustering, the cluster centers can be obtained. The clustering loss function is as follows: ;

[0169] ;

[0170] ;

[0171] ;

[0172] where is the vector representation of the th packet in the feature matrix represents the calculation of the KL divergence, represents the clustering distribution, represents the sharpened clustering distribution. The overall clustering optimization loss function is , , weighted average of

[0173] ;

[0174] In the formula, , and are hyperparameters that control the contribution of each term to the loss.

[0175] The quality assessment module is optimized through a clustering - oriented quality evaluation index :

[0176] ;

[0177] where and represent the th and th cluster centers respectively, represents belonging to the th cluster center and the a packet vector. First, use the state to perform quality assessment to obtain the number of clusters , and perform k-means clustering on the feature matrix . Then calculate and obtain to update the model. After the model is updated, obtain the new feature matrix at time . Then continue to use as the number of clusters, and perform k-means clustering on the feature matrix to obtain the cluster centers. According to these cluster centers and the feature matrix , calculate and obtain the cluster quality evaluation . Finally, the quality assessment optimizes the loss function as follows:

[0178] ;

[0179] Up to this point, the training optimization process of the present invention has been calculated. All experiments were carried out on a server running Windows 11 (64-bit), equipped with an NVIDIA GeForce GTX 2080 Ti graphics processing unit (GPU) and 64GB of memory. Implemented using PyTorch and Python, and during the training process, the Adam optimizer was used. To evaluate the present invention, 6 protocol packets were captured and preprocessed from a real network environment, including DNS, HTTP, Modbus, OICP, s7, and smtp. A part of the packets from each protocol was selected to form a dataset for testing. Compared with six baseline methods, including: UPGMA, K-means, DBSCAN, IDFD, DSIMVC, DML. The present invention uses the Adam optimizer with a learning rate of 0.0001. The vector dimension mapped by bytes is 50, the preset segmentation scheme is [50, 100, 200, 300, 500], the number of rows after packet vector alignment is 400, and the number of experts is 5. The number of Laplace filtering times is 2. and are set to 0.025, is set to 1, and are set to 0.5, is set to 1. The homogeneity, completeness, and v-measure scores are used as evaluation indicators to compare the performance of different methods.

[0180] (1)Performance comparison experiment. The homogeneity, completeness, and v-measure values of the clustering results of the present invention and each baseline on the same dataset are shown in Table 1. The method (UPACM) proposed in the present invention performs best in the three evaluation indicators of homogeneity, completeness, and V-measure, reaching 94.58%, 92.17%, and 93.36% respectively. Further analysis reveals that the methods based on deep learning (IDFD and DML) outperform the traditional machine learning methods (UPGMA, K-means, and DBSCAN), indicating that deep learning has more advantages in feature extraction, can effectively mine the potential information of the packets, and highlights the effectiveness of the deep learning methods. Among the deep learning methods, DSIMVC focuses on filling in the missing view data, so the effect is slightly inferior; while DML and the present method that focus on optimizing the feature representation perform better, highlighting the importance of optimizing the feature representation. Compared with DML, the present method eliminates redundant information during packet vectorization, and the superiority of the results proves that after removing redundant information, the model can capture the core features of the packets more accurately, thereby improving the learning effect.

[0181] (2)Ablation experiment. As shown in Table 2, ablation studies are conducted to analyze the roles of the main modules in the present invention. In the table, w / o-contrastive represents the model with the contrastive learning module removed, w / o-vector alignment means that after packet vectorization, all two-dimensional packet vectors are filled to the same size, and w / o-feature enhancement means not using the self-attention mixture of experts model to enhance features. It can be seen from Table 2 that after losing the contrastive learning module, the overall effect of the model decreases, indicating that the contrastive learning module can capture the internal structure and features of the data, and these features can reflect more comprehensive data patterns. Also, if redundant information removal is not considered and simple filling is used to align the packet vectors, the overall effect of the model will also decrease. Thus, it can be seen that the redundant information contained in the filled vectors affects feature extraction. Finally, the removal of the feature enhancement module also leads to a decrease in the overall effect of the model, which proves that the feature enhancement module can reduce the feature interference between different categories and significantly improve the discriminative ability of feature representation. Generally speaking, compared with UPACM, the other three models are slightly lacking in the ability to model packet features, and their performances all decline.

[0182] (3) Visualization analysis experiment. To further demonstrate the effectiveness of the protocol type adaptive clustering proposed in the present invention and visualize it, the clustering results are mapped to a two-dimensional space through the t-SNE dimensionality reduction algorithm (t-SNE Visualization), and each type of protocol (class0 to class5) is displayed in a different color, as Figure 3 shown. The horizontal and vertical coordinate dimensions 1 and 2 represent the mapped positions of the protocol in the two-dimensional space. As can be seen from Figure 3 , the protocol messages of the same type closely form clusters, and there are obvious boundaries between the clusters, which illustrates the distinguishability between different protocol types, thus fully verifying the effectiveness of the proposed method.

[0183] In summary, the present invention proposes a new protocol type adaptive clustering method for unknown message formats. First, the message byte sequence is vectorized to obtain a high-dimensional and dense vector representation to serve the subsequent message feature enhancement module. Secondly, through the message vector alignment technology and the mixture of experts self-attention mechanism, the message vectors of different dimensions are unified, and the category specialization ability of the model in feature extraction is enhanced. Then, by calculating the message vector similarity and the message format distance, an adjacency matrix is generated, and Laplacian filtering is performed using these matrices to generate a contrast view, and the message vectors for clustering are obtained through contrast learning. Finally, the clustering quality is evaluated based on the message vectors to generate the optimal number of clusters, and the message vectors are clustered according to the optimal number of clusters to obtain the clustering results. Finally, a large number of experiments of the SOTA method are carried out on a data set composed of 6 protocol messages to verify the effectiveness of the present invention.

[0184] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of the present application can be implemented in various computer languages, for example, object-oriented programming languages such as Java and interpreted scripting languages such as JavaScript.

[0185] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions run by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0186] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0187] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are run on the computer or other programmable device to generate a computer-implemented process, and thus the instructions running on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0188] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present application.

[0189] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these changes and modifications.

Claims

1. A protocol type adaptive clustering method for unknown message formats, characterized in that: The method comprises the following steps: Vectorize the message byte sequence to obtain a two-dimensional vector corresponding to each message segment; Construct a message feature enhancement module, unify message vectors of different dimensions through message vector alignment technology and hybrid expert self-attention mechanism, and enhance the model's category specialization ability when extracting features; Adaptive clustering based on contrastive learning generates an adjacency matrix by calculating the message vector similarity and message format distance, and uses the adjacency matrix to perform Laplace filtering to generate a contrast view. The message vector used for clustering is obtained through contrastive learning. The optimal number of clusters is generated based on the clustering quality evaluation of the message vectors, and the message vectors are clustered based on the optimal number of clusters to obtain the clustering results; The process of generating the optimal number of clusters by clustering quality evaluation based on the message vectors and clustering the message vectors based on the optimal number of clusters to obtain the clustering results includes the following steps: The current state G t Represented as feature matrix X t and cluster status C t The combination of features matrix X t is the state of the feature matrix X generated after contrastive learning at time t, and the cluster state C t Then the feature matrix X t The mean of Through the quality assessment module according to G t Output the quality score vector q of different cluster numbers t ; Based on the quality score vector q t , select the number of clusters K′ at time t through a greedy strategy t : where argmax(·) represents the selection of q t The item with the largest median value, θ represents the random probability, and θ gradually increases as the training progresses; Combined with the number of clusters K′ t , use k-means clustering on the feature matrix X to get the clustering results.

2. The protocol type adaptive clustering method for unknown message formats according to claim 1 is characterized in that: The process of vectorizing the message byte sequence to obtain the two-dimensional vector corresponding to each message segment includes the following steps: For message m={b1,b2,…,b l }, where l is the number of bytes in the message, in bytes b j ∈[00,ff] as a unit, by embedding the function f b (·) Map each byte to a one-dimensional vector z j ∈R d : z j =f b (b j ); Where f b (·) indicates that the size is R c×d The mapping relationship between bytes and vectors in the randomly initialized embedding matrix; c represents the size of the input space, and d represents the dimension size of the one-dimensional vector embedding; the obtained multiple one-dimensional vectors are stacked to obtain the two-dimensional vector Z∈R corresponding to each message segment l×d Z=Stack(z1,z2,…,z l ); Here, Stack(·) represents a stacking operation, which concatenates multiple one-dimensional vectors into a two-dimensional vector by row.

3. The protocol type adaptive clustering method for unknown message formats according to claim 1 is characterized in that: The process of building a message feature enhancement module and unifying message vectors of different dimensions through message vector alignment technology and hybrid expert self-attention mechanism includes the following steps: Construct the first hybrid expert model, which consists of the first gating layer and n e The proposed method is composed of independent multi-layer perceptron models; after filling the two-dimensional vector Z corresponding to each message segment to a uniform size, the filled message vector is divided multiple times through multiple preset segmentation schemes; after each division, the part that does not contain valid information is discarded, and the results obtained by each segmentation scheme are respectively input into the corresponding multi-layer perceptron model for processing: Z′=Padding(Z); S i =MLP i (Seg i (WITH')) Among them, MLP i (·) represents the operation function of the multi-layer perceptron model, Padding(·) represents the padding operation, Segment i (·) represents the segmentation operation of the i-th segmentation scheme, S i represents multiple segments processed according to the i-th segmentation scheme, all S i All fragments in are of dimension R h×d , h represents the size of the hidden layer, and d represents the dimension size of the one-dimensional vector embedding; Each padded two-dimensional vector Z′ is calculated through the first gating layer and outputs a Softmax probability distribution: Gate(Z′)=Softmax(Z′ W g ) Among them, Gate(·) represents the gate layer, is a trainable weight matrix; each element in Gate(Z′) corresponds to a multilayer perceptron model e i The weight of the processing result, i∈0,1,…,n e ; Based on the scores of the gating layer, the processing results of the multi-layer perceptron model are weighted and combined, and the aligned message vector is calculated as: Wherein, MoE(·) is the operation function of the first hybrid expert model.

4. The protocol type adaptive clustering method for unknown message formats according to claim 1 is characterized in that: The process of increasing the model’s ability to specialize in categories when extracting features involves the following steps: For the aligned message vector α, three different linear transformation layers W are used respectively. Q ∈R d×d , W K ∈R d×d , W V ∈R d×d To generate the query matrix Q, key matrix K and value matrix V: Q=αW Q ,K=αW K ,V=αW V ; The magnitudes of Q, K and V are all R h×d , Q and K perform matrix multiplication to obtain a matrix of size R h×h The feature weight matrix of ; after Softmax activation of the feature weight matrix, matrix multiplication with the value matrix V is performed to obtain the attention matrix: Where Softmax(·) represents the activation function, δ is a hyperparameter used to measure the attention weight, and T represents the transposition operation; Construct a second hybrid expert model, which consists of the second gating layer and n e The attention matrix of each input is calculated through the second gating layer, and the feedforward layer that best matches the current input characteristics is selected. The gating mechanism calculates the activation score of each feedforward layer based on the input. P=(p1,p2,p3,…): P=Softmax(W P Attention(Q,K,V)+b P ) Where W P and b P is the learnable parameter of the gating layer, and the element p in the activation score i Corresponding to the score of the i-th feed-forward layer; Let the set of feed-forward layers be represented as Each feed-forward layer specializes the input attention matrix to generate intermediate outputs for a specific category: o i =e i (Attention(Q,K,V)); n e The feed-forward layer generates n e After the intermediate outputs, we get the intermediate output set The activation scores calculated by the second gating layer are then used to perform weighted aggregation on the intermediate outputs to generate the final feature representation: in The size is R h A one-dimensional vector of .

5. The protocol type adaptive clustering method for unknown message formats according to claim 1 is characterized in that: Adaptive clustering based on contrastive learning generates an adjacency matrix by calculating the message vector similarity and message format distance, and uses the adjacency matrix to perform Laplace filtering to generate a contrast view. The process of obtaining a message vector for clustering through contrastive learning includes the following steps: The message pair similarity matrix is ​​calculated by cosine similarity. The formula is as follows: The · represents the inner product operation, and ||·|| represents the modular operation. Indicates The characteristics of a message are represented by Indicates The characteristics of a message are represented by After calculating the cosine similarity between each pair of message feature representations, a shape of R is formed. N×N The message similarity matrix, N represents the number of messages; The proportion of similarity matrix mapping to 1 is controlled by the adjustable threshold parameter λ1, and the message pair similarity matrix is ​​mapped to the adjacency matrix A; specifically, the value in the message pair correlation matrix is ​​compared with the adjustable threshold parameter λ1, and if the correlation value is greater than or equal to the adjustable threshold parameter λ1, the element at the corresponding position is set to 1, indicating that there is an edge connection; if the correlation value is less than the adjustable threshold parameter λ1, the element at the corresponding position is 0, indicating that there is no edge connection; By calculating the minimum operation cost required for conversion between all message pairs, a distance matrix is ​​constructed, and the proportion of the distance matrix mapping to 1 is controlled by the adjustable threshold parameter λ2, and the distance matrix is ​​mapped to the adjacency matrix B; specifically, the value in the distance matrix is ​​compared with the adjustable threshold parameter λ2, and if the distance value is greater than or equal to the adjustable threshold parameter λ2, the element at the corresponding position is set to 0, indicating no edge connection; if the distance value is less than the adjustable threshold parameter λ2, the element at the corresponding position is 0, indicating an edge connection; The embedding vectors of all messages are stacked to form a feature matrix: Where Stack(·) represents the stacking operation; Use adjacency matrix A and adjacency matrix B, and their corresponding degree matrix D A and D B , respectively, for the feature matrix Perform Laplace filtering to generate two view feature matrices with different semantic association characteristics and Where n represents the number of filtering times, and I represents the identity matrix; In both view feature matrices and Perform comparative learning between them, and then use linear combination to calculate the final feature matrix of the message:

6. The protocol type adaptive clustering method for unknown message formats according to claim 1, characterized in that: Through the quality assessment module according to G t Output the quality score vector q of different cluster numbers t The process includes the following steps: Through the linear transformation layer Lin X and Lin C For the feature matrix X t and cluster status C t Perform a linear transformation to extract more representative features; normalize the transformed result, activate it through the activation function σ(·), and connect the two activation results together: With t =Concat(σ(Norm(Lin X (X t ),σ(Norm(Lin C (C t )); Where Concat(·) represents the concatenation operation; Use the output linear layer Lin out And Softmax function calculates the quality score vector q t : q t =Softmax(Lin out (Con t ))。 7. The protocol type adaptive clustering method for unknown message formats according to claim 6 is characterized in that: The method further comprises: Use cluster status G t Perform quality assessment to obtain the number of clusters K′ t , for the feature matrix X t Use k-means clustering to calculate the overall clustering optimization loss function To update the byte embedding function f b (·), the parameters in the first hybrid expert model, the second hybrid expert model, and the contrastive learning module; after the parameters are updated, a new feature matrix X is obtained t+1 ; Based on the number of clusters K′ t For the feature matrix X t+1 Use k-means clustering to get the cluster center; according to the obtained cluster center and feature matrix X t+1 , calculate the clustering quality evaluation index Y G , through the clustering quality evaluation index Y G To optimize the quality assessment module: Center i″ and Center k″ Respectively represent the i″th and k″th cluster centers, x i″j″ Represents the j'th message vector belonging to the i'th cluster center; the quality assessment optimization loss function is calculated 8. The protocol type adaptive clustering method for unknown message formats according to claim 7, characterized in that: The overall clustering optimization loss function is The calculation process includes the following steps: Let Div represent the number of preset partitioning schemes, then U = (u1,u2,u3,,u Div ) represents the effective information ratio of different segmentation schemes, and combined with the weight Gate(Z′) generated by the gating layer, the byte sequence information redundancy loss function is calculated Here, ||·||2 represents the two-norm; Calculate the contrastive learning loss function for: where x′ i and x′ i′ Respectively represent the i′th message in two different views and The vector representation in , τ is a hyperparameter for similarity scaling, N represents the number of messages, there are 2N message vector representations in the two views, x j′ Refers to the division vector x i′ The j′th vector representation among the remaining 2N-1 message vector representations; Combined with the number of clusters K′ at time t t , use k-means clustering on the feature matrix X to get the cluster center Clustering loss function for: where x i′ is the vector representation of the i′th message in the feature matrix X, KL(·) represents the calculation of KL divergence, H i′j′ represents the cluster distribution, H′ i′j′ represents a sharp clustered distribution; Perform weighted averaging to obtain the overall clustering optimization loss function Optimize the loss function based on the overall clustering To update the model: Where λ, μ, and η are hyperparameters that control the contribution of each term in the loss.

9. A protocol type adaptive clustering system for unknown message formats based on the method according to any one of claims 1 to 8, characterized in that: The system includes a module for determining the number of clusters and a clustering module; The clustering module vectorizes the message byte sequence to obtain a two-dimensional vector corresponding to each message segment; a message feature enhancement module is constructed through message vector alignment technology and hybrid expert self-attention mechanism; Adaptive clustering based on contrastive learning obtains message vectors for clustering through contrastive learning, and performs k-means clustering on the message vectors to obtain clustering results according to the optimal number of clusters generated by the cluster number module; the cluster number module performs clustering quality assessment on the message vectors generated in the clustering module, and generates the optimal cluster according to the quality assessment result.

Citation Information

Patent Citations

  • Message sequence clustering method of unknown binary private protocol

    CN109951464A

  • Unknown network protocol classification method based on transfer learning

    CN117874599A