Unknown protocol identification method and system based on semi-supervised clustering

By adopting a semi-supervised clustering method in unknown protocol recognition, using the statistical features and fingerprint features of traffic samples, and combining constraint information to improve the K-means algorithm, the problem of scarcity and insufficient adaptability in unknown protocol recognition is solved, and the recognition effect of high accuracy and efficiency is achieved.

CN119946150AActive Publication Date: 2025-05-06SOUTHEAST UNIV +1

Patent Information

Application Number
CN202510037066.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-05-06
Estimated Expiration
2045-01-09

AI Technical Summary

Technical Problem

It is difficult for the prior art to effectively identify unknown protocols, especially in the absence of prior knowledge and scarcity of labeled samples, traditional methods such as port-based identification and deep packet retrieval have problems with inefficient efficiency and inadequate adaptability.

Method used

Using a semi-supervised clustering method, by extracting the statistical features and fingerprint features of the traffic samples, constrained information is constructed to guide the clustering process, clustering is performed using the improved K-means algorithm, and the initial clustering center and allocating samples are selected in combination with the constraint information to form more accurate and robust clustering results.

Benefits of technology

It effectively alleviates the problem of scarcity of data labeling, improves the accuracy and efficiency of unknown protocol identification, and can achieve high accuracy and purity recognition under a small number of labeled samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119946150A_ABST
    Figure CN119946150A_ABST
Patent Text Reader

Abstract

The invention discloses an unknown protocol identification method and system based on semi-supervised clustering. The method comprises the following steps: collecting marked and unmarked flow data of different protocols and extracting a flow statistical feature vector and a fingerprint feature vector; constructing constraint information by using marked and unmarked data according to flow correlation to obtain a necessary connection constraint set, a non-connection constraint set and an equivalence class set; calculating a Laplacian score of each flow statistical feature by using the necessarily connected constraint set and the not connected constraint set to perform feature selection, and fusing the flow statistical features after feature selection with the fingerprint features to obtain a single-flow feature vector; taking an equivalence class set and do not link constraint information as guidance, and mixing marked and unmarked single streams for semi-supervised clustering; and constructing a classifier by using the clustered traffic clusters. According to the method, the multi-dimensional characteristics of the network traffic can be represented more accurately through feature selection and feature fusion. Under the condition of data scarcity, the potential information in the unmarked data is mined by constructing the constraint information, so that the identification effect on the unknown protocol traffic is improved, and the dependence on the marked data is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of network protocol identification, and specifically to a method and system for identifying unknown protocols based on semi-supervised clustering, which is suitable for scenarios such as network traffic analysis, network security protection, and dynamic identification of unknown protocols. Background Art

[0002] In scenarios such as industrial control, the Internet of Things, and military communications, many network applications use customized private unknown protocols to meet specific functional, transmission performance, and privacy protection requirements. However, due to the lack of prior knowledge, the network traffic generated by them is difficult to monitor through conventional methods, and these protocols are often not fully tested for security and may have design flaws, making them easy targets for hacker attacks and posing a potential threat to cyberspace.

[0003] Unknown protocol identification aims to determine the category of known protocol traffic without prior knowledge or with only a small amount of prior knowledge, and to mine private unknown protocol traffic. This is of great value for applications such as network intrusion detection, malware detection, user behavior analysis, and protocol reverse engineering. However, traditional port-based identification methods and deep packet retrieval techniques are no longer sufficient to meet this challenge. Port-based identification methods mainly rely on detecting the traffic transport layer port to identify the protocol type of the traffic, and the identification speed is very fast. However, many modern protocols use dynamic port allocation strategies, and protocol traffic may be transmitted on any port. The port information cannot accurately reflect the protocol type, and malicious traffic and some encrypted protocols will disguise themselves as common ports (such as 80, 443, etc.) to transmit data, evading security detection and making port-based identification invalid. Deep packet inspection requires parsing the payload content of each data packet. This method has a large computational overhead and is difficult to process massive data in real time, especially in a dynamic network environment. As new protocols continue to emerge, deep packet retrieval technology usually relies on rule bases or signature bases for parsing. However, the lag in updating the rule base and the lack of ability to parse unseen protocol content limit the adaptability of deep packet retrieval technology to unknown protocols. Although machine learning and deep learning technologies have achieved remarkable results in identifying known protocol traffic, they often rely on the need for a large amount of labeled data, which is expensive to obtain. In view of this, the problem of unknown protocol identification has received widespread attention. The current mainstream methods are divided into unsupervised methods based on direct clustering, semi-supervised methods that combine known protocol information, and methods based on self-supervised learning.

[0004] Unsupervised methods cluster samples based on similarities, without labeling information or prior knowledge, but lack label guidance. The results require subsequent analysis to explain, which increases the complexity of practical applications. Some studies combine self-supervised learning and supervised and unsupervised hybrid methods for recognition. The features generated by self-supervised learning already contain a lot of contextual information in the pre-training stage, which can improve the feature separation effect of unsupervised clustering and supervised classifiers and promote the recognition accuracy of protocol types. However, its effect depends on whether the pre-training task design is reasonable. If the constructed task cannot capture the key features between protocols, it will directly affect the effect of subsequent steps. Semi-supervised methods guide clustering through a small number of labeled samples or constraint information. However, existing studies have difficulty in constructing the most representative features, insufficient use of the prior knowledge contained in the labeled samples, and cannot achieve good results when the number of clustering categories is set small or the parameters are not selected appropriately. Summary of the invention

[0005] The present invention proposes a method and system for identifying unknown protocols based on semi-supervised clustering, the main purpose of which is to solve the problems of difficulty in identifying unknown protocols, scarcity of labeled samples, and insufficient utilization of prior knowledge.

[0006] The present invention discloses a method for identifying unknown protocols based on semi-supervised clustering, comprising the following steps:

[0007] Step 1: Extract data features of traffic samples of labeled and unlabeled protocols;

[0008] Step 2: Construct constraint information based on data characteristics;

[0009] Step 3: Use constraint information to filter and fuse the data features of each traffic sample to obtain new data features;

[0010] Step 4: Use the K-means clustering algorithm to perform semi-supervised clustering using the new data features as input. At the same time, use constraint information to improve the K-means clustering algorithm during the clustering process to obtain clustered traffic clusters.

[0011] Step 5: Use the clustered traffic clusters as training data to construct a classifier.

[0012] Preferably, the step 1 specifically includes:

[0013] Collect labeled and unlabeled traffic samples of different protocols;

[0014] The data packets with the same source IP address, destination IP address, source port, destination port, and transport layer protocol in the traffic sample are divided into the same flow;

[0015] Extract data features of each stream.

[0016] Preferably, the data features of each stream include a statistical feature vector and a fingerprint feature vector;

[0017] Preferably, the statistical features include at least basic flow information (source IP, destination IP, protocol type, flow duration, source port, destination port), data packet size (maximum, minimum, average, standard deviation of a single flow and the size of upstream and downstream data packets therein), number of data packets (number of data packets in a single flow, number of upstream and downstream data packets in a single flow and their ratio), data packet time interval (maximum, minimum, average, standard deviation of a single flow and the time interval between arrival of upstream and downstream data packets therein), rate (flow rate of a single flow, rate of upstream and downstream data packets in a single flow and their ratio, in bytes / second), data packet rate (flow rate of a single flow, rate of upstream and downstream data packets in a single flow and their ratio, in number / second), bytes (number of bytes in a single flow, number of upstream and downstream bytes in a single flow and their ratio), data packet fragmentation (maximum, minimum, average, standard deviation of the number of upstream and downstream fragments), and these statistical features are combined into a statistical feature vector.

[0018] Preferably, the fingerprint feature is a set of frequent items that can represent a specific protocol communication mode, including the location information and content information of the frequent items. These frequent items reflect the commonality and regularity of the data traffic generated under the same protocol during the communication process, such as the recurrence of specific application layer protocol field values, which usually contain some semantic information, such as type field, Flag field, protocol identifier, version number and other fields.

[0019] Preferably, the constraint information includes a must-connect constraint set, a must-connect constraint set, and an equivalence class set.

[0020] Preferably, the step 2 specifically includes:

[0021] For any two marked flows with the same label or any two related flows in the traffic sample, add them as a must-connect constraint to the must-connect constraint set;

[0022] For two marked flows with different labels, add them as a no-connect constraint to the no-connect constraint set;

[0023] The flows with indirect association in the must-connect constraint set are classified into the same equivalence class to obtain an equivalence class set.

[0024] Preferably, the correlation is: for two flows, if their destination IP, destination port and transport layer protocol are the same and the time interval between them is within a certain range, then the two flows are correlated.

[0025] Preferably, the step three specifically includes:

[0026] Using the traffic sample statistical feature vector and the must-connect constraint set, the k-nearest neighbor algorithm is used to construct a neighbor graph and calculate the weight matrix of the neighbor graph.

[0027] Using the traffic sample statistical feature vector and the unconnected constraint set, a negative constraint graph is constructed and the weight matrix of the negative constraint graph is calculated;

[0028] The Laplacian matrices of the neighbor graph and the negative constraint graph are calculated respectively through the weight matrices of the neighbor graph and the negative constraint graph;

[0029] According to the Laplace matrix and the statistical feature vector of the traffic sample, the Laplace score of each feature is calculated and the traffic statistical features below the threshold are removed to obtain the statistical feature vector after feature selection;

[0030] The statistical feature vector after feature selection is fused with the fingerprint feature vector to obtain new data features for each traffic sample.

[0031] Preferably, the fusion of the statistical feature vector after feature selection and the fingerprint feature vector is performed by using a serial fusion and splicing method.

[0032] Preferably, the traffic sample statistical feature vector and the must-connect constraint set are used to construct a neighbor graph using a k-nearest neighbor algorithm and calculate a weight matrix of the neighbor graph, specifically including:

[0033] For any two samples i and j, if the sample pair (i, j) belongs to the must-connect constraint set or i is one of the k neighbors with the smallest distance to j, then i and j have a common edge in the neighbor graph;

[0034] For samples i and j that have common edges in the neighbor graph, calculate their value w in the weight matrix ij The calculation formula is:

[0035]

[0036] where x i and x j are the statistical feature vectors of sample i and sample j respectively, ||x i -x j || 2 is the square of the Euclidean distance between sample i and sample j, indicating their similarity. λ is a constant parameter used to control the degree of attenuation of similarity. Usually, λ is a hyperparameter selected through experiments or cross-validation;

[0037] For samples i and j that have no common edges in the neighbor graph, their values ​​in the weight matrix are 0.

[0038] Preferably, the negative constraint graph is constructed using the traffic sample statistical feature vector and the unconnected constraint set, and the weight matrix of the negative constraint graph is calculated, specifically including:

[0039] For any two samples i and j, if the sample pair (i, j) belongs to the unconnected constraint set, then i and j have a common edge in the negative constraint graph;

[0040] For samples i and j that have common edges in the negative constraint graph, their values ​​in the weight matrix are 1;

[0041] For samples i and j that have no common edges in the negative constraint graph, their values ​​in the weight matrix are 0.

[0042] Preferably, the step 4 specifically includes:

[0043] Randomly select a flow from each equivalence class as the initial cluster center;

[0044] The next cluster center is selected according to the weighted probability distribution of the square of the distance from each flow in the traffic sample to the selected cluster center, and this step is repeated until a predetermined number of initial cluster centers are selected;

[0045] Assign each flow in the traffic sample to its own cluster center;

[0046] Update the cluster centers, and then continue to assign each flow in the traffic sample to its respective cluster center until convergence or the number of iterations reaches a predetermined amount.

[0047] Preferably, the next cluster center is selected by weighted probability distribution according to the square of the distance from each data point to the selected cluster center, and the probability of each stream being selected as the next cluster center is:

[0048]

[0049] Among them, v i is the new data feature of traffic sample i, D(v i ) is the Euclidean distance from traffic sample i to its nearest cluster center, and X is the new data feature set of all traffic samples.

[0050] Preferably, each flow in the traffic sample is assigned to its own cluster center, specifically including:

[0051] For the first single flow, the traffic cluster with the closest cluster center is selected as the traffic cluster to which the first single flow will be assigned;

[0052] Determine whether the first single flow and all second single flows in the traffic cluster to be allocated violate the no-connection constraint. If so, find the next nearest cluster center until the no-connection constraint is not violated.

[0053] When the no-connection constraint is not violated, the first single flow is assigned to the corresponding traffic cluster, and all third single flows in the equivalence class to which the first single flow belongs are directly classified into the same traffic cluster as the first single flow.

[0054] The present invention also provides a system corresponding to the unknown protocol identification method based on semi-supervised clustering, the system comprising a traffic preprocessing module, a constraint information construction module, a multi-dimensional feature construction module, a semi-supervised clustering module, and a classifier construction module.

[0055] The traffic preprocessing module is used to perform flow splitting processing on the collected marked and unmarked traffic samples, and extract statistical features and fingerprint features of each flow respectively;

[0056] The constraint information building module is used to build a must-connect constraint set, a do-not-connect constraint set, and an equivalence class set according to the label of the marked single flow and the correlation between two different single flows;

[0057] The multi-dimensional feature construction module is used to calculate the constrained Laplace score of each statistical feature according to the constraint information, retain the statistical features with scores higher than the threshold and merge them with the fingerprint features to obtain the feature vector of each traffic sample;

[0058] The semi-supervised clustering module uses constraint information to improve the K-means algorithm clustering process and clusters traffic samples to obtain multiple traffic clusters;

[0059] The classifier construction module is used to construct an NCC classifier according to the traffic cluster.

[0060] The present invention also provides a computer device, characterized in that the device includes: a storage medium storing a computer program for executing unknown protocol identification; a processor for executing the computer program; and the computer program enables the processor to implement the steps of any of the methods described above.

[0061] Preferably, the storage medium includes a hard disk, a flash disk or other non-volatile storage devices.

[0062] Beneficial effects:

[0063] The present invention adopts a semi-supervised clustering method, combines known protocol samples and traffic correlation to build constraint information, and can make full use of unlabeled samples for training on the basis of a small number of labeled samples, thereby effectively alleviating the problem of scarce data annotation. Compared with the prior art, the semi-supervised learning method can use constraint information to enhance the clustering process, especially when facing unknown protocol traffic, making the clustering results more accurate and robust.

[0064] Statistical features can capture the basic data patterns of traffic, while fingerprint features can identify protocol-specific behavior patterns. The present invention extracts traffic statistical features and protocol fingerprint features and combines a small number of known protocol samples with flow correlation to construct constraint information, and uses Laplace scores to measure the "importance" of each feature to retain those features with higher discriminative power, which can more comprehensively characterize the behavior of network traffic. Moreover, this method takes into account the interrelationships and constraint information between features, and can refine and enhance the characterization capabilities of features. It is more sophisticated than the method of removing redundant or irrelevant features by simple metrics (such as information gain, variance, etc.), thereby improving the accuracy of clustering.

[0065] In terms of the initial cluster center selection strategy, if the initial cluster center is not properly selected, the K-means algorithm may fall into a local optimal solution and fail to find a global optimal solution. The present invention uses the constructed equivalence class set to change the previous K-means algorithm's random selection of the initial cluster center. It randomly selects samples from each equivalence class with labeled samples as the initial cluster center, and then selects the next cluster center according to the weighted probability distribution of the square of the distance from each data point to the selected cluster center. This improves the representativeness of the initial cluster center, allowing the clustering algorithm to find a suitable cluster center in fewer iterations, while reducing the algorithm's dependence on the initial center selection and enhancing the robustness of the model. In the sample allocation process, a distance and constraint-based allocation mechanism is adopted to ensure that each traffic sample will not violate the constraint conditions when allocated, avoiding incorrect cluster allocation. At the same time, the division of equivalence classes is used to ensure consistent clustering of similar traffic samples, effectively reducing the possibility of misclassification. This constraint-based sample allocation method not only enhances the relevance of clustering results, but also better copes with complex network traffic, especially the identification of unknown protocols and variant traffic, and has stronger adaptability and robustness.

[0066] In summary, the present invention can effectively utilize the prior information in the labeled samples, reduce the dependence on the labeled samples and improve the accuracy and efficiency of unknown protocol recognition. This scheme can achieve high accuracy and purity recognition of unknown protocols under the conditions of a small number of labeled samples and a low number of clusters. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Figure 1 This is a flow chart of the method described in the present invention.

[0068] Figure 2 This is the framework diagram of the system corresponding to the unknown protocol identification method based on semi-supervised clustering.

[0069] Figure 3 is a structural schematic diagram of a computer device of the present invention,

[0070] Figure 4 Schematic diagram for comparing the clustering method of the present invention with other clustering methods,

[0071] Figure 5 Schematic diagram of clustering purity comparison between the clustering method of the present invention and other semi-supervised clustering methods,

[0072] Figure 6 This is a schematic diagram showing the comparison of the accuracy of the classifier of the present invention with other semi-supervised classifiers.

[0073] Figure 7 Schematic diagram of the effect of the number of labeled streams on the classifier accuracy.

[0074] Specific implementation method

[0075] The following embodiments of the present invention are described in conjunction with the accompanying drawings. The following embodiments are used to illustrate the present invention, but are not intended to limit the scope of the present invention.

[0076] Embodiment: A method for identifying unknown protocols based on semi-supervised clustering, Figure 1 : is a flow chart of the unknown protocol identification method based on semi-supervised clustering of the present invention, the method comprising the steps of:

[0077] S1: Extract data features of traffic samples of labeled and unlabeled protocols;

[0078] The step S1 specifically includes the following steps:

[0079] S11: Collect marked and unmarked traffic samples of different protocols. This step uses Wireshark traffic collection software to collect marked and unmarked traffic data of different protocols. Since the duration of the data stream generated by most unknown protocols usually depends on the duration of user behavior, and the statistical characteristics of the streams in the same time period are very similar, a timeout is set for all sample streams, that is, for a data stream, the messages whose duration exceeds the timeout are removed. In this embodiment, the timeout is 120s.

[0080] S11: Divide the data packets with the same source IP address, destination IP address, source port, destination port, and transport layer protocol in the traffic sample into the same flow. This step uses the Splitcap tool to divide the collected data according to the source IP address, destination IP address, source port, destination port, and transport layer protocol.

[0081] S12: Extract the data features of each stream. The data features of each stream include statistical feature vectors and fingerprint feature vectors. At least include basic stream information (source IP, destination IP, protocol type, stream duration, source port, destination port), data packet size (maximum, minimum, average, standard deviation of single stream and the size of uplink and downlink data packets therein), number of data packets (number of data packets of single stream, number of uplink and downlink data packets in single stream and their ratio), data packet time interval (maximum, minimum, average, standard deviation of single stream and the time interval between uplink and downlink data packets arrival therein), rate (stream rate of single stream, rate of uplink and downlink data packets in single stream and their ratio, in bytes / second), data packet rate (stream rate of single stream, rate of uplink and downlink data packets in single stream and their ratio, in number / second), bytes (number of bytes of single stream, number of uplink and downlink bytes in single stream and their ratio), data packet fragmentation (maximum, minimum, average, standard deviation of the number of uplink and downlink fragmentation), and these statistical features are combined into a statistical feature vector. Table 1 below is a flow statistics feature table of each single flow extracted in this embodiment.

[0082] Table 1 Flow statistics characteristics table

[0083]

[0084] The fingerprint feature is a set of frequent items that can represent a specific protocol communication mode, including the location information and content information of the frequent items, which are combined into a fingerprint feature vector. These frequent items reflect the commonality and regularity of the data traffic generated under the same protocol during the communication process, such as the repeated appearance of a specific application layer protocol field value. Usually, these fields contain some semantic information, such as type field, Flag field, protocol identifier, version number and other fields. This step extracts the frequent items with a length of 1 byte in the first 32 bytes of the application layer load, and records the position of the frequent items. Subsequently, the frequent items are sorted according to the number of times they appear simultaneously in the data stream, and representative frequent items and their positions are selected to construct the final fingerprint feature vector.

[0085] S2: Construct constraint information based on data characteristics;

[0086] The step S2 specifically includes the following steps:

[0087] S21: For any two marked flows with the same label or any two related flows in the traffic sample, add them as a must-connect constraint to the must-connect constraint set. The two related single flows are two single flows, if the destination IP, destination port and transport layer protocol are the same and the time between them is within a predetermined range, then the two flows are related. In this embodiment, the predetermined range is one minute.

[0088] S22: For two marked flows with different labels, add them as a no-connection constraint to the no-connection constraint set;

[0089] S23: Classify the flows with indirect association in the must-connect constraint set into the same equivalence class to obtain an equivalence class set.

[0090] S3: Use constraint information to filter and fuse the data features of each traffic sample to obtain new data features.

[0091] The step S3 specifically comprises the following steps:

[0092] S31: Use the traffic sample statistical feature vector and the must-connect constraint set to construct a neighbor graph using the k-nearest neighbor algorithm and calculate the weight matrix of the neighbor graph. In this step, for any two samples i and j, if the sample pair (i, j) belongs to the must-connect constraint set or i is one of the k neighbors with the smallest distance to j, then i and j have a common edge in the neighbor graph; for samples i and j that have a common edge in the neighbor graph, calculate their value w in the weight matrix. ij The calculation formula is:

[0093]

[0094] where x i and x j are the statistical feature vectors of sample i and sample j respectively, ||x i -x j || 2 is the square of the Euclidean distance between sample i and sample j, indicating their similarity. λ is a constant parameter used to control the degree of attenuation of similarity. Usually, λ is a hyperparameter selected through experiments or cross-validation. For samples i and j that have no common edges in the neighbor graph, their values ​​in the weight matrix are 0.

[0095] S32: Use the traffic sample statistical feature vector and the unconnected constraint set to construct a negative constraint graph and calculate the weight matrix of the negative constraint graph. For any two samples i and j, if the sample pair (i, j) belongs to the unconnected constraint set, i and j have a common edge in the negative constraint graph; for samples i and j that have a common edge in the negative constraint graph, the value in the weight matrix is ​​1, and for samples i and j that do not have a common edge in the negative constraint graph, the value in the weight matrix is ​​0.

[0096] S33: Calculate the Laplacian matrices of the neighbor graph and the negative constraint graph respectively through the weight matrices of the neighbor graph and the negative constraint graph. The calculation formula is as follows:

[0097] L kn =D kn -S kn

[0098] LCL =D CL -S CL

[0099] Among them, S kn is the neighbor graph weight matrix, S CL is the negative constraint graph weight matrix, D is a diagonal matrix, L kn and L CL are the Laplacian matrices of the neighbor graph and the negative constraint graph respectively.

[0100] S34: According to the Laplace matrix and the statistical feature vector of the traffic sample, the Laplace score of each feature is calculated and the flow statistical features below the threshold are removed to obtain the statistical feature vector after feature selection. The formula for calculating the constrained Laplace score in this step is:

[0101]

[0102] Among them, f r is a vector composed of the values ​​of feature r in all samples. The threshold is selected according to the influence of the number of features on the experimental effect. In this embodiment, the constrained Laplace score of each feature is arranged from high to low, and the constrained Laplace score of the 40th feature is selected as the threshold.

[0103] S35: The statistical feature vector after feature selection is merged with the fingerprint feature vector to obtain new data features for each traffic sample. In this step, the feature fusion is performed by the concatenation method.

[0104] S4: Use the K-means clustering algorithm to perform semi-supervised clustering with the new data features as input. At the same time, use constraint information to improve the K-means clustering algorithm during the clustering process to obtain the clustered traffic clusters.

[0105] The step S4 specifically comprises the following steps:

[0106] S41: Randomly select a flow from each equivalence class as the initial cluster center.

[0107] S42: Select the next cluster center by weighted probability distribution based on the square of the distance from each flow in the traffic sample to the selected cluster center. The distance in this step is the Euclidean distance, and the probability of each single flow being selected as the next cluster center is:

[0108]

[0109] Among them, v i is the new data feature of traffic sample i, D(v i) is the Euclidean distance from traffic sample i to its nearest cluster center, and X is the new data feature set of all traffic samples.

[0110] S43: Repeat step S42 until a predetermined number of initial cluster centers are selected.

[0111] S44: Assign each flow in the traffic sample to its own cluster center. In this step, for each single flow v i , select the traffic cluster C with the smallest distance to the cluster center j The formula for the traffic cluster to which this single flow will be assigned is:

[0112]

[0113] where μ j C j The cluster center, μ g is the cluster center of other traffic clusters, and p is the number of cluster centers. Before allocation, first determine whether the single flow violates the no-connection constraint with all the single flows in the traffic cluster to be allocated. If it violates the no-connection constraint, find the next nearest cluster center until it does not violate the no-connection constraint; if it does not violate the no-connection constraint, allocate the single flow to the corresponding traffic cluster, and at the same time, assign the equivalence class to which the currently successfully allocated single flow belongs and directly classify all single flows in the equivalence class into the same traffic cluster as the current sample.

[0114] S45: Update the cluster center. The specific formula for updating the cluster center μ is as follows:

[0115]

[0116] S46: Repeat steps S44-S45 until convergence or the number of iterations reaches a predetermined amount.

[0117] S5: Use the clustered traffic clusters as training data to construct a classifier. Specifically, first obtain each traffic cluster C i The number of labeled single-stream feature vectors n in i If n i The value of is 0, indicating that the traffic cluster does not contain any marked single-flow feature vector, so the traffic cluster C i Mapped as an unknown traffic cluster; if n i If the value of is not 0, the statistic is marked as y j The number of single-stream eigenvectors n ij After the statistics are completed, calculate the traffic cluster C i The posterior probability P(Y=y j |C i )=n ij / n i Finally, traffic cluster Ci Mapped to a traffic cluster with label y, where

[0118] Assume that the traffic class is represented by Ω = {ω1, ..., ω q}, these categories are generated by mapping between clustering results and predefined categories. Each traffic category ω i Use the cluster centroid set M belonging to this category i To indicate: M i ={m j :C j ∈ω i} Among them, C j represents a cluster, m j is the centroid of the cluster. For a test flow x, the classification rule is:

[0119]

[0120] In this way, the method provided by the present invention abandons the traditional method of unknown protocol identification based on ports and deep packet retrieval, and provides a new idea for using semi-supervised learning to identify unknown protocols based on the K-means clustering algorithm. In view of the complexity of unsupervised learning applications, the rationality of the design of self-supervised learning pre-training tasks, and the insufficient use of prior knowledge in semi-supervised learning, this solution extracts traffic statistical features and protocol fingerprint features and combines a small number of known protocol samples with flow correlation to construct constraint information, calculates the constrained Laplace score for feature selection and fusion, and refines and enhances the characterization ability of features. By optimizing the K-means initial clustering center selection and sample allocation process through constraint information, the prior information in the labeled samples can be effectively utilized, the dependence on labeled samples can be reduced, and the accuracy and efficiency of unknown protocol identification can be improved.

[0121] like Figure 2 As shown, the present invention provides an unknown protocol identification system based on semi-supervised clustering, including: a traffic preprocessing module, a constraint information construction module, a multi-dimensional feature construction module, a semi-supervised clustering module, and a classifier construction module.

[0122] The traffic preprocessing module is used to perform flow splitting processing on the collected marked and unmarked data, and extract statistical features and fingerprint features of each single flow respectively;

[0123] The constraint information building module is used to build a must-connect constraint set, a do-not-connect constraint set, and an equivalence class set according to the label of the marked single flow and the correlation between two different single flows;

[0124] The multi-dimensional feature construction module is used to calculate the constrained Laplace score of each statistical feature according to the constraint information, retain the features with scores higher than the threshold and merge them with the fingerprint features to obtain the feature vectors of the marked and unmarked single streams;

[0125] The semi-supervised clustering module is used to mix the marked and unmarked single flows, use the constraint information to improve the K-means algorithm clustering process and cluster the mixed single flows to obtain multiple traffic clusters;

[0126] The classifier construction module is used to construct an NCC classifier according to the traffic cluster.

[0127] Figure 3 An example of a physical structure diagram of an electronic device is provided, wherein the device includes a processor and a storage medium. The storage medium stores a computer program for performing unknown protocol recognition, and the processor is used to execute the computer program to perform the above method.

[0128] like Figure 4 , Figure 5 , Figure 6 , Figure 7 The figure shows some experimental data in the embodiments of the present invention. The improved clustering method in the present invention is compared with four other traditional clustering methods, and the classifier obtained by the present invention is compared with the classifiers of other semi-supervised methods.

[0129] Figure 4 The horizontal axis in the figure corresponds to the clustering purity, standard mutual information, and adjusted Rand coefficient of each method in the bar graph from left to right. Each method from left to right is the improved clustering method of the present invention, DBSCAN, K-means, GMM, and BRICH.

[0130] Figure 5 The comparison of clustering effects between the present invention and other semi-supervised methods in different unknown protocol situations is shown. Even if some samples lack label information, the present invention can not only correctly classify the unlabeled traffic, but also classify the traffic of unknown protocols.

[0131] Figure 6 The performance comparison results of the classifiers of the present invention and other semi-supervised methods are shown.

[0132] Figure 7The dependence of each classifier on the labeled data is further analyzed. When a small number of samples of known protocols are labeled, the clustering effect can be increased, and the effective identification and differentiation of unknown protocol traffic can be achieved. As the number of labeled samples increases, the evaluation indicators of each classifier will increase, because there are more labeled samples to accurately identify more known clusters. The evaluation indicators of the present invention are better than those of the comparative classifiers in all cases of sample labeling numbers, because the flow correlation is used to improve the constraint information when establishing the constraint information, so that when clustering, not only more samples of the same protocol can be assigned to the same cluster, but it is also conducive to cluster mapping.

[0133] It should be noted that, in this document, relational terms such as first and second, etc. are merely used to distinguish one entity or one operation from another entity or another operation, and do not necessarily require or imply any actual relationship or order between these entities or operations.

[0134] The above implementation modes are only used to illustrate the present invention, but not to limit the present invention. Ordinary technicians in the relevant technical field can make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, all equivalent technical solutions also belong to the scope of the present invention. The patent protection scope of the present invention should be defined by the claims.

Claims

1. A method for identifying unknown protocols based on semi-supervised clustering, characterized in that: The method comprises the following steps: S1: Extract data features of traffic samples of labeled and unlabeled protocols; S2: Construct constraint information based on data characteristics; S3: Use constraint information to filter and fuse the data features of each traffic sample to obtain new data features; S4: Use the K-means clustering algorithm to perform semi-supervised clustering using the new data features as input. At the same time, use the constraint information to improve the K-means clustering algorithm during the clustering process to obtain the clustered traffic clusters. S5: Use the clustered traffic clusters as training data to construct a classifier. Wherein, the step S1 specifically includes: S11: Collect labeled and unlabeled traffic samples of different protocols; S12: Divide the data packets with the same source IP address, destination IP address, source port, destination port, and transport layer protocol in the traffic sample into the same flow; S13: extracting data features of each flow, where the data features of each flow include a statistical feature vector and a fingerprint feature vector.

2. The unknown protocol identification method based on semi-supervised clustering according to claim 1 is characterized in that: The statistical features at least include basic flow information (source IP, destination IP, protocol type, flow duration, source port, destination port), data packet size (maximum, minimum, average, standard deviation of single flow and upstream and downstream data packet sizes), number of data packets (number of data packets in a single flow, number of upstream and downstream data packets in a single flow and their ratio), data packet time interval (maximum, minimum, average, standard deviation of single flow and the time interval between arrival of upstream and downstream data packets), rate (flow rate of a single flow, upstream and downstream data packet rates in a single flow and their ratio, in bytes / second), data packet rate (flow rate of a single flow, upstream and downstream data packet rates in a single flow and their ratio, in number / second), bytes (number of bytes in a single flow, number of upstream and downstream bytes in a single flow and their ratio), data packet fragmentation (maximum, minimum, average, standard deviation of the number of upstream and downstream fragments), and these statistical features are combined into a statistical feature vector; Fingerprint features are a set of frequent items that can represent a specific protocol communication mode, including the location information and content information of frequent items, which are combined into a fingerprint feature vector. The constraint information includes a must-connect constraint set, a must-connect constraint set, and an equivalence class set.

3. The unknown protocol identification method based on semi-supervised clustering according to claim 1 is characterized in that: The step S2 specifically includes: S21: For any two marked flows with the same label or any two related flows in the traffic sample, add them as a must-connect constraint to the must-connect constraint set; S22: For two marked flows with different labels, add them as a no-connection constraint to the no-connection constraint set; S23: Classify the flows with indirect association in the must-connect constraint set into the same equivalence class to obtain an equivalence class set; The correlation is as follows: for two flows, if their destination IP, destination port and transport layer protocol are the same and the time interval between them is within a certain range, then the two flows are correlated.

4. The unknown protocol identification method based on semi-supervised clustering according to claim 1 is characterized in that: The step S3 specifically includes: S31: Using the traffic sample statistical feature vector and the must-connect constraint set, a k-nearest neighbor algorithm is used to construct a neighbor graph and calculate the weight matrix of the neighbor graph; S32: construct a negative constraint graph using the traffic sample statistical feature vector and the no-connect constraint set and calculate a weight matrix of the negative constraint graph; S33: Calculate the Laplacian matrices of the neighbor graph and the negative constraint graph respectively through the weight matrices of the neighbor graph and the negative constraint graph; S34: according to the Laplace matrix and the statistical feature vector of the traffic sample, the Laplace score of each feature is calculated and the traffic statistical features below the threshold are removed to obtain the statistical feature vector after feature selection; S35: The statistical feature vector after feature selection is merged with the fingerprint feature vector to obtain a new data feature of each traffic sample; Among them, the statistical feature vector after feature selection is fused with the fingerprint feature vector by using a serial fusion and splicing method.

5. The unknown protocol identification method based on semi-supervised clustering according to claim 4 is characterized in that: Using the traffic sample statistical feature vector and the must-connect constraint set, the k-nearest neighbor algorithm is used to construct a neighbor graph and calculate the weight matrix of the neighbor graph, including: For any two samples i and j, if the sample pair (i, j) belongs to the must-connect constraint set or i is one of the k neighbors with the smallest distance to j, then i and j have a common edge in the neighbor graph; For samples i and j that have common edges in the neighbor graph, calculate their value w in the weight matrix ij , the calculation formula is: where x i and x j are the statistical feature vectors of sample i and sample j respectively, ||x i -x j || 2 is the square of the Euclidean distance between sample i and sample j, indicating their similarity. λ is a constant parameter used to control the degree of attenuation of similarity. λ is a hyperparameter selected through experiments or cross-validation. For samples i and j that have no common edges in the neighbor graph, their values ​​in the weight matrix are 0.

6. The unknown protocol identification method based on semi-supervised clustering according to claim 5 is characterized in that: The negative constraint graph is constructed using the traffic sample statistical feature vector and the unconnected constraint set, and the weight matrix of the negative constraint graph is calculated, including: For any two samples i and j, if the sample pair (i, j) belongs to the unconnected constraint set, then i and j have a common edge in the negative constraint graph; For samples i and j that have common edges in the negative constraint graph, their values ​​in the weight matrix are 1; For samples i and j that have no common edges in the negative constraint graph, their values ​​in the weight matrix are 0.

7. The unknown protocol identification method based on semi-supervised clustering according to claim 1 is characterized in that: The step S4 specifically includes: S41: Randomly select a flow from each equivalence class as the initial cluster center; S42: Select the next cluster center by weighted probability distribution based on the square of the distance from each flow in the traffic sample to the selected cluster center; S43: Repeat step S42 until a predetermined number of initial cluster centers are selected; S44: assigning each flow in the traffic sample to its own cluster center; S45: Update cluster center; S46: Repeat steps S44-S45 until convergence or the number of iterations reaches a predetermined amount.

8. The unknown protocol identification method based on semi-supervised clustering according to claim 7 is characterized in that: The next cluster center is selected by weighted probability distribution based on the square of the distance from each data point to the selected cluster center. The probability of each flow being selected as the next cluster center is: Among them, v i is the new data feature of traffic sample i, D(v i ) is the Euclidean distance from traffic sample i to its nearest cluster center, and X is the new data feature set of all traffic samples; Assign each flow in the traffic sample to its own cluster center, including: For the first single flow, the traffic cluster with the closest cluster center is selected as the traffic cluster to which the first single flow will be assigned; Determine whether the first single flow and all second single flows in the traffic cluster to be allocated violate the no-connection constraint. If so, find the next nearest cluster center until the no-connection constraint is not violated. When the no-connection constraint is not violated, the first single flow is assigned to the corresponding traffic cluster, and all third single flows in the equivalence class to which the first single flow belongs are directly classified into the same traffic cluster as the first single flow.

9. An unknown protocol identification system based on semi-supervised clustering, comprising: Traffic preprocessing module, constraint information construction module, multi-dimensional feature construction module, semi-supervised clustering module, classifier construction module, The traffic preprocessing module is used to perform flow splitting processing on the collected marked and unmarked data, and extract statistical features and fingerprint features of each single flow respectively; The constraint information building module is used to build a must-connect constraint set, a do-not-connect constraint set, and an equivalence class set according to the label of the marked single flow and the correlation between two different single flows; The multi-dimensional feature construction module is used to calculate the constrained Laplace score of each statistical feature according to the constraint information, retain the features with scores higher than the threshold and merge them with the fingerprint features to obtain the feature vectors of the marked and unmarked single streams; The semi-supervised clustering module is used to mix the marked and unmarked single flows, use the constraint information to improve the K-means algorithm clustering process and cluster the mixed single flows to obtain multiple traffic clusters.

10. A computer device, characterized in that: The equipment includes: A storage medium storing a computer program for performing unknown protocol recognition; A processor, configured to execute the computer program; The computer program causes the processor to implement the steps of the method according to any one of claims 1 to 8, The storage medium includes a hard disk, a flash disk or other non-volatile storage devices.

Citation Information

Patent Citations

  • Classification detection method facing network abnormal data flow

    CN106060039A

  • Network traffic classification method based on constraint fuzzy clustering and granular computing

    CN111786903A

  • Application flow automatic classification method based on semi-supervised learning

    CN112187664A

  • Creating and using multiple packet traffic profiling models to profile packet flows

    US20130100849A1

Cited By

  • Encrypted traffic classification method and device, electronic equipment, computer readable storage medium and computer program product

    CN121125631A

  • Encrypted traffic classification method and apparatus, electronic device, computer readable storage medium and computer program product

    CN121125631B

  • Multi-task semi-supervised TSK fuzzy system and method for MI electroencephalogram signal recognition

    CN122346737A