Mining and Utilization Methods of New Types of Encrypted Network Traffic Packets in Distributed Scenarios

By collaborating between multiple network traffic monitoring nodes, a new type of network traffic packet detection model is trained, and globally consistent category division and label allocation is carried out, the problem of monitoring and management of new type encrypted network traffic packets in distributed scenarios is solved, and efficient detection and management of new type of traffic packets is achieved.

CN115134128BActive Publication Date: 2025-05-16湖南工商大学
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210665404.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-11
Publication Date
2025-05-16
Estimated Expiration
2042-05-11

AI Technical Summary

Technical Problem

In distributed scenarios, it is difficult for the existing technology to effectively monitor and manage new types of encrypted network traffic packets, especially when a single network monitoring node has limited capacity and new types of traffic packets frequently occur.

Method used

By collaborating between multiple network traffic monitoring nodes, a new type of network traffic packet detection model is trained, and a globally consistent category division and label allocation is adopted in distributed scenarios. These labeled new type of traffic packet samples are used to quickly update the existing model.

Benefits of technology

The global consistent category division and label allocation of new types of encrypted network traffic packets distributed on different network traffic monitoring nodes is realized, which expands the model's category identification capabilities and improves the detection and management capabilities of new types of traffic packets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115134128B_ABST
    Figure CN115134128B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for mining and utilizing new types of encrypted network traffic packets in a distributed scenario. New types of traffic packets detected on different network nodes in a distributed scenario contain valuable pattern information. The scheme designed by the present invention can perform globally consistent classification and category label assignment for new types of encrypted network traffic packets distributed on different network traffic monitoring nodes. The scheme can also use these labeled new types of traffic packet samples to quickly update various existing global models (such as feature vector extraction models, new types of encrypted network traffic packet detection models, existing types of encrypted network traffic classification models, etc.) to expand the category recognition capabilities of these models.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network security, and more particularly to a method for monitoring and managing network traffic packets in a distributed scenario. Background Art

[0002] Network traffic packet classification is a vital task in network management and cyberspace security. Network management departments usually need to classify network traffic packets into different categories, and then use different routing or firewall configuration strategies for different types of network traffic packets. For example, we can classify network traffic packets according to application categories and assign different priorities to different categories of network traffic packets to ensure the network quality of service (QoS) of high-priority services. For another example, network packet classification can be used for network intrusion detection. By classifying network data packets into benign traffic packets and malicious traffic packets, the purpose of network anomaly detection can be achieved.

[0003] Currently, most network traffic is encrypted. Secure communication protocols, such as SSL (Secure Sockets Layer) and TLS (Transport Layer Security), have been introduced in most network applications to improve their respective security performance. At the same time, many malware encrypt their network traffic packets to evade detection by firewalls and network intrusion detection systems. Since the payload of encrypted network traffic packets is in an encrypted state, this poses a challenge to traditional traffic classification methods such as deep packet inspection (DPI). Network traffic classifiers based on machine learning usually require manual feature design and selection, which is difficult to implement and has low classification accuracy.

[0004] In recent years, deep learning technology has been introduced into encrypted network traffic classification scenarios. However, network encrypted traffic classification schemes based on deep learning face many challenges that are out of touch with real scenarios.

[0005] First, the training of deep learning models requires a large number of sample support, otherwise it is easy to induce overfitting problems. Deep learning models are generally complex, with many parameters to be trained. Building a high-precision deep learning-based encrypted traffic classifier requires the support of a large number of well-labeled training samples. However, it is not easy to collect a large amount of correctly labeled encrypted traffic. Since the traffic packet payload is in an encrypted state, the cost of encrypted network traffic type analysis and labeling is very high. The capacity of a single monitoring node is limited, and the number of encrypted traffic packet samples that can be labeled is limited.

[0006] Secondly, a classification model with application value should be able to identify as many traffic categories as possible. However, a single network monitoring node has limited coverage and can collect limited sample types, which limits the model's recognition capabilities. The distribution of network traffic packets usually has certain regional characteristics. For example, the types of network traffic generated by different types of network users are not completely the same. For another example, network viruses usually break out in a certain area and then spread to other areas.

[0007] In addition, new types of traffic packets emerge in an endless stream, and the models trained based on existing traffic packet samples cannot correctly classify them. In real application scenarios, the types of network traffic are not fixed, and we often encounter a large number of new types of network traffic packets. There are many reasons for the frequent appearance of new types of network traffic packets. On the one hand, various new network applications emerge in an endless stream, and new network applications will inevitably lead to new network traffic patterns. On the other hand, in order to evade network monitoring, malicious network users usually change their behavior patterns, which leads to changes in malicious network traffic patterns.

[0008] Therefore, it is necessary to study the distributed network monitoring scenario, where existing and new types of encrypted network traffic packets coexist, and the monitoring and management problems of encrypted network traffic packets that are closer to the actual situation. In the scenario studied by the present invention, there are multiple network monitoring nodes. Multiple network monitoring nodes (referred to as nodes) are distributed at the entrance of different network areas to monitor the network traffic in the area. Each node has accumulated a certain amount of labeled network traffic samples. We refer to the network traffic types corresponding to these labeled samples as existing types. Correspondingly, the new type refers to the fact that no samples of this category have been labeled. In this scenario, existing and new types of encrypted network traffic packets coexist. The newly received traffic packet sample may be an existing type sample or a new type sample. Summary of the invention

[0009] The technical problem to be solved by the present invention is to propose a method for mining and utilizing new types of encrypted network traffic packets in a distributed scenario in view of the shortcomings of the existing technology. The technical solution of the present invention is:

[0010] A method for mining and utilizing new types of encrypted network traffic packets in a distributed scenario, characterized in that it includes the following steps:

[0011] (1) Preparation stage: multiple network traffic monitoring nodes (referred to as "nodes") monitor the network traffic of different network areas they are responsible for respectively; each node independently collects a certain number of network traffic packet samples that have been labeled by category (referred to as "labeled samples"); multiple network traffic monitoring nodes cooperate with each other to train a new type of network traffic packet detection model; the new type of network traffic packet refers to a network traffic packet sample of this type that has not yet been labeled by category;

[0012] (2) Detection of new types of traffic packets: Each node detects new types of network traffic packets from its newly received network traffic packets. The mining and utilization of new types of traffic packets are performed in a periodic manner, and each round of mining and utilization operations is based on all new types of traffic packets detected in the current period.

[0013] (3) Subcategory discovery: Each node independently performs local clustering operations on the new type of traffic packets detected in this round (within the current cycle time); each node independently assigns labels to each subcategory sample in its own clustering result; in the local clustering results, new type traffic packet samples of the same subcategory will be assigned the same local label; labels of different local subcategories are different;

[0014] (4) Extraction of local subcategory feature vectors: Each node selects a globally unified benchmark; based on the globally unified benchmark, each node extracts a globally consistent category feature vector for each local subcategory; each node uploads the feature vectors of the local subcategory, together with their corresponding local subcategory labels, to the aggregation node;

[0015] (5) Globally consistent category labeling: The sink node collects subcategory feature vectors and local label information from different nodes; the sink node performs global clustering based on all collected subcategory feature vectors; the sink node assigns global labels to each subcategory in the global clustering result; the sink node establishes a mapping scheme between local labels and global labels for each node, and returns the mapping scheme to the corresponding node; each node uses the received mapping scheme to assign a global label to each subcategory sample;

[0016] (6) Model update: Multiple network traffic monitoring nodes expand the model and use the samples collected by each node and assigned global labels as described in step (5) to train the expanded model in a collaborative manner until the model converges or reaches a pre-set error threshold.

[0017] As a further optimization, the specific steps of step (4) are as follows:

[0018] (4.1) Design of globally consistent benchmark model:

[0019] The globally consistent benchmark model is defined as: y = f μ (x) = f e (f θ (x)) = argmax(softmax(f θ (x))); sub-model f θ Extract the feature model for encrypted network traffic packets; each node uses the globally optimal model parameter θ * Pair Model f θ Initialize the submodel f e It does not contain the parameters to be optimized and does not need to be initialized;

[0020] (4.2) Training of subcategory incremental models: Each node independently trains an incremental model for its own different local subcategory samples; the optimization equation for incremental training is:

[0021] (4.3) Subcategory feature extraction based on the incremental model: Each node selects a parameter subset from each subcategory model parameter according to the same rule as the feature vector of the subcategory; each node uploads each local subcategory feature vector and local subcategory label to the aggregation node;

[0022] (4.4) Globally consistent subcategory label assignment: The aggregation node globally clusters all local subcategories based on the collected subcategory feature vectors, and assigns different global labels to each global subcategory according to the global clustering results; the aggregation node establishes a mapping scheme between local subcategory labels and global labels for the local subcategories of each node based on the collected local labels of the subcategories and the reallocated global labels, and feeds back the mapping relationship to the corresponding nodes; each node modifies the local category label of its own sample into the global category label according to the received mapping scheme.

[0023] As a further optimization, the specific steps of step (6) are as follows:

[0024] (6.1) Model expansion: The model is expanded according to the total number of newly added categories and the total number of newly added samples. When the number of newly added categories and samples is small, the number of neurons in the output layer of the model is increased. When the number of newly added categories and samples is very large, the number of layers in the middle layer or the number of neurons in each layer needs to be increased.

[0025] (6.2) Model initialization: The original basic model parameters in the extended model are initialized using the existing optimal feature parameters; the parameters of each neuron in the extended part of the model are initialized using random numbers;

[0026] (6.3) Optimization equation: The optimization equation is defined as where f' θ It is an expanded model;

[0027] (6.4) Model training: Multiple network traffic monitoring nodes use their own collected samples that have been assigned global labels to train the expanded model in a collaborative manner until the model converges or reaches a pre-set error threshold.

[0028] Beneficial effects:

[0029] The solution adopted by the present invention designs a method for mining and utilizing new types of encrypted network traffic packets in distributed scenarios. New types of traffic packets detected on different network nodes in distributed scenarios contain valuable pattern information. The solution designed by the present invention can perform globally consistent classification and category label assignment for new types of encrypted network traffic packets distributed on different network traffic monitoring nodes. The solution can also use these labeled new types of traffic packet samples to quickly update various existing global models (such as: feature vector extraction models, new types of encrypted network traffic packet detection models, existing types of encrypted network traffic classification models, etc.) to expand the category recognition capabilities of these models. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 Schematic diagram of feature extraction model structure

[0031] Figure 2 (a) Distribution of top-3 elements of feature vector of new type traffic samples

[0032] Figure 2(b) Distribution of top-3 elements of feature vectors of existing traffic samples

[0033] Figure 3 High confidence new type traffic packet sample extraction model

[0034] Figure 4 Incremental model parameter category expressiveness

[0035] Figure 5 (a) Two-dimensional spatial view of network traffic packets (the first and second largest element dimensions of the feature vector)

[0036] Figure 5(b) Two-dimensional spatial view of network traffic packets (the first and third largest element dimensions of the feature vector)

[0037] Figure 5(c) Two-dimensional spatial view of network traffic packets (the second and third largest element dimensions of the feature vector)

[0038] Figure 6 (a) The expressive power of the bias parameter of the first layer

[0039] Figure 6(b) The expressive power of kernel parameters in the first layer

[0040] Figure 7 (a) The expressive power of the bias parameter of the last layer

[0041] Figure 7(b) The expressive power of kernel parameters in the last layer

[0042] Figure 8 Globally consistent class label assignment DETAILED DESCRIPTION

[0043] The specific implementation process of the present invention is as follows:

[0044] This paper studies a problem of monitoring and managing encrypted network traffic packets in a distributed scenario that is similar to a real-life scenario. There are multiple network monitoring nodes (referred to as "nodes") in this problem scenario, and they independently monitor and manage the encrypted network traffic packets in the area under their jurisdiction. Each network monitoring node has accumulated some labeled encrypted network traffic packet samples. The number and type of labeled samples on each node are limited, and it is impossible to complete the training of complex deep learning models alone.

[0045] The newly received encrypted network traffic packets of each node include both existing type traffic packets and new type traffic packets. Existing type encrypted network traffic packets (referred to as "existing type traffic packets") refer to some encrypted network traffic packet samples of this type that have been assigned the correct category labels. Such labeled samples are called labeled samples, or simply labeled samples. New type encrypted network traffic packets (referred to as "new type traffic packets") refer to no encrypted network traffic packet samples of this type that have been assigned category labels. We assume that different network monitoring nodes assign the same category labels to the same type of encrypted traffic packet samples.

[0046] In order to solve the problem of limited number and type of labeled network traffic packets in a single network node, different nodes will use their own labeled samples to collaboratively train models. By integrating sample resources from multiple nodes, we can increase the number of labeled samples used for model training to avoid overfitting problems, and enable the trained model to learn the differences in traffic pattern characteristics in different network areas.

[0047] In response to the research problem, the inventor has mainly carried out specific research in the following three aspects.

[0048] (1) Feature extraction method for encrypted network traffic packets in distributed scenarios: This feature extraction model can be used in other methods such as new type of traffic packet detection, new type of traffic packet labeling, and existing type of traffic packet classification.

[0049] (2) Detection method of new types of encrypted network traffic packets in distributed scenarios: In the newly received network traffic, existing types and new types of encrypted network traffic packets coexist. If we directly classify the newly received network traffic, the new type of traffic packets will be mistakenly classified as an existing type, resulting in classification errors. Therefore, it is necessary to detect and separate the new type of encrypted network traffic packets from the newly received network traffic at different nodes.

[0050] (3) Mining and utilizing new types of encrypted network traffic packets in distributed scenarios: New types of traffic packets detected on different network nodes contain valuable pattern information. We will study how to mine this information and use it to update existing models.

[0051] 1. Feature extraction method of encrypted network traffic packets in distributed scenarios

[0052] This section introduces the feature extraction method of encrypted network traffic packets in distributed scenarios. This method can be used in other methods such as new type traffic detection, new type traffic labeling, existing type traffic classification, etc. The feature extraction method of encrypted network traffic packets in distributed scenarios mainly includes the following steps: First, design a feature extraction model. This model is used to directly convert the original encrypted traffic packet into a feature vector. Then, design a training method for the feature extraction model in a distributed scenario. The specific steps are as follows:

[0053] (1) Preparation stage: Multiple network traffic monitoring nodes monitor the network traffic of different network areas they are responsible for respectively; each node independently collects a certain number of network traffic packet samples that have been labeled (assigned category labels) (referred to as “labeled samples”);

[0054] (2) Construction of feature extraction model: Network traffic packet feature extraction model f θ It can be expressed as v = f θ (x), where x is the encrypted network traffic packet, and v is the feature vector extracted by the model; the feature extraction model f θ It includes at least one one-dimensional convolution (1D CNN) layer and one attention (Attention) layer; the output of the Attention layer is transformed into a set of weights; the set of weights is used as the weights of different channels of the one-dimensional convolution layer to change the original output value of the one-dimensional convolution layer; as an optimization, the feature extraction model f θ It can also include one-dimensional pooling layers, fully connected layers, and activation layers; Figure 1The combination of one-dimensional convolutional layer and attention layer in the feature extraction model is demonstrated. In specific implementation, the feature extraction model generally includes multiple convolutional layers. The commonly used convolutional layer structures in the field of computer vision are mainly two-dimensional convolutional layer and three-dimensional convolutional layer. Some researchers have applied two-dimensional convolutional layer and three-dimensional convolutional layer to encrypted traffic classification scenarios. However, network traffic is essentially sequential data, which is a one-dimensional byte stream. Therefore, this feature extraction model will use one-dimensional convolutional layer and one-dimensional pooling layer as the basic components of convolutional neural network. At the same time, the Attention layer is also introduced in the feature extraction model. The Attention layer uses the output of a certain convolutional layer as its input to capture the feature differences of different channels of the convolutional layer. The Attention layer converts the captured difference information into a set of weights by combining with Softmax. This set of weights will be used as the weights of different channels of the convolutional layer to change the weights of the original output values ​​of the convolutional layer, thereby realizing dynamic weighting of different output features of the convolutional layer. As an optimization, the feature extraction model f θ It can also include one-dimensional pooling layers, fully connected layers, and activation layers; Figure 1 In the above figure, “1D CNN” stands for a subnetwork of artificial neural networks with one-dimensional convolutional layers as the main component. Figure 1 The "other layers" in CNN usually consist of components such as one-dimensional convolutional layers, pooling layers, and fully connected layers.

[0055] The depth and structure of the layers in the middle layer of the feature extraction model need to be determined comprehensively based on factors such as the number of training samples and the performance of the machine. According to deep learning theory, in general, when the number of samples is large enough, the more complex the model structure and the deeper the layers, the stronger the model's expressive power. The number of neurons in the output layer of the feature extraction model is kept at the same order of magnitude as the number of existing encrypted network traffic packet categories, and can generally be set to be equal to or slightly larger than the number of existing types. The following is a relatively simple implementation example of a feature extraction model. This implementation example is suitable for scenarios with a limited number of training samples. The input of the implementation example model is a one-dimensional encrypted network traffic packet x. The model output is the corresponding feature vector v. The model consists of 7 layers. It includes 3 convolutional layers, two pooling layers, and two fully connected layers. The Attention layer is inserted after the second convolutional layer in the form of a bypass. The results calculated by the Attention layer are used to dynamically weight the output features of the second convolutional layer. Each output of the second convolutional layer will be multiplied by the corresponding weight in the weight group provided by the Attention layer to obtain a weighted output. The weighted output of the second convolutional layer will continue to be input into subsequent modules for processing.

[0056] In order to extract the feature model f θTo conduct training, we need to solve the following two problems. First, the input of the feature extraction model is the traffic packet sample, and the output is the feature vector. However, we do not have prior knowledge about the optimal feature vector and cannot directly guide the training process. Second, the sample resources are distributed on multiple independent nodes, so it is necessary to construct a collaborative training mechanism.

[0057] In theory, the feature extractor should not change the category affiliation of traffic packet samples. In other words, the positions of traffic packet samples of the same type in the feature space should be close. Therefore, we can use the category labels of the labeled samples to supervise and guide the training process and optimization direction of the feature extractor. To achieve the above idea, we construct the interface model f based on the feature vector v e .

[0058] (3) Construction of interface model: Interface model f e It is composed of two nested modules, Softmax and Argmax; the interface model can be expressed as y = f e (v) = argmax(softmax(v));

[0059] (4) Construction of optimization equation: The optimization equation can be expressed as Where l is the loss function;

[0060] The feature extraction model is constructed based on deep neural network technology and requires a large number of labeled training samples as support. In order to increase the number of samples used for model training and improve the accuracy of the model, we construct a distributed training solution for the model to achieve model-level sharing of sample resources accumulated by each node.

[0061] (5) Distributed training of the model: Multiple network traffic monitoring nodes (referred to as “nodes”) use the labeled network traffic packet samples collected by each node in step (1) and work together to train the feature extraction model f in step (2) according to the optimization equation given in step (4). θ Training is performed until the model converges or reaches a pre-set error threshold. The specific steps of distributed training of the model are as follows:

[0062] (5.1) Model initialization: A node is selected as the sink node. The sink node first extracts the feature model f θ The parameters of are randomly initialized to θ0, and then θ0 is sent to other nodes;

[0063] (5.2) Local model training: Node i uses the received θ0 to train f θ Initialize and construct local optimization equations Node i uses the locally accumulated and labeled encrypted network traffic data set to optimize the model f based on the above optimization equation θ Optimize and get the optimized model parameters Node i feeds back the optimization result to the sink node

[0064] (5.3) Model parameter generation in this stage: The aggregation node receives feedback results from each participating node Calculate its mathematical expectation The model parameters for this round of distributed training are The sink node sends the model parameter θ1 of this stage to other nodes;

[0065] (5.4) Repeat steps (5.2)-(5.3) until the model converges or reaches a preset error threshold, thereby obtaining the current optimal model parameter θ * ;

[0066] (5.5) All nodes obtain the current optimal model parameter θ from the sink node * , and construct the current optimal feature extraction model f θ , which is used to extract the feature vector of network traffic packets.

[0067] 2. Detection Method of New Types of Encrypted Network Traffic Packets in Distributed Scenarios

[0068] In the newly received network traffic, existing types and new types of encrypted network traffic packets coexist. If the newly received network traffic is directly classified, the new type of traffic packet will be mistakenly classified as an existing type, resulting in a classification error. Therefore, it is necessary to detect and separate the new type of encrypted network traffic packets from the newly received network traffic from different nodes. To this end, we have designed a new type of encrypted network traffic packet detection method for distributed scenarios, which is used to detect and separate new types of encrypted network traffic packets from the newly received network traffic from different nodes. The inventors proposed a new type of network traffic packet detection method in a distributed scenario, comprising the following steps:

[0069] (1) Preparation stage: multiple network traffic monitoring nodes monitor the network traffic of different network areas they are responsible for respectively; each node independently collects a certain number of network traffic packet samples that have been tagged (assigned category labels) (referred to as "tagged samples"); the new type of network traffic packet refers to a network traffic packet sample of this type that has not yet been tagged;

[0070] (2) Feature extraction model training: The feature extraction model can be expressed as v = f θ(x), where x is the network traffic packet, and v is the feature vector extracted by the model; multiple network traffic monitoring nodes (referred to as "nodes") use their own collected labeled network traffic packet samples and adopt a collaborative approach to train an optimal parameter θ for the feature extraction model * ;

[0071] (3) Acquisition of positive samples for detection model training: Each network node uses a feature vector extraction model to extract a feature vector from a newly received network traffic packet, and compares the vector with a preset vector to determine whether the network traffic needs to be used as a positive sample in the new type of network traffic packet detection model training sample; the labels of all positive samples are set to the same value; the specific steps for obtaining positive samples for detection model training are as follows:

[0072] (3.1) Define the threshold vector [α1, α2, .., α k ]:

[0073] The length of the threshold vector is k; each element α of the threshold vector i Each element α of the threshold vector i It consists of two parts: flag and value. Flag∈{+,-}, value∈(0,1). If flag is negative, it means that value defines the right boundary and its left boundary is 1. If flag is positive, it means that value defines the left boundary and its right boundary is 0.

[0074] The threshold vector [α1, α2, .., α k] is determined based on historical data. The following is an example. We input the same number of existing and new types of traffic packet samples into the feature extraction model to obtain their respective feature vectors. In order to form a threshold result that can be quantitatively compared for use in subsequent solutions, we use the Softmax module to normalize the feature vector. For the feature vector of each sample, we sort the elements of the feature vector in descending order. Most of the element values ​​in each vector are close to 0, which is at the same order of magnitude as the noise error, and there is little significance in analyzing and comparing them. Therefore, we only record the top-k elements of each vector. For existing and new types of traffic packets, we draw histograms of their top-k feature vector elements for comparison. Figure 2 is the histogram statistics of the top-3 elements of the feature vector. The horizontal axis is the value of the element, and the vertical axis is the number of samples corresponding to the element value. Figure 2(a) shows the distribution of the top three elements of the new type of traffic feature vector (corresponding to k=1, 2, 3 in the figure), and Figure 2(b) shows the distribution of the top three elements of the existing type of traffic feature vector (corresponding to k=1, 2, 3 in the figure). The first column corresponds to the distribution of the top-1 vector elements, and the second and third columns correspond to the distribution of the second and third largest vector elements.

[0075] Although there is some overlap between the distribution intervals of new and existing samples, we can still select a distribution interval to obtain new samples with very high credibility. Taking Figure 2 as an example, when the top 3 elements of the sample output vector are in the intervals [0, 0.75], [0.2, 1], and [0.1, 1], the credibility of this type of sample being a new sample is very high. For the example in Figure 2, the threshold vector can be expressed as [0.75, -0.2, -0.1]. With the help of the above threshold vector, we can construct Figure 3 The high-confidence new type traffic packet sample extraction model shown.

[0076] (3.2) Extract the top-k elements of the network traffic packet feature vector: extract the feature vector of the network traffic packet; the length of the feature vector is greater than or equal to k; sort the extracted feature vectors and retain the largest k elements (top-k elements), denoted as [v′1, v′2, .., v′ k ];

[0077] (3.3) Obtain positive samples for detection model training: By comparing the threshold vector [α1, α2, .., α k ] and the top-k elements of the eigenvector [v′1, v′2, .., v′ k ] to determine whether it is a new type of sample with high confidence; compare [v′1, v′2.., v′ k] are located in [α1, α2, .., α k ] is within the interval indicated by the k elements of [v′1, v′2, .., v′ k ] are located in [α1, α2, .., α k ] is within the interval indicated by the k elements of ], a positive sample label is set for the sample and it is added to the positive sample set used for detection model training. The samples in the positive sample set represent new types of samples with very high credibility. The specific algorithm is as follows:

[0078]

[0079] (4) Obtaining negative samples for detection model training: Negative samples represent existing types and are derived from existing labeled network traffic packet sample sets; each network node randomly selects a certain number of samples from the labeled network traffic packet samples described in step (1) as negative samples; the number of selected negative samples is the same as or similar to the number of positive samples; the labels of all negative samples are set to the same value; the labels of negative samples should be different from those of positive samples, for example, they can be set to 0 and 1 respectively.

[0080] (5) Construction of new type traffic packet detection model: New type traffic packet detection model f n By b and f θ The two sub-models are combined in series and can be expressed as y′=f n (x) = f b (f θ (x));

[0081] (6) Construction of optimization equation: The optimization equation is expressed as Where l is the loss function;

[0082] (7) Distributed training of the model: Multiple network traffic monitoring nodes (referred to as “nodes”) use the samples collected by each node as described in steps (3)-(4) to cooperate with each other and train the new type of traffic packet detection model f in step (5) according to the optimization equation given in step (6). n Training is performed until the model converges or reaches a pre-set error threshold. The specific steps of distributed training of the model are as follows:

[0083] (7.1) Model initialization: The model initialization parameters are defined as n0 = [b0, θ * ]; where θ * is the sub-model f θThe optimal model parameters already exist in each node; the aggregation node only needs to adjust the sub-model f b The parameter b0 is randomly initialized and the initialization result is sent to each node;

[0084] (7.2) Model construction: Each node builds a new type of traffic packet detection model y = f n (x) = f b (f θ (x)), using the received model initialization parameters b0 and sub-model f θ The current optimal parameter θ * , initialize the model parameters;

[0085] (7.3) Local model training:

[0086] First, node i constructs the local optimization equation

[0087] Then, node i uses the local training sample set to train model f n Optimize and get the optimized model parameters The training set includes new type samples (i.e., positive samples) and existing type samples (i.e., negative samples); after the training is completed, node i will feedback the optimization results to the sink node

[0088] (7.4) Generation of model training results for this round: The aggregation node receives feedback results from each participating node Calculate its mathematical expectation Then we can get the optimization result of this round of distributed training. The aggregation node sends the optimization result n1 of this round to each node;

[0089] (7.5) Repeat steps (7.3)-(7.4) until the model converges or reaches a preset error threshold, thereby obtaining the final model parameter n * .

[0090] After the distributed training is completed, all nodes obtain the current optimal model parameter n from the aggregation node. * , and construct the current optimal new type of traffic packet detection model f n* (The model subscript is “n*”). This model detects and identifies all newly received network traffic packets in this round of time interval to separate new types of traffic packets. New type traffic packet detection model f n* In the training process, the existing type and the new type are introduced as a comparison. This allows the model to learn the difference feature information between the new type and the existing type more comprehensively. Therefore, compared with the previous simple threshold segmentation method, the new type traffic packet detection model fn* The detection capability will be greatly improved.

[0091] 3. Mining and Utilization Methods of New Types of Encrypted Network Traffic Packets in Distributed Scenarios

[0092] This section mainly includes two aspects. The first is a globally consistent category label assignment method. The new encrypted network traffic packets on different nodes are divided into different subclasses, and globally unified labels are assigned to samples of different subclasses. The second is an existing model update method. The existing model is updated in a distributed mode using samples that have been globally consistently labeled on each node.

[0093] There are challenges in assigning globally consistent category labels. New types of traffic packets can be further divided into different types. For new types of traffic packets with similar pattern features but distributed on different nodes, they should have the same category label. The most direct way to achieve global unified labeling is to let each node upload the acquired new type of traffic packets to the server, and the server will perform unified labeling. However, due to the large amount of data in the original traffic packets, this method is not suitable. If each node performs dimensionality reduction operations independently and uploads the dimensionality reduction results, the communication overhead can be reduced. However, the dimensionality reduction results of different nodes are not globally unified and cannot be directly compared to achieve global unified category labeling.

[0094] In the method proposed by the present invention, the globally consistent category labeling of new-type traffic packets consists of three processes: (1) local sub-category division and local sub-category labeling of new-type traffic packets. Each node clusters its own new-type traffic packets into different sub-categories, and assigns appropriate local labels to each new-type sample based on the category division results. It should be noted that since each node independently labels its samples, the local labels of similar samples located at different nodes are usually not the same. (2) Globally consistent feature extraction of local categories. Each node extracts globally consistent category features for each local category and uploads them to the server. Ensuring the global consistency of the extracted category features is the premise and basis for the next step. (3) Globally consistent category label assignment. The server divides the local categories into different global categories based on the similarity between the feature data uploaded by each node, and assigns corresponding global labels to each global category. Each local node will replace its own local category label with a global label, thereby achieving global consistency of the category label. A method for mining and utilizing new-type encrypted network traffic packets in a distributed scenario comprises the following steps:

[0095] (1) Preparation stage: multiple network traffic monitoring nodes (referred to as "nodes") monitor the network traffic of different network areas they are responsible for respectively; each node independently collects a certain number of network traffic packet samples that have been labeled (referred to as "labeled"); multiple nodes train a new type of encrypted network traffic packet detection model through the above-mentioned new type of encrypted network traffic packet detection method in a distributed scenario; the new type of network traffic packet refers to a network traffic packet sample of this type that has not yet been labeled;

[0096] (2) Detection of new types of traffic packets: Each node detects new types of network traffic packets from its newly received network traffic packets. The mining and utilization of new types of traffic packets are performed in a periodic manner, and each round of mining and utilization operations is based on all new types of traffic packets detected in the current period.

[0097] (3) Local sub-category label assignment: Each node independently performs local clustering operations on the new type of traffic packets detected in this round (within the current cycle time); each node independently assigns labels to each sub-category sample in its clustering result; in the local clustering results, new type traffic packet samples of the same sub-category will be assigned the same local label; labels of different local sub-categories are different;

[0098] (4) Extraction of local subcategory feature vectors: Each node selects a globally unified benchmark; based on the globally unified benchmark, each node extracts a globally consistent category feature vector for each local subcategory; each node uploads the feature vectors of the local subcategory, together with their corresponding local subcategory labels, to the aggregation node; the specific steps are as follows:

[0099] (4.1) Design of globally consistent benchmark model:

[0100] The globally consistent benchmark model is defined as: y = f μ (x) = f e (f θ (x)) = argmax(softmax(f θ (x))); sub-model f θ Extract the feature model for encrypted network traffic packets; each node uses the globally optimal model parameter θ * Pair Model f θ Initialize the submodel f e It does not contain the parameters to be optimized and does not need to be initialized;

[0101] (4.2) Training of subcategory incremental models: Each node independently trains an incremental model for its own different local subcategory samples; the optimization equation for incremental training is:

[0102] It should be noted that although the incremental model training process uses the traditional deep learning model training method, its training sample composition, training purpose and training cost are different. Traditional deep learning training data contains samples of multiple different categories, and its purpose is to learn the feature information contained in samples of different categories to improve model accuracy. Due to the large differences between samples of different categories, the convergence speed of the model is slow, and the training cost is high. In the incremental model training process designed in this scheme, the training samples come from the same local subcategory, and its purpose is to learn the feature information of the single category data. Due to the small differences between samples of the same category, the model converges very quickly. Experiments show that even after several rounds of epoch training, the parameters can be guaranteed to achieve very good category representative performance. Taking the VPNNonVPN dataset as an example, 100 single-category training sample sets are generated by random sampling, and the incremental model training test is performed independently. When epoch>2, the training accuracy of all 100 times is close to 1.

[0103] (4.3) Subcategory feature extraction based on the incremental model: Each node selects a parameter subset from each subcategory model parameter according to the same rule as the feature vector of the subcategory; each node uploads each local subcategory feature vector and local subcategory label to the aggregation node;

[0104] Since different nodes use the same baseline model for incremental model training, these subcategory features have global consistency. At the same time, the subcategory incremental model parameters have good subcategory discrimination. Figure 4 It is the category feature expression capability of the incremental model parameters. Figure 4 Each node in corresponds to a parameter vector of an incremental model. The incremental model is trained by a sampling subset of a certain category of data. For ease of display, we use principal component analysis to reduce the dimension of the parameter vector and display it in two-dimensional space. As can be seen from the figure, the parameter vectors of the incremental models of different types of samples have good discrimination. The incremental models obtained from the same type of sample subsets are located in adjacent positions in the parameter space. However, these incremental model parameters are not suitable for direct use as local subcategory feature vectors. On the one hand, some incremental models with different local category numbers overlap in the model parameter space. For example, Figure 4 The two models of categories numbered 8 and B overlap in parameter space, which makes the data of these two categories inseparable.

[0105] To avoid the problem of overlapping feature spaces of some subcategories mentioned above, and to reduce the dimension of feature vectors and communication overhead, we will extract a representative subset of optimized parameters from the model parameters as the feature vectors of the subcategories. This solution selects the bias parameter of the last layer of the model as the final parameter. First, in the layers from the input layer to the output layer of a deep neural network, the later the layer, the higher its abstract ability and the stronger its category expression ability. Second, from the perspective of the back propagation algorithm, the parameters of the layers close to the output are adjusted first. Therefore, the parameters close to the output layer are most susceptible to the incremental training process and can best capture the feature information in the training data of the relevant category. In addition, in most deep learning models, the number of nodes in the layers close to the output layer is relatively small, so the number of parameters in these layers is also smaller. Taking the classic letNet-5 model as an example, the total number of parameters of the model exceeds 40,000, while the feature parameters selected according to our solution are 10. The category expression ability of the optimized parameter subset can be referred to the "Performance Evaluation" section.

[0106] (5) Globally consistent subcategory label assignment: The aggregation node collects subcategory feature vectors and local label information from different nodes; the aggregation node globally clusters these local subcategory feature vectors based on the collected subcategory feature vectors, and assigns different global labels to each global subcategory according to the global clustering results; the aggregation node assigns global labels to each subcategory in the global clustering results, and the aggregation node establishes a mapping scheme between local subcategory labels and global labels for the local subcategory of each node based on the collected local labels of the subcategory and the reallocated global labels, and feeds back the mapping relationship to the corresponding node; each node modifies the local category label of its own sample into a global category label according to the received mapping scheme. The whole process is as follows: Figure 8 shown.

[0107] Through the above operations, we have obtained many new types of sample sets. Next, we will use the new types of sample sets as training data to update the feature vector extraction model, the new type of encrypted network traffic packet detection model, the existing type of encrypted network traffic classification model and other models. The principles of updating the aforementioned models are basically similar. The following takes the update of the feature extraction model as an example to explain.

[0108] (6) Model update: Multiple network traffic monitoring nodes expand the model and use the samples collected by each node and assigned global labels as described in step (5) to train the expanded model in a collaborative manner until the model converges or reaches a pre-set error threshold. The specific steps are as follows:

[0109] (6.1) Model expansion: The model is expanded according to the total number of newly added categories and the total number of newly added samples. When the number of newly added categories and samples is small, the number of neurons in the output layer of the model is increased. When the number of newly added categories and samples is very large, the number of layers in the middle layer or the number of neurons in each layer needs to be increased.

[0110] When the model is updated, as new types of training samples are added, the total number of categories and the total number of samples increase. We need to appropriately expand the model so that the complexity of the model is adapted to the number of samples and the number of categories. Specifically, when the number of newly added categories and samples is much smaller than the number of existing types of samples and categories during previous training, we only need to modify the output layer of the model and add the same number of neurons as the new types. When the number of newly added categories and samples is very large, we need to expand the intermediate layers. The intermediate layers can be expanded by either increasing the number of levels or expanding the dimensions of the existing intermediate layers.

[0111] (6.2) Model initialization: The original basic model parameters in the extended model are initialized using the existing optimal feature parameters; the parameters of each neuron in the extended part of the model are initialized using random numbers;

[0112] (6.3) Optimization equation: The optimization equation is defined as where f' θ It is an expanded model;

[0113] (6.4) Model training: Multiple network traffic monitoring nodes use their own collected samples that have been assigned global labels to train the expanded model in a collaborative manner until the model converges or reaches a pre-set error threshold.

[0114] 4. Performance Evaluation

[0115] 1. Experimental setup

[0116] 1) Datasets used for evaluation: The datasets used in the experiment consist of two parts. One part comes from the ISCXVPNnonVPN dataset. This dataset includes different types of conventional encrypted traffic and protocol encapsulated traffic. Samples with controversial categories were deleted. The final dataset has 12 categories in total. These include 6 conventional encrypted traffic categories (i.e., Chat, Email, File, P2p, Streaming, Voip) and 6 protocol encapsulated traffic categories (i.e., Vpn_Chat, Vpn_Email, Vpn_File, Vpn_P2p, Vpn_Streaming, Vpn_Voip). If not otherwise specified, we will use the 10 numbers from 0 to 9 to identify the first 10 categories in the order listed above, and use the two letters A and B to identify the last two categories in order. However, the number of samples in some categories of this dataset is too small. For example, the total number of samples in the Vpn_Email category is only 253. And the distribution of samples between categories is extremely unbalanced. For example, the number of samples in the Chat category is 5257, which is much larger than the total number of samples in the Vpn_Email category. To this end, we have expanded the samples of each category in the dataset and ensured that the number of samples in each category of the dataset is basically equal.

[0117] 2) Platform and model: The deep learning framework used in the experiment is TensorFlow. The federated learning platform used in the experiment is TensorFlow Federated (TFF) framework. In the specific experiment, the federated learning related mechanisms are implemented in local mode, that is, the client node and the server are implemented in a virtual way, and they are actually located on the same device. Several models are mentioned in the proposed scheme, such as feature extraction scheme and new detection scheme. The main parts of these models used in the experiment are similar to LeNet-5, and they are modified in the following three aspects. First, all 2D modules are replaced by 1D modules. For example, the 2D convolution module of the convolution layer is replaced by a 1D convolution module. Second, an attention layer is added to the model. Third, the input layer and the output layer are adaptively modified according to the number of samples and traffic categories. The main part of the model consists of seven layers. It includes three convolutional layers, two pooling layers, and two fully connected layers. The convolutional layer and the pooling layer are implemented in one-dimensional form. The attention layer is inserted after the second convolutional layer in the form of a bypass, and the output of the attention layer is used to dynamically weight the output of the second convolutional layer.

[0118] 3) Evaluation indicators: The evaluation indicators used in the experiment include accuracy, precision, recall and F1 score.

[0119] 2. Model component selection and performance comparison.

[0120] The model adopted in the proposed scheme is built based on 1D-CNN and attention mechanism. Table 1 is the performance comparison between different model strategies, namely ours (model based on 1D-CNN and attention mechanism), 1D-CNN based model and 2D-CNN based model. According to the experimental results, the classification performance of the model based on 1D-CNN and attention mechanism is better than that of the individual 1D-CNN and 2D-CNN models. It should be pointed out that it is meaningless to compare the results of this experiment with those of other subsequent experiments. On the one hand, the number of samples in different categories is extremely unbalanced. The datasets of subsequent experiments are balanced in categories. On the other hand, the number of samples used for model training in this experiment is much larger than the number of samples in subsequent experiments.

[0121] Table 1 Performance comparison of different model strategies

[0122] plan Accuracy Precision Recall F1-score The present invention 0.949 0.957 0.938 0.947 1D-CNN 0.921 0.945 0.933 0.939 2D-CNN 0.845 0.852 0.846 0.848

[0123] 3. Performance of new type traffic packet detection model.

[0124] As can be seen from Figure 2, the output vectors of new traffic packets and known traffic packets have large feature differences in each dimension corresponding to the top3 vector elements. In order to more intuitively show the differences, we use the top3 vector elements of the output vector as three different dimensions to construct a three-dimensional space. We mark these samples in this three-dimensional space. Samples marked as 0 represent existing type samples, and samples marked as 1 represent new type samples. In order to observe the distribution characteristics of existing type samples and new type samples in space, we project the three-dimensional space graphics onto three different two-dimensional planes to facilitate viewing the effect. The results are shown in Figures 5(a), 5(b), and 5(c).

[0125] It can be seen from Figure 5(a), Figure 5(b) and Figure 5(c) that the vast majority of existing type samples and new type samples are distributed in different areas of the space, with relatively obvious distribution differences. By selecting appropriate threshold parameters, we can easily separate most of the new type samples. At the same time, there is a certain overlap between existing type samples and new type samples in some areas. For example, there are a large number of samples marked as 0 or 1 densely distributed in the lower right corner area of ​​Figure 5(a) and Figure 5(b), and the same situation also occurs in the lower left corner area of ​​Figure 5(c). Therefore, the threshold-based segmentation scheme will completely separate them from this area. We compared the performance of the new type sample detection scheme proposed in the present invention with the threshold-based segmentation scheme. Table 2 is the comparison results. During the experiment, the segmentation threshold of the first dimension was set to 0.9, and the segmentation thresholds of the other two dimensions were set to 0.1.

[0126] Table 2 Performance comparison of new sample detection schemes

[0127] plan Accuracy Precision Recall F1-score Threshold segmentation scheme 0.683 0.996 0.613 0.759 The present invention 0.942 0.986 0.906 0.944

[0128] As shown in Table 2, the accuracy of the threshold segmentation scheme is very high, but the recall and accuracy are relatively low. The accuracy of the threshold segmentation scheme is very high, mainly because the existing type samples are very concentrated in the top3 dimensions. When the threshold segmentation scheme is used, the existing type samples can be correctly identified with a high probability, and the probability of misidentifying the existing type samples as new type samples is low.

[0129] The distribution of new type samples is relatively scattered, and many new type samples overlap with existing type samples, and cannot be directly separated by simple threshold segmentation. Therefore, the ability of scheme L1 to identify new type samples is relatively poor, and its recall rate and accuracy are relatively low.

[0130] The recall rate (Recall) and precision (Accuracy) of the scheme of the present invention are greatly improved. However, its accuracy (Precision) is slightly reduced compared with the threshold segmentation scheme. The new type samples (labeled as 1) in the training samples of the scheme of the present invention come from the threshold segmentation scheme. Under the specified threshold, some existing type samples are mistakenly divided into new type samples by the threshold segmentation scheme. The training process of the scheme of the present invention will inevitably learn the pattern of such mislabeled samples, which will increase the probability of dividing existing type samples into new type samples. Therefore, its accuracy (Precision) is reduced to a certain extent compared with the threshold segmentation scheme.

[0131] 4. Performance of Globally Consistent Feature Extraction Scheme

[0132] Although the incremental model trained with different categories of data has a certain category expression ability, due to the large number of model parameters and the need for further improvement in expression ability, the incremental model parameters are not suitable for directly using as category feature data to upload to the server.

[0133] The last layer of parameters of the incremental model may have extremely strong expressive power. The following experiment verifies this. The experimental data comes from the VPNnonVPN dataset. The five categories numbered [1, 2, 3, 4, 5] are regarded as existing types, and the other 7 categories are regarded as new types. We construct a training sample set from the existing type dataset and train a basic model. Then, 56 training sample sets are randomly generated from the new type dataset, and each training sample set only contains samples of one of the 7 new types of data. Then, based on each dataset, incremental model training is performed based on the aforementioned basic model. Since the model converges very quickly when incremental training is performed on samples of a single category, in order to reduce the training cost, we set the epoch to 3 when performing incremental training on each type of sample. The training process of a single incremental model takes no more than 1s. At the end of the training, the training accuracy of most models reaches 1.

[0134] We extract two sets of parameters (kernel and bias) from the first and last layers of the model, and perform visual comparison after dimensionality reduction through PCA. The clustering effects of the two sets of parameters in the first layer in 2D space are shown in Figure 6(a) and Figure 6(b), respectively. The clustering effects of the two sets of parameters in the last layer in 2D space are shown in Figure 7(a) and Figure 7(b), respectively.

[0135] As shown in Figures 6 and 7, the clustering effect of the two sets of parameters in the first layer of the model is obviously worse than that of the two sets of parameters in the last layer of the model. In Figure 7, all categories do not overlap. In Figure 6, the distribution of nodes in each category is scattered and there is overlap. The parameter bias ( Figure 6a and Figure 7a ) has a significantly better clustering effect than the kernel parameter ( Figure 6b and Figure 6b ). Taking Figure 7 as an example, although the kernel and bias parameters of the last layer have good clustering effects, the bias parameter ( Figure 7a ) is obviously more concentrated.

[0136] like Figure 4 As shown in the figure, it is not a wise choice to directly use all model parameters as feature data. Not only are there too many feature vectors, but the category distinction effect is not ideal, and some categories cannot be distinguished. If the bias parameter of the last layer is used as feature data, the feature vector length is greatly reduced, and the category distinction effect is also significantly improved ( Figure 7a ).

[0137] 5. Performance comparison before and after model update

[0138] In the performance analysis experiment of adaptive updating global model, we designed 3 scenarios. All three scenarios include 9 known traffic categories. The number of unknown traffic categories in the three scenarios is 1, 2 and 3 respectively. The specific scenario description is shown in Table 3. In each experimental scenario, we randomly select 2000 samples from each existing type as training samples to train a basic classification model G1. Then we randomly select 2000 samples from each existing type and each new type as new traffic samples for the operation in the latter stage of the proposed scheme to obtain an updated classification model G2. The traffic samples used in the two stages are different. Finally, we analyze the performance of the two classification models before and after the update in different scenarios based on the experimental results. The experimental results are shown in Table 4. It can be seen from the results that in the three scenarios, the performance of the new model G2 is slightly lower than that of G1, and the more new types there are, the greater the performance degradation. This is because as the number of new types increases, the proportion of new type samples in the new sample set also increases. Due to the existence of errors in the identification and labeling of new types of samples, the greater the proportion of new types of samples, the greater the impact of the errors on the final results.

[0139] Table 3 Experimental scenario description

[0140] Scenario Existing Types New Types Scenario 1 [1,2,3,4,5,6,7,8,9] [A] Scenario 2 [2,3,4,5,6,7,8,9,A] [0,B] Scene 3 [0, 2, 3, 4, 6, 7, 8, A, B] [1,5,9]

[0141] Table 4 Performance analysis of adaptive update model

[0142] Scenario Model Accuracy Accuracy Recall F1-score 1 G1 0.948 0.949 0.948 0.948 1 G2 0.942 0.942 0.942 0.942 2 G1 0.952 0.953 0.952 0.952 2 G2 0.942 0.943 0.942 0.942 3 G1 0.968 0.969 0.968 0.968 3 G2 0.897 0.898 0.897 0.896

Claims

1. A method for mining and utilizing new types of encrypted network traffic packets in a distributed scenario, characterized in that: The following steps are involved: (1) Preparation stage: multiple network traffic monitoring nodes monitor the network traffic of different network areas they are responsible for respectively; each network traffic monitoring node independently collects a certain number of network traffic packet samples that have been labeled by category; multiple network traffic monitoring nodes cooperate with each other to train a new type of network traffic packet detection model; the new type of network traffic packet refers to a network traffic packet sample of this type that has not yet been labeled by category; (2) Detection of new types of traffic packets: Each node detects new types of network traffic packets from its newly received network traffic packets. The mining and utilization of new types of traffic packets are performed in a periodic manner, and each round of mining and utilization operations is based on all new types of traffic packets detected in the current period. (3) Subcategory discovery: Each node independently performs local clustering operations on the new type of traffic packets detected in this round; each node independently assigns labels to each subcategory sample in its clustering result; In the local clustering results, new type traffic packet samples of the same subcategory will be assigned the same local label; labels of different local subcategories are different; (4) Extraction of local subcategory feature vectors: Each node selects a globally unified benchmark; based on the globally unified benchmark, each node extracts a globally consistent category feature vector for each local subcategory; each node uploads the feature vectors of the local subcategory, together with their corresponding local subcategory labels, to the aggregation node; the specific steps are as follows: (4.1) Design of globally consistent benchmark model: The globally consistent benchmark model is defined as: y = f μ (x) = f e (f θ (x)) = argmax(softmax(f θ (x))); sub-model f θ Extract the feature model for encrypted network traffic packets; each node uses the globally optimal model parameter θ * Pair Model f θ Initialize the submodel f e It does not contain the parameters to be optimized and does not need to be initialized; (4.2) Training of subcategory incremental model: Each node independently trains an incremental model for its own local subcategory samples; the optimization equation for incremental training is: (4.3) Subcategory feature extraction based on the incremental model: Each node selects a parameter subset from each subcategory model parameter according to the same rule as the feature vector of the subcategory; each node uploads each local subcategory feature vector and local subcategory label to the aggregation node; (4.4) Globally consistent subcategory label assignment: The aggregation node globally clusters all local subcategories based on the collected subcategory feature vectors, and assigns different global labels to each global subcategory according to the global clustering results; The aggregation node establishes a mapping scheme between local subcategory labels and global labels for the local subcategories of each node based on the collected local labels of the subcategories and the reallocated global labels, and feeds back the mapping relationship to the corresponding nodes; each node modifies the local category labels of its own samples into global category labels based on the received mapping scheme. (5) Globally consistent category labeling: The sink node collects subcategory feature vectors and local label information from different nodes; the sink node performs global clustering based on all collected subcategory feature vectors; the sink node assigns global labels to each subcategory in the global clustering result; the sink node establishes a mapping scheme between local labels and global labels for each node, and returns the mapping scheme to the corresponding node; each node uses the received mapping scheme to assign a global label to each subcategory sample; (6) Model update: Multiple network traffic monitoring nodes expand the model and use the samples collected by each node and assigned global labels to train the expanded model in a collaborative manner until the model converges or reaches a pre-set error threshold. The specific steps are as follows: (6.1) Model expansion: The model is expanded according to the total number of newly added categories and the total number of newly added samples. When the number of newly added categories and samples is small, the number of neurons in the output layer of the model is increased. When the number of newly added categories and samples is very large, the number of layers in the middle layer or the number of neurons in each layer needs to be increased. (6.2) Model initialization: The original basic model parameters in the extended model are initialized using the existing optimal feature parameters; the parameters of each neuron in the extended part of the model are initialized using random numbers; (6.3) Optimization equation: The optimization equation is defined as where f′ θ It is an expanded model; (6.4) Model training: Multiple network traffic monitoring nodes use their own collected samples that have been assigned global labels to train the expanded model in a collaborative manner until the model converges or reaches a pre-set error threshold.

Citation Information

Patent Citations

  • Open set category mining and extending method based on depth neural network and device thereof

    CN107506799A

  • Network traffic identification method based on self-supervised convolution subspace clustering network

    CN114006870A