Flow classification model continuous learning method for dynamic network scene

By constructing the teacher-student model architecture and testing time domain adaptation technology in dynamic network scenarios, high-quality pseudo-labels are generated and empirical playback, the problems of data distribution offset and continuous learning adaptability of deep learning models in dynamic network scenarios are solved, and the efficient adaptation and accurate classification of the model are achieved.

CN120045999APending Publication Date: 2025-05-27UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510120559.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-25
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

In dynamic network scenarios, traditional deep learning models cause data distribution offset problems due to the assumption that the distribution of training and test data is consistent, resulting in a significant decline in the performance of the model in unknown environments, and it is difficult to continuously learn and adapt to changes in the network environment.

Method used

A continuous learning method of traffic classification model for dynamic network scenarios is proposed. By constructing a teacher-student model architecture, using test time domain adaptation technology, unsupervised learning, generating high-quality pseudo-labels, and consolidating the model's memory of past knowledge through an experience replay mechanism to avoid forgetting.

Benefits of technology

This method can reduce dependence on static data sets in dynamic network scenarios, enhance the adaptability and robustness of the model, ensure the accuracy and reliability of classification results, and avoid the risk of forgetting the model in the continuous learning process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045999A_ABST
    Figure CN120045999A_ABST
Patent Text Reader

Abstract

The invention discloses a traffic classification model continuous learning method for a dynamic network scene, and belongs to the field of network traffic classification. The method provided by the invention aims to provide a model continuous learning framework, so that the pre-trained network traffic classification model can still maintain relatively good performance in a complex and changeable network environment. According to the method provided by the invention, a teacher-student model architecture is constructed based on a test time domain adaptation technology, each data batch arriving step by step is learned, and the model can adapt to continuous change of data distribution, so that the dependence of the model on a static data set is reduced, and the adaptability of the model in a real scene is enhanced. Through the learning mechanism, the model can better cope with the evolution of a network flow mode in a real scene, and the accuracy and reliability of a classification result are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network traffic classification, and particularly to a continuous learning method for a traffic classification model oriented to dynamic network scenarios. Background Art

[0002] Network traffic classification refers to the analysis of data packets or data streams in a network and their classification into different categories. This process is not just a simple data division; it forms the basis of network management, optimization, and security protection and is a core part of network maintenance. Network traffic classification technology can effectively help network administrators such as operators understand the structure, usage, and potential risks of traffic in the network, thereby making more scientific and reasonable decisions, effectively optimizing network resource allocation, enhancing network protection capabilities, and promoting the development of intelligent network management.

[0003] Traditional network traffic analysis techniques mainly include port-based analysis and deep packet inspection, etc. With the continuous progress and rapid iteration of Internet technology, the limitations of traditional traffic analysis methods have become increasingly obvious due to the wide application of emerging technologies such as random ports, port spoofing, and encryption protocols (such as SSL / TLS). With the continuous development of deep learning, deep neural networks have achieved remarkable results in network traffic classification tasks. However, these results are only realized on the premise that the training data and test data meet the assumption of data distribution consistency in deep learning.

[0004] A dynamic network scenario refers to a network environment that is constantly changing, where network traffic characteristics, traffic patterns, etc. evolve over time and due to various factors. These changes can be influenced by multiple factors, such as the heterogeneity of the network environment caused by differences in infrastructure, network devices, communication protocols, bandwidth, etc., and the changes in traffic characteristics brought about by changes in user behavior, application version updates, and the application of emerging technologies such as new encryption methods and new transmission protocols. In such a dynamic network scenario, the premise of the traditional deep learning assumption that the training and test data distributions are consistent no longer holds, and the problem of data distribution shift frequently occurs, resulting in a significant decline in the performance of existing models when generalizing to data in unknown environments. Traditional training methods are difficult to effectively address these problems and usually require frequent retraining of the model, which not only consumes a large amount of time and computing resources but also requires a large amount of manpower to label the data. In addition, when retraining the model, for highly private data such as network traffic, the source domain data usually cannot be directly accessed, which further limits the applicability of traditional methods. For example, assume that a deep learning model is trained based on the traffic data of a certain region, but when the model needs to be deployed to another region, due to privacy protection policies or laws and regulations, the source domain data cannot be shared, which makes it impossible to implement methods such as direct transfer learning or joint training. Additionally, the change of the network environment is a dynamic and continuous process, and the model needs to be continuously updated and adapted to cope with this dynamic change. These challenging practical situations pose two key problems for network traffic classification research: 1) How to construct a target domain classifier as a proxy for the source domain without true labels and only with the source domain classifier available; 2) How to avoid forgetting the learned knowledge in a dynamic network scenario. Summary of the Invention

[0005] This application provides a continuous learning method for a traffic classification model for dynamic network scenarios, aiming to provide a model continuous learning framework so that a pre-trained network traffic classification model can still maintain good performance in a complex and changing network environment.

[0006] The technical solution adopted in this application is as follows:

[0007] A continuous learning method for a traffic classification model for dynamic network scenarios, the method comprising the following steps:

[0008] Step 1, obtain a network traffic data set with labeled category information as the source domain data set, and perform data preprocessing on the source domain data set to construct a feature vector of the network traffic data; wherein, the data preprocessing includes: data standardization, extraction of feature information, and construction of corresponding feature vectors;

[0009] Step 2: Build a convolutional neural network model in the training phase. Its input is the feature information of the network traffic data in the source domain dataset, and its output is the predicted classification of network traffic classification.

[0010] Step 3: Initialize the convolutional neural network model built in Step 2, and train the convolutional neural network model based on the network traffic dataset in the source domain dataset.

[0011] Step 4: Perform test-time domain adaptation on the convolutional neural network model trained in Step 3 based on the collected target dataset. Through the student model f s , the teacher model f t and the anchor model f a perform unsupervised learning to obtain a network traffic classifier for the network traffic data in the target domain based on the learned student model f s .

[0012] Among them, the initial models of the student model f s , the teacher model f t and the anchor model f a are all the convolutional neural network models trained in Step 3. The teacher model f t is used to provide pseudo-labels to guide the update of the model parameters of the student model f s . The model parameters of the anchor model f a are fixed and are used to restrict the update of the model parameters of the student model f s in a direction deviating from the source domain dataset.

[0013] Furthermore, the data preprocessing in Step 1 specifically includes:

[0014] For each packet file in the network traffic dataset, divide each data stream in the file according to the five-tuple; among them, the five-tuple refers to the source IP address, destination IP address, source port, destination port, and protocol.

[0015] For each data stream with the number of packets greater than or equal to K, extract the feature information of its first K packets. Among them, the feature information of each packet includes: packet size and packet arrival interval time; where K is an integer greater than or equal to 2.

[0016] Based on the feature information of the K packets of each data stream, construct its feature vector ((l 1 , t 1 ), (l 2 , t 2 ),..., (l K , t K ))), where l k represents the packet size of the k-th packet, and t kDenote the inter-arrival time of the k-th data packet, where the data packet number k of each data stream is k = 1, 2,..., K;

[0017] Normalize the feature vector of each data stream. Preferably, the min-max normalization method can be used.

[0018] Furthermore, in step 2, the convolutional neural network model successively includes: an input layer, a stacked structure alternating with convolutional layers and max-pooling layers, and a classification output layer. The classification output layer is used to output the prediction probabilities of each network traffic classification category, so as to determine the predicted classification of network traffic classification based on the maximum prediction probability (prediction confidence).

[0019] Furthermore, the classification output layer includes at least two fully connected layers, and the last fully connected layer is a fully connected layer with a Softmax function.

[0020] Furthermore, the stacked structure alternating with convolutional layers and max-pooling layers altogether includes three convolutional layers and three max-pooling layers, which are respectively set as:

[0021] The first convolutional layer uses Num one-dimensional convolutional kernels with a size of 3, the stride of the convolutional kernel is 1, and the ReLU function is used for activation;

[0022] The first max-pooling layer uses a 2*1 max-pooling operation to perform time-dimensional downsampling on the output of the first convolutional layer;

[0023] The second convolutional layer uses 2*Num one-dimensional convolutional kernels with a size of 3, the stride of the convolutional kernel is 1, and the ReLU function is used for activation. At the same time, a residual connection is introduced. After adjusting the dimension of the output of the first convolutional layer through a 1*1 convolutional layer, it is added to the output of the second convolutional layer to obtain the output of the second convolutional layer with a residual connection and send it to the second max-pooling layer;

[0024] The second max-pooling layer uses a 2*1 max-pooling operation to perform time-dimensional downsampling on its input;

[0025] The third convolutional layer uses 4*Num one-dimensional convolutional kernels with a size of 3, the stride of the convolutional kernel is 1, and the ReLU function is used for activation. At the same time, a residual connection is introduced. After adjusting the dimension of the output of the second convolutional layer with a residual connection through a 1*1 convolution, it is added to the output of the third convolutional layer to obtain the output of the third convolutional layer with a residual connection and send it to the third max-pooling layer;

[0026] The third max-pooling layer uses a 2*1 max-pooling operation to perform time-dimensional downsampling on its input.

[0027] Further, the classification output layer includes two fully connected layers. The number of neurons in the first fully connected layer is set to 2 * Num, and the ReLU function is used for activation. The number of neurons in the second fully connected layer is consistent with the number Cum of preset network traffic classification categories, which is used to map the output features of the first fully connected layer to the output dimensions of Cum categories. The Softmax of the second fully connected layer is used to perform Softmax normalization on this output dimension to obtain the probability distributions of Cum categories, that is, the prediction probabilities of Cum categories, so as to determine the predicted classification of network traffic classification based on the maximum prediction probability (i.e., prediction confidence).

[0028] Further, in step 3, when training the convolutional neural network model based on the network traffic dataset of the source domain dataset, the loss function adopts the cross-entropy loss function, and the training stops when the value of the loss function converges.

[0029] Further, step 4 includes:

[0030] Collect network traffic data in different network environments as the target dataset for testing time-domain adaptation, and the service categories corresponding to the target dataset are within the service categories corresponding to the network traffic classification categories of the source domain dataset;

[0031] Adopt the same data preprocessing method as the source domain dataset to preprocess the target dataset to obtain the feature vectors of each target data in the target dataset;

[0032] Based on the teacher model f t Obtain high-quality pseudo-labels:

[0033] For the current time step T, define the current batch of data in the target dataset input to the teacher model f t where N represents the number of samples in the current batch, and is the i-th sample in the current batch of data with the sample index i = 1,..., T;

[0034] Input the data into the teacher model f t , and obtain the prediction confidence of each data sample through forward inference

[0035] Add a dropout operation to the teacher model f t , and then based on the teacher model f t perform N forward inferences on each data sample in the current batch of data , and calculate the variance σ of the prediction confidences obtained from the N forward inferences T ;

[0036] If σ T is less than the preset threshold γ, and the prediction confidence of the data sample is greater than the threshold τ of the set time step T T , then is used as a high-quality pseudo label; where the value of the threshold γ is (0, 1);

[0037] Based on all high-quality pseudo labels and the corresponding data samples the student model f s is updated; and after each round of updating the student model f s , based on the model parameters s of the updated student model f the current model parameters t of the teacher model f (i.e., the model parameters updated at time step T - 1) are updated to obtain the model parameters of the updated teacher model f t : where μ is the preset update weight, and its value is (0, 1);

[0038] where the loss function L of the student model f s during model update includes: the cross-entropy loss L 1 between the student model and the high-quality pseudo label, and s the difference loss L a between the model parameters of the student model f 2 and the anchor model f

[0039] In practical applications, based on the student model, the traffic classification result of the target data to be classified currently collected is obtained.

[0040] Furthermore, the loss function L of the student model f s during training also includes the cross-entropy loss L 3 corresponding to the selected representative data samples.

[0041] Furthermore, the setting method of the threshold τ of time step T T is as follows:

[0042] At time step T, define the batch data t input to the teacher model f where N represents the number of samples in the current batch;

[0043] At time step T, calculate the prediction distribution of the target data set to obtain the global confidence mean:

[0044]

[0045] Among them, is the prediction probability distribution of the teacher model f t for the data sample at the i'-th time step, indicating the prediction confidence of the i-th data sample at the i'-th time step;

[0046] According to the formula calculate the global confidence variance

[0047] Calculate the threshold where the control parameter k is the current training round, and K is the total number of rounds;

[0048] Calculate the global threshold at time step T: where α is a preset parameter less than 1;

[0049] Calculate the local confidence mean at time step T:

[0050] Calculate the local confidence variance:

[0051] Calculate the local threshold at time step T

[0052] Based on the weighted fusion result of the global threshold and the local threshold at time step T, obtain the threshold τ T .

[0053] Preferably, the threshold τ T can be set as: where the value range of the parameter β is (0, 1).

[0054] Furthermore, the specifically selected representative data samples are:

[0055] Define M as the storage area for storing representative data samples, and its capacity is defined as K;

[0056] In the storage area M, maintain two attribute quantities A and U for each data sample. Among them, A is the time when the data sample exists in M, which increases with the number of time steps. U is the confidence of the teacher model f t in predicting it, that is, the classification category probability value predicted by the model;

[0057] The update rule of M is set as follows: for each data sample reached (i.e., in each epoch, when taking out one data sample from its batch data), if M is not full, the current data sample is stored in M; otherwise, calculate the screening score of the current data sample and the screening scores of all data samples in M. If the maximum value among the screening scores of all data samples in M is greater than the screening score of the currently reached data sample, then replace the data sample with the highest screening score in M with the currently reached data sample and store the currently reached data sample in M.

[0058] Among them, the calculation method of the screening score of the data sample is:

[0059]

[0060] Among them, A(x) and U(x) are respectively the attribute A and U of the data sample x in M, the initial value of A(x) is 0, sim(x, y) represents the cosine similarity between the data sample x and y, and δ 1 、δ 2 and δ 3 are preset parameters (i.e., hyperparameters), and satisfy δ 2 = 3δ 1 、δ 3 = 2δ 1 . Further, the values of δ 1 、δ 2 and δ 3 are all in the range of (0, 1).

[0061] The technical solution provided by this application at least brings the following beneficial effects:

[0062] The method proposed in this application is a continuous learning method for a network traffic classification model based on test-time domain adaptation. Based on the test-time domain adaptation technology, a teacher-student model architecture is constructed. By learning each data batch that arrives step by step, the model can adapt to the continuous change of the data distribution, thereby reducing the model's dependence on static data sets and enhancing its adaptability in real-world scenarios. Through this learning mechanism, the model can better cope with the evolution of network traffic patterns in real-world scenarios and ensure the accuracy and reliability of classification results. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] The above and / or additional aspects and advantages of this application will become obvious and easy to understand from the following description of the embodiments in conjunction with the drawings, where:

[0064] Figure 1This is a schematic diagram of the processing procedure of the continuous learning method for the traffic classification model facing dynamic network scenarios in the embodiments of this application. Among them, EMA refers to Exponential Moving Average, Rs refers to the output of the student model, Rt refers to the output of the teacher model, prob is the prediction confidence of the teacher model output, and R1 to Rn refer to the prediction confidence of each of the n repeated forward inferences after adding the Dropout operation to the teacher model. Detailed implementation manners

[0065] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this application will be described in detail and completely below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the embodiments described by referring to the accompanying drawings are exemplary and are intended to explain this application, rather than being construed as a limitation on this application.

[0066] The embodiments of this application provide a continuous learning method for the traffic classification model facing dynamic network scenarios. First, in order to construct a classifier on the target domain using the source domain classifier without labeled target domain data, existing solutions usually adopt methods based on pseudo-label generation and transfer learning. These methods generate pseudo-labels for the target domain using the classifier trained on the source domain, and then jointly train the model with the target domain data and pseudo-labels. However, this method faces the problem that pseudo-labels are prone to be noisy. Especially when the data distribution of the target domain is quite different from that of the source domain, incorrect pseudo-labels may lead to noise accumulation and affect the performance of the classifier. This application effectively screens out high-quality pseudo-labels by combining the uncertainty estimation of the model and the global and local confidence analysis, avoiding the noise accumulation problem caused by relying on incorrect pseudo-labels, thereby ensuring the accurate learning of the target domain representation. Second, existing methods usually focus on the optimization of the target domain and lack an effective mechanism to retain and consolidate the source domain knowledge. In order to avoid the model forgetting the knowledge learned in the past during the continuous learning process, this application proposes an experience replay mechanism for representative data samples considering timeliness, uncertainty, and diversity, effectively consolidating the model's memory of important features in the past and reducing the risk of forgetting during the continuous learning process. Finally, this application also introduces a parameter soft alignment mechanism. By softly aligning the target model parameters with the initial pre-trained model parameters, it avoids the model from overly updating in the direction deviating from the source domain when adapting to new data, thus achieving a balance between learning new knowledge and maintaining old knowledge. Based on the above design, this application can flexibly cope with the continuous change of data distribution in dynamic network scenarios and significantly improve the accuracy and robustness of the network traffic classification model.

[0067] In one embodiment, the specific implementation steps of the continuous learning method for the traffic classification model facing dynamic network scenarios provided by the embodiments of this application include:

[0068] Step 1: Obtain a network traffic dataset with labeled category information, preprocess the network traffic dataset, and divide the dataset (training set, validation set, and test set). Among them, the preprocessing includes data standardization, feature extraction, and construction of feature vectors for the convolutional neural network;

[0069] Step 2: Build a convolutional neural network model in the training stage;

[0070] Step 3: Initialize the parameters of the convolutional neural network model built in Step 2; input the divided dataset (training set, validation set) into the convolutional neural network for training;

[0071] Step 4: Perform test-time domain adaptation on the convolutional neural network trained in Step 3, that is, perform test-time domain adaptation based on the target dataset.

[0072] Further, Step 1 includes the following steps:

[0073] Step 1-1: Obtain the ISCX-VPN-NonVPN dataset, select an unprocessed pcap file, and divide each data stream in the file according to the five-tuple (source IP address, destination IP address, source port, destination port, protocol);

[0074] Step 1-2: Based on the data streams divided in Step 1-1, use the dpkt toolkit to extract the feature information of the first K (such as 20) packets of each data stream. Among them, the feature information refers to the packet size and the packet arrival interval time. For data streams with less than 20 packets, they are skipped;

[0075] The packet size refers to the total number of bytes of the effective payload of the network packet;

[0076] The packet arrival interval time refers to the difference between the timestamps of two consecutive packets;

[0077] Step 1-3: Based on the feature information, construct a feature vector ((l 1 , t 1 ), (l 2 , t 2 ),..., (l 20 , t 20 ))), where l k represents the packet size of the kth packet, and t k represents the difference between the timestamps of the kth packet and the (k - 1)th packet;

[0078] Step 1-4: Perform normalization processing on the feature vector. Among them, the normalization method uses the maximum-minimum normalization method to make the feature values evenly distributed between (0, 1). If there are still data streams not processed, return to Step 1-2 and execute sequentially. Otherwise, execute Step 1-5;

[0079] Steps 1-5: Prepare two files respectively for storing data and the corresponding labels. The data is sourced from the feature vectors processed in Steps 1-2 to 1-4, and the labels are the service types corresponding to the pcap files processed in this round. If all pcap files in the ISCX-VPN-NonVPN dataset have been processed, divide the data and label files into a training set, a validation set, and a test set in a ratio of 8:1:1 correspondingly. Otherwise, return to Step 1-1 and execute sequentially.

[0080] In one embodiment, in Step 2, the hierarchical structure of the convolutional neural network model is as follows:

[0081] Input layer: Receive input features with a size of (batchsize, 2, 20);

[0082] First convolutional layer: Consist of 64 one-dimensional convolutional kernels of size 3, with a stride of 1 for the convolutional kernels, and use the ReLU function for activation;

[0083] First max-pooling layer: Apply a 2*1 max-pooling operation to the output of the first convolutional layer for temporal dimension downsampling;

[0084] Second convolutional layer: Consist of 128 one-dimensional convolutional kernels of size 3, with a stride of 1 for the convolutional kernels, and use the ReLU function for activation. At the same time, introduce a residual connection. After adjusting the dimension of the output of the first convolutional layer through a 1*1 convolution, add it to the output of this layer;

[0085] Second max-pooling layer: Apply a 2*1 max-pooling operation to the output of the second convolutional layer for temporal dimension downsampling;

[0086] Third convolutional layer: Consist of 256 one-dimensional convolutional kernels of size 3, with a stride of 1 for the convolutional kernels, and use the ReLU function for activation. At the same time, introduce a residual connection. After adjusting the dimension of the output of the second convolutional layer through a 1*1 convolution, add it to the output of this layer;

[0087] Third max-pooling layer: Apply a 2*1 max-pooling operation to the output of the third convolutional layer to complete the temporal dimension compression of the feature vectors;

[0088] Fully connected layer: Flatten the output of the third max-pooling layer and input it into a fully connected layer with 128 units, and use the ReLU function for activation;

[0089] Output layer: Consist of a fully connected layer that maps high-order features to the output dimension of 5 categories;

[0090] Softmax layer: Normalize the result of the output layer through Softmax to obtain the probability distribution of 5 categories:

[0091] P 类别 = softmax(z),

[0092] In one embodiment, step 3 includes the following steps:

[0093] Step 3-1, set the learning rate to 10 -4 , the batch size is 256, and the total number of training iterations is 50. The optimizer uses Adam to accelerate gradient convergence, and set the smoothing factors to β 1 = 0.9, β 2 = 0.999, and the weight decay coefficient is 1*10 -5 ,, for regularization to prevent overfitting. The loss function uses the cross-entropy loss function. During training, the learning rate is dynamically adjusted. Through the StepLR learning rate scheduler, the learning rate is decayed to 0.1 of the original every 10 epochs.

[0094] Step 3-2, input the preprocessed ISCX-VPN-NonVPN dataset into the convolutional neural network for training. The model is trained using the standardized data to optimize the input feature distribution. At the same time, the performance of the validation set is evaluated after each epoch to adjust the model hyperparameters. During the training process, the value of the loss function gradually decreases. When its value decreases to the set threshold and shows a stable trend, it is determined that the model has converged. At this time, it is used as the pre-trained model required for step 4.

[0095] In one embodiment, step 4 includes the following steps:

[0096] Step 4-1, initialize the student model f from the model obtained in step 3 s , the teacher model f t and the anchor model f a . The goal of step 4 is to obtain a student model that can, while inferring data in the actual network environment, synchronously update its own parameters, so that it can gradually adapt to the complex and changeable network environment in the real scenario. This is an unsupervised process. In this process, the role of the teacher model f t is to provide pseudo-labels to guide the update of the student model parameters. The parameters of the anchor model f a remain fixed during this process. Its role is to restrict the update of the student model parameters in the direction deviating from the original data domain, so as to retain the discriminative ability of the student model in the original data domain;

[0097] Step 4-2: Use Wireshark software to collect network traffic data in different network environments as the target data for testing time-domain adaptation. The data service categories collected are within the scope of the service categories in the ISCX-VPN-NonVPN dataset, that is, the service categories of the target dataset belong to the service category set in the training stage. Filters can be set in Wireshark software to collect data streams of specified services;

[0098] Step 4-3: Preprocess the collected data in the same way as in Step 1 to form the data type that the pre-trained model in Step 3 can receive;

[0099] Step 4-4: Dynamic threshold setting. At time step T, the mini-batch data input to the model

[0100] Step 4-4-1: Global information collection: At time step T, calculate the predicted distribution of the entire target dataset before, and obtain the global confidence mean:

[0101]

[0102] where, is the predicted probability distribution of the teacher model for the data sample at the i'-th time step, represents the predicted classification of the i'-th data sample at the i'-th time step. For example, for a 5-class traffic classification task, the predicted probability distribution of each can be described as (0.1, 0.2, 0.3, 0.9, 0.1). Based on the maximum predicted probability (0.9), its predicted classification label is the 4th class, and the predicted confidence of the current data sample is 0.9.

[0103] Calculate the global confidence variance to evaluate the stability of the model's prediction for the overall data before time step T:

[0104]

[0105] Calculate the threshold where λ is used to control the influence of variance on the threshold, and is obtained by the following formula:

[0106]

[0107] where, k is the current training epoch, and K is the total number of epochs.

[0108] Finally, obtain the global threshold at time step T through the following formula:

[0109]

[0110] Among them, the hyperparameter α ∈ (0, 1), and in the embodiments of the present application, it is set to α = 0.1.

[0111] Step 4-4-2, Local information collection:

[0112] For the mini-batch data B at time step T, calculate the predicted distribution of B to obtain the local confidence mean:

[0113]

[0114] Calculate the local confidence variance to evaluate the stability of the model prediction at time step T:

[0115]

[0116] Finally, calculate the local threshold at time step T

[0117] Step 4-4-3, Calculate the dynamic threshold. After obtaining the global threshold and the dynamic threshold, the threshold at time step T can be obtained:

[0118]

[0119] Among them, β is a hyperparameter, and its value range is (0, 1). In the embodiments of the present application, β = 0.3 is set;

[0120] Step 4-5, Obtain high-quality pseudo-labels: At time step T, evaluate the stability of the teacher model's prediction at this time. Add the dropout operation to the teacher model and let it perform N inferences on the data samples of the current batch Since the dropout operation randomly discards some neurons, the results of these N inferences will be different. Calculate the variance σ of the results of these N inferences T , if σ T < γ and then is provided to the student model for learning as a pseudo-label, otherwise skip the data sample Among them, γ is a hyperparameter, and its value range is (0, 1);

[0121] Step 4-6, Update the student model: After obtaining high-quality pseudo-labels from Step 4-4 to Step 4-5, use the student model to perform inferences on the data of the current batch, and at the same time calculate the cross-entropy loss between the student model and the pseudo-labels:

[0122]

[0123] Steps 4-7, Anti-forgetting Training: Since the distribution of the target data samples in each batch is continuously changing, if the update of the student model parameters is not restricted, the student model will update in the direction deviating from the original data domain, resulting in a decline in the discrimination ability of the student model in the original data domain. To solve this problem, the embodiments of the present application align the parameters of the student model and the anchor model softly, thereby restricting the student model from updating in the direction deviating from the original data domain. This method is implemented through the following loss function:

[0124] L 2 =∑||θ s -θ a || 2

[0125] where θ s and θ a are the parameter vectors of the student model and the anchor model respectively.

[0126] During the process of testing time-domain adaptation, the student model needs to continuously adapt to consecutive multiple batches of target data. However, due to the continuous change of the distribution of the target data samples in each batch, the student model will inevitably forget the knowledge learned from the previous batch of data when adapting to the new batch of data. To solve this problem, the embodiments of the present application maintain a set of high-quality representative data samples. By periodically performing replay training on the model, while adapting to the new distribution, it can effectively retain the important knowledge learned previously, thereby achieving the dynamic integration and stable transfer of knowledge and avoiding forgetting.

[0127] Among them, the high-quality representative data samples are specifically:

[0128] Define M as the storage area for storing the above-mentioned representative data samples, with a capacity of K. For each data sample in M, two attributes A and U are maintained. Among them, A is the time when the data sample exists in M, which increases with the time step, and U is the confidence level predicted by the teacher when the sample exists in M. M is updated through the following rules: For each arriving sample, if M is not full, the sample is put into M. Otherwise, calculate the score of the sample and the scores of all samples in M. If the maximum value of the scores of all samples in M is greater than the score of the newly arriving sample, replace the sample with the highest score in M and store the newly arriving sample in M. The above-mentioned score is obtained by the following formula:

[0129]

[0130] where A(x) and U(x) are the attributes A and U of the sample x in M respectively, the initial value of A(x) is 0, sim(x, y) represents the cosine similarity between the sample x and y, δ 1 、δ 2 and δ 3is a hyperparameter, i.e., the weight coefficient of the corresponding term. In the embodiments of the present application, δ 1 is set to 0.2, δ 2 is set to 0.6, and δ 3 is set to 0.4.

[0131] With such a set of representative data samples, experience replay can be performed each time the student model is updated, thereby effectively retaining the key information in the previous learning process and alleviating the catastrophic forgetting phenomenon that may occur when the model adapts to the new data distribution. This method can help the student model balance the learning of new knowledge and the retention of old knowledge under the dynamically changing target data distribution, and improve the long-term adaptability and stability of the model. Among them, the experience replay of the student model is performed through the following loss function:

[0132]

[0133] where M represents the set of representative data samples, x represents the target data, represents the predicted confidence of the target data x output by the student model f s (i.e., the class prediction probability corresponding to the predicted classification output by the model), represents the predicted confidence of the target data x output by the teacher model f t .

[0134] So far, the total loss function can be obtained:

[0135] L = η 1 L 1 + η 2 L 2 + η 3 L 3 ,

[0136] where η 1 = η 2 = η 3 = 1.0, and it can also be set to other values based on the application scenario requirements.

[0137] Step 4-8, Teacher model update. After the student model is updated through Steps 4-1 to 4-7, the teacher model is updated by the following formula:

[0138]

[0139] where and are the parameters of the teacher model and the student model at time step T respectively. The value range of the parameter μ is (0, 1). In this embodiment, μ is set to 0.1.

[0140] In actual application, based on the student model, obtain the traffic classification result of the target data to be classified by traffic currently collected.

[0141] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the application, rather than limiting them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A continuous learning method for traffic classification model for dynamic network scenarios, characterized in that: The following steps are involved: Step 1, obtaining a network traffic data set with labeled category information as a source domain data set, and performing data preprocessing on the source domain data set to construct a feature vector of the network traffic data; wherein the data preprocessing includes: data standardization, extracting feature information, and constructing a corresponding feature vector; Step 2: Build a convolutional neural network model in the training phase, whose input is the feature information of the network traffic data of the source domain data set, and the output is the predicted classification of the network traffic classification; Step 3: Initialize the convolutional neural network model built in step 2, and train the convolutional neural network model based on the network traffic dataset of the source domain dataset; Step 4: Based on the collected target data set, the convolutional neural network model trained in step 3 is adapted to the test time domain, and the student model f s , teacher model f t and anchor model f a Perform unsupervised learning based on the learned student model f s obtaining a network traffic classifier for network traffic data of a target domain; Among them, the student model f s , teacher model f t and anchor model f a The initial models are all convolutional neural network models trained in step 3, and the teacher model f t Used to provide pseudo labels to guide the student model f s Update of model parameters, anchor model f a The model parameters are fixed and used to limit the student model f s The model parameters are updated in the direction that deviates from the source domain dataset.

2. The method according to claim 1, characterized in that The data preprocessing in step 1 specifically includes: For each data packet file in the network traffic data set, each data flow in the file is divided into five tuples; the five tuples refer to the source IP address, destination IP address, source port, destination port and protocol; For each data stream with a number of packets greater than or equal to K, extract characteristic information of the first K packets, wherein the characteristic information of each packet includes: packet size and packet arrival interval; wherein K is an integer greater than or equal to 2; Based on the feature information of K packets of each data stream, construct its feature vector ((l1,t1),(l2,t2),…,(l K ,t K )), where l k represents the packet size of the kth packet, t k represents the inter-arrival time of the kth data packet, and the data packet number of each data stream is k = 1, 2, ..., K; The feature vector of each data stream is normalized.

3. The method according to claim 1, characterized in that In step 2, the convolutional neural network model includes in sequence: an input layer, a stacked structure of alternating convolutional layers and maximum pooling layers, and a classification output layer. The classification output layer is used to output the prediction probability of each network traffic classification category to determine the prediction category of the network traffic classification based on the maximum prediction probability therein, and the maximum prediction probability is the corresponding prediction confidence.

4. The method according to claim 3, characterized in that The stacked structure of alternating convolutional layers and maximum pooling layers includes three convolutional layers and three maximum pooling layers, which are set as follows: The first convolutional layer uses Num one-dimensional convolution kernels of size 3, the stride of the convolution kernel is 1, and the ReLU function is used for activation; The first maximum pooling layer uses a 2*1 maximum pooling operation to downsample the output of the first convolutional layer in the time dimension; The second convolutional layer uses 2*Num one-dimensional convolution kernels of size 3, with a step size of 1, and uses the ReLU function for activation. At the same time, a residual connection is introduced. The output of the first convolutional layer is adjusted in dimension through a 1*1 convolutional layer, and then added to the output of the second convolutional layer to obtain the output of the second convolutional layer with residual connection and send it to the second maximum pooling layer; The second maximum pooling layer uses a 2*1 maximum pooling operation to downsample its input in the time dimension; The third convolutional layer uses 4*Num one-dimensional convolution kernels of size 3, the stride of the convolution kernel is 1, and the ReLU function is used for activation. At the same time, the residual connection is introduced. The output with residual connection of the second convolutional layer is adjusted by 1*1 convolution, and then added to the output of the third convolutional layer to obtain the output with residual connection of the third convolutional layer and send it to the third maximum pooling layer; The third maximum pooling layer uses a 2*1 maximum pooling operation to downsample its input in the time dimension.

5. The method according to claim 3, characterized in that The classification output layer includes two fully connected layers. The neurons of the first fully connected layer are set to 2*Num and activated by the ReLU function. The neurons of the second fully connected layer are consistent with the preset number of network traffic classification categories Cum, which is used to map the output features of the first fully connected layer to the output dimension of Cum categories. The Softmax of the second fully connected layer is used to perform Softmax normalization on the output dimension to obtain the predicted probability of Cum categories, so as to determine the predicted category of the network traffic classification based on the maximum predicted probability.

6. The method according to claim 1, characterized in that Step 4 includes: Collect network traffic data under different network environments as the target data set for testing time domain adaptation, and the business category corresponding to the target data set is within the business category range corresponding to the network traffic classification category of the source domain data set; The target data set is preprocessed using the same data preprocessing method as the source domain data set to obtain the feature vector of each target data in the target data set; Based on the teacher model f t Get high-quality pseudo-labels: For the current time step T, define the input to the teacher model f t The current batch of data of the target dataset Where N represents the number of samples in the current batch. For the current batch of data The i-th sample in , sample index i=1,…,T; The data Input teacher model f t , the prediction confidence of each data sample is obtained through forward reasoning In the teacher model t Add dropout operation, and then based on the teacher model f t For the current batch of data Each data sample Perform N forward inferences and calculate the variance σ of the prediction confidence obtained by N forward inferences T ; If σ T is less than the preset threshold γ, and the data sample The prediction confidence Greater than the threshold τ of the set time step T T , then As a high-quality pseudo label; where the threshold γ is (0,1); Based on all high-quality pseudo labels and corresponding data samples For student models s Update the model; and each round of the student model f s After the update is completed, all are based on the updated student model f s Model parameters For the teacher model f t Current model parameters (i.e., the updated model parameters at time step T-1) are updated to obtain the updated teacher model f t Model parameters: Among them, μ is the preset update weight, and its value is (0,1); Among them, the student model f s The loss function L during model update includes: the cross entropy loss L1 between the student model and the high-quality pseudo-label, and the student model f s With anchor model f a The difference loss between the model parameters is L2.

7. The method according to claim 6, characterized in that Student Model s The loss function L during training also includes the cross entropy loss L3 corresponding to the selected representative data samples.

8. The method according to claim 7, characterized in that The cross entropy loss L3 is specifically: Among them, M represents the representative data sample set, x represents the target data, and f s T (x) represents the student model f s The prediction confidence of the output target data x, f t T (x) represents the teacher model f t The prediction confidence of the output target data x.

9. The method according to claim 6, characterized in that The threshold τ at time step T T The setting method is: At time step T, define the input to the teacher model f t Batch data Where N represents the number of samples in the current batch; At time step T, calculate the predicted distribution of the target data set and obtain the global confidence mean: in, is the teacher model f t The predicted probability distribution of the data sample at the i′th time step, represents the prediction confidence of the i-th data sample at the i′th time step; According to the formula Calculate the global confidence variance Calculating the threshold The control parameters k is the current training round number, K is the total round number; Compute the global threshold at time step T: Among them, α is a preset parameter less than 1; Compute the local confidence mean at time step T: Calculate the local confidence variance: Compute the local threshold at time step T Global threshold based on time step T and local threshold The weighted fusion result of T .

10. The method according to claim 9, characterized in that The representative data samples selected are as follows: Define M as a storage area for storing representative data samples, and its capacity is defined as K; In the storage area M, two attributes A and U are maintained for each data sample, where A is the time the data sample exists in M, and U is the teacher model f when the data sample exists in M. t The confidence level in the prediction of the data evolution, i.e. the probability value of the classification category predicted by the model; The update rule of M is set as follows: every time a data sample arrives, if M is not full, the current data sample is stored in M; otherwise, the screening score of the current data sample and the screening scores of all data samples in M ​​are calculated. If the maximum value of the screening scores of all data samples in M ​​is greater than the screening score of the currently arriving data sample, the data sample with the highest screening score in M ​​is swapped out, and the currently arriving data sample is stored in M; The calculation method of the screening score of the data sample is: Among them, A(x) and U(x) are the attributes A and U of the data sample x in M ​​respectively, the initial value of A(x) is 0, sim(x,y) represents the cosine similarity between data samples x and y, δ1, δ2 and δ3 are preset parameters, and satisfy δ2=3δ1, δ3=2δ1.