Noisy network flow representation method and network flow identification method and system based on double representation networks
Through the noisy network traffic representation method based on the dual-characterized network, the network traffic data is characterized from the two-dimensional label and feature dimensions, and the limitations of processing noise-containing samples in the prior art are solved, and the efficient identification effect is achieved in a high-noise environment.
Patent Information
- Application Number
- CN202510209748.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-30
AI Technical Summary
Existing noisy learning methods for network traffic mostly process noisy samples from a single dimension, which has limitations and cannot completely eliminate noise labels and make full use of the correct information in the original label.
The noisy network traffic representation method based on dual-characterized network is adopted. The network traffic data characteristic representation is mined from the two-dimensional label characteristics and traffic characteristics through label characteristics, and optimization functions such as KL divergence, cross entropy and self-cross entropy are used to enhance noise immunity.
Through the noisy network traffic representation method of the dual-characterized network, the effectiveness of deep neural networks when processing noisy data sets is improved, and can maintain excellent recognition accuracy in scenarios with high label noise, and even maintain more than 85% recognition accuracy under extreme conditions with a training data noise rate of up to 80%.
Smart Images

Figure CN120074908A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security technology, and particularly to a method for representing noisy network traffic based on a dual-representation network, a method for identifying network traffic, and a system therefor. Background Art
[0002] For a long time, traffic identification technology, as a core component of a Network Intrusion Detection System (NIDS), has played a key role in identifying malicious attacks and has been widely studied. With the advent of the era of full encryption and the rapid development of deep learning, the encrypted traffic identification technology based on deep learning has become a research hotspot in the field of traffic identification because it can automatically mine the deep features of encrypted traffic without decrypting data packets and constructing a feature set, effectively avoiding the risk of information leakage and saving resources.
[0003] Currently, most deep learning methods with good generalization are supervised learning, which requires training data to be pre-labeled. However, whether it is manual labeling or automatic labeling, it is easy to introduce label noise. In this context, Learning with Noisy Labels (LNL) emerged, aiming to reduce the impact of label noise on the model's generalization ability and thus enhance the robustness of the model in a noisy label environment.
[0004] Existing noisy learning methods mostly transform noisy samples from a single label or feature dimension to reduce the interference of noise on model training. DLDL, PRNCI, and GLLE train the recognition model by representing labels as probability distributions. Although they reduce the impact of labeling errors to a certain extent, they cannot completely eliminate the noise in the label distribution. There are also explorations of noise-resistant representations of traffic features in unsupervised scenarios, but they ignore a large amount of correct label information and its reflected classification features in the noise dataset. Although these methods reduce the impact of noise by enhancing labels or features, they each have their own defects: label enhancement cannot completely eliminate noisy labels, and feature enhancement cannot fully utilize the correct information in the original labels. Summary of the Invention
[0005] Therefore, the present invention provides a method and a system for identifying noisy network traffic based on a dual-representation network, to solve the problem of limitations existing in the existing noisy learning of network traffic when processing noisy samples from a single dimension.
[0006] According to the design scheme provided by the present invention, on the one hand, a method for representing noisy network traffic based on a dual-representation network is provided, including:
[0007] Obtain network traffic sample data, where the network traffic sample data is network traffic data with original sample labels, and the original sample labels contain noise labels;
[0008] Use the network traffic sample data to train a dual-representation network, enabling the dual-representation network to mine the feature representations of network traffic data from the dual dimensions of label features and traffic features. The dual-representation network includes a label representation network for learning the label distribution in the original sample labels and outputting label representations, and a feature representation network for combining label representations to mine the noise-resistant traffic features in the original network traffic features and output feature representations.
[0009] As the noisy network traffic representation method based on the dual-representation network of the present invention, further, the label representation network includes: a fully connected layer for updating the sample label distribution, and a softmax layer for ensuring that the label distribution conforms to the probability distribution vector distribution characteristics.
[0010] As the noisy network traffic representation method based on the dual-representation network of the present invention, further, the feature representation network adopts a multi-layer fully connected autoencoder to perform noise-resistant feature representation mapping on the features of the samples in the feature representation space.
[0011] As the noisy network traffic representation method based on the dual-representation network of the present invention, further, training the dual-representation network using the network traffic sample data includes:
[0012] Adopt the KL divergence function to optimize the classification of the label distribution, use cross-entropy to quantify the similarity between the original label vector and the label distribution vector, and use the self-cross-entropy function to keep the label distribution with a single peak, so as to set the optimization function of the label representation network based on the KL divergence function, cross-entropy, and self-cross-entropy;
[0013] Set the feature representation loss by minimizing the information entropy difference before and after representation in the information layer, and set the feature encoding loss based on the information entropy difference between the decoded output and the encoded input; in the classification layer, fuse the label representation to transform the feature distribution of network traffic data into a spatial distribution, and strengthen the classification boundary between each feature category by increasing the difference in the distribution of similar and dissimilar features in the feature space, so as to set the optimization function of the feature representation network through the constraints of the information layer and the classification layer;
[0014] For the original sample labels, based on the optimization function of the label representation network and through an iterative process, enable the label representation network to learn the label distribution of the network traffic data and output the label distribution as the label representation;
[0015] For the network traffic data, based on the optimization function of the feature representation network and through an iterative process, combine the label representation to perform noise-resistant feature representation on the network traffic data features and output the traffic feature representation.
[0016] As the method for representing noisy network traffic based on the dual representation network of the present invention, further, the optimization function of the label representation network is expressed as: I label (x) = α 1 ·I class (x) + α 2 ·I one (x) + α 3 ·I label (x), where α 1 , α 2 and α 3 are hyperparameters, x represents the input of the label representation network, I label (x) represents the optimization function of the label representation network, I class (x) represents the label distribution classification optimization function set based on the KL divergence function, I clean (x) represents the optimization function for quantifying the similarity between the original label and the label distribution and performing label local purity, I one (x) represents the single-peak optimization function for keeping the label distribution with a single peak based on the self-cross entropy.
[0017] As the method for representing noisy network traffic based on the dual representation network of the present invention, further, the optimization function of the feature representation network is expressed as: where β 1 and β 2 are hyperparameters, x is the input of the feature representation network, represents the optimization function of the feature representation network, represents the information layer constraint, represents the classification layer constraint, and represents the representation loss constraint function, represents the coding loss constraint function, ω represents the relative magnitude of the classification layer constraint, and Normal() is the normalization function, d inside represents the average distance within the traffic feature cluster, d outside represents the average distance outside the traffic feature cluster.
[0018] On the other hand, the present invention also provides a network traffic recognition method, including:
[0019] Training a network traffic recognition model by using the network traffic data feature representation obtained by the above method to obtain a network traffic recognition target model;
[0020] Inputting the network traffic data to be recognized into the network traffic recognition target model, and using the network traffic recognition target model to obtain the type of the network traffic data to be recognized.
[0021] On the other hand, the present invention also provides a network traffic identification system, comprising: a model training module and a traffic identification module, wherein,
[0022] The model training module is used to train a network traffic identification model by using the network traffic data feature representation obtained based on the above method to obtain a network traffic identification target model;
[0023] The traffic identification module is used to input the network traffic data to be identified into the network traffic identification target model, and use the network traffic identification target model to obtain the type of the network traffic data to be identified.
[0024] Advantages of the present invention:
[0025] The present invention expands the sample information of network traffic identification from one-dimensional enhancement to high-dimensional information enhancement, realizes noisy traffic representation based on a dual-representation network, represents traffic from both the label and feature dimensions to enhance its anti-label noise ability, and improves the efficiency of deep neural networks in processing noisy data sets. By training the model using the representation data, superior performance can be demonstrated in scenarios with high label noise. Experimental results show that even under extreme conditions where the training data noise rate is as high as 80%, the solution of this case can still maintain an identification accuracy of over 85%, further indicating that the solution of this case has good potential and value in practical applications and can meet the traffic identification requirements in the field of network intrusion detection. Description of the Drawings
[0026] Figure 1 Schematic diagram of the noisy network traffic representation process based on the dual-representation network in the embodiment;
[0027] Figure 2 Schematic diagram of the principle of the noisy traffic identification algorithm DuRe based on the dual-representation network in the embodiment;
[0028] Figure 3 Schematic diagram of the label representation process in the embodiment;
[0029] Figure 4 Schematic diagram of the change curve of the self-cross entropy value of the probability distribution vector in the embodiment;
[0030] Figure 5 Schematic diagram of each constraint of the feature representation in the embodiment;
[0031] Figure 6 Schematic diagram of the change curve of the representation distance and relative value with the number of rounds in the embodiment. Detailed Embodiments
[0032] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the drawings and technical solutions.
[0033] As a core component of network intrusion detection systems, traffic recognition technology plays a crucial role in identifying malicious attacks. With the advent of the era of full encryption and the rapid development of deep learning, deep learning-based encrypted traffic recognition technology has become a research hotspot. Due to the susceptibility of supervised deep learning methods to label noise, noisy learning techniques have emerged. Existing noisy learning methods mostly process noisy samples from a single dimension, which has limitations. Therefore, in the embodiments of the present invention, see Figure 1 as shown, a noisy network traffic representation method based on a dual-representation network is provided, including:
[0034] S101. Obtain network traffic sample data, where the network traffic sample data is network traffic data with original sample labels, and the original sample labels contain noisy labels;
[0035] S102. Use the network traffic sample data to train the dual-representation network, so that the dual-representation network mines the feature representations of network traffic data from the dual dimensions of label features and traffic features. The dual-representation network includes a label representation network for learning the label distribution in the original sample labels and outputting label representations, and a feature representation network for combining the label representations to mine the noise-resistant traffic features in the original network traffic features and output feature representations.
[0036] Specifically, the label representation network can be designed to include a fully connected layer for updating the sample label distribution and a softmax layer for ensuring that the label distribution conforms to the probability distribution vector distribution characteristics. The feature representation network can be designed to use a multi-layer fully connected autoencoder to perform noise-resistant feature representation mapping on the features of the samples in the feature representation space.
[0037] In the embodiments of this case, the idea of one-dimensional enhancement of existing sample information to high-dimensional information enhancement is adopted to help the neural network more effectively learn the classification features in the samples. High-dimensional information enhancement mines the potential classification features in each dimension of the sample vector through the neural network and combines these features with the samples to generate sample representations rich in category information. During the training process, the neural network reconstructs the information of the samples based on the classification attributes, thereby improving the data information analysis ability of the neural network. Such as Figure 2In the (A method for noisy traffic identification based on a dual - representation network, DuRe) algorithm architecture shown, a dual - dimensional representation of sample labels and features is implemented based on a dual - representation network. Through the dual - dimensional representation, the anti - noise capabilities of both the sample labels and features are enhanced. During the label transformation process, the influence of mislabeling is weakened by enhancing label information. Subsequently, the features are transformed in combination with the label representation results to enhance the anti - noise ability of feature representation, so that the represented samples are robust to label noise at both the label and feature levels.
[0038] Among them, training the dual - representation network using network traffic sample data can be designed to include:
[0039] The KL - divergence function is used to optimize the classification of the label distribution. The cross - entropy is used to quantify the similarity between the original label vector and the label distribution vector, and the self - cross - entropy function is used to keep the label distribution with a single peak, so as to set the label representation network optimization function based on the KL - divergence function, cross - entropy, and self - cross - entropy.
[0040] At the information layer, the feature representation loss is set by minimizing the difference in information entropy before and after representation, and the feature encoding loss is set according to the difference in information entropy between the decoded output and the encoded input. At the classification layer, the label representation is fused to transform the feature distribution of network traffic data into a spatial distribution, and the classification boundary between different feature categories is strengthened by increasing the difference in the distribution of similar and dissimilar features in the feature space, so as to set the feature representation network optimization function through the constraints of the information layer and the classification layer.
[0041] For the original sample labels, based on the label representation network optimization function and through an iterative process, the label representation network learns the label distribution of network traffic data and takes the label distribution as the label representation output.
[0042] For network traffic data, based on the feature representation network optimization function and through an iterative process in combination with the label representation, the anti - noise feature representation of network traffic data features is carried out and the traffic feature representation is output.
[0043] Such as Figure 2In the shown algorithm, a neural network is utilized to learn the noise-resistant label vector representation of traffic data in a label noise environment, thereby improving the recognition performance. DuRe consists of a label representation network and a feature representation network: the former learns the label distribution as the label representation, and the latter combines this label representation to map the noise-resistant feature representation of samples in a low-dimensional feature space. Finally, DuRe outputs the noise-resistant representation for training a classifier to mitigate the impact of the original label noise on the generalization ability of the classifier. Its training process is divided into two stages: label representation and feature representation. In the label representation stage, the label information is enhanced by learning the label distribution, and the label representation network updates the label distribution under the action of the network. In the feature representation stage, the noise-resistant representation of samples is mapped in a low-dimensional space by combining the label representation. The low-dimensional mapping is achieved using the feature representation network, and the mapping result is constrained and optimized using the information layer and the classification layer. The pseudo-code of the algorithm training process is shown in Algorithm 1:
[0044] Algorithm 1. DuRe Training Process.
[0045] Input: Training set Number of epochs for label representation training epoch 1 , Number of epochs for feature representation training epoch 2 , Label representation network model 1 , Feature representation network model 2
[0046] Output: Represented dataset
[0047] 1. First stage: Label representation training
[0048] 2. Initialize the label representation y d = onehot(y), initialize the label representation network model model = model 1 ;
[0049] 3. FOR i = 1 to epoch 1 DO
[0050] 4. Calculate the label representation optimization function:
[0051] I label = α 1 ·I class + α 2 ·I one + α 3 ·I class ;
[0052] 5. Apply the optimization function to the label representation network and update the label representation:
[0053] model ← I label ,
[0054]
[0055] 6. If i = epoch 1 THEN:
[0056] Obtain the labeled representation dataset
[0057] End the labeled representation and enter the feature representation stage. Initialize the feature representation network model = model 2 ;
[0058] 7. FOR i = epoch 1 to epoch 1 + epoch 2 DO
[0059] 8. Calculate the two-dimensional constraints of the feature representation
[0060] 9. Apply the two-dimensional constraints to the feature representation network and update the feature representation:
[0061] x = model(x);
[0062] RETURN
[0063] Among them, the purpose of the labeled representation is to learn the deep classification features of the samples through a deep neural network, strengthen the label information, and reduce noise interference. The traditional method of predicting labels maps high-dimensional features to a one-dimensional label space, which easily leads to information loss and noise amplification. In the embodiments of this case, a neural network is used to learn the label distribution as the labeled representation to reflect the association strength between the samples and each label.
[0064] Such as Figure 3 A brief diagram of the labeled representation method shown. The update criterion of the label distribution can be described as follows:
[0065] For the sample x in the dataset D, denote its label as y, and the label distribution of x output by the labeled representation network in the t-th iteration is Then in the (t + 1)-th iteration, the label distribution The calculation formula is as follows:
[0066]
[0067] Among them, η is the learning rate of the labeled representation network, ε is a hyperparameter, pred(x t ) is the predicted label of x output in the t-th iteration, onethot(pred(x t )) is pred(xt )'s one-hot processing result, I label (x) is the optimization function, grad t+1 ((I label (x))) is the optimization gradient for the (t + 1)-th round.
[0068] It should be noted that in the first round of iteration, the input label is one-hot processed. This process converts the label from the scalar form y to the one-hot vector form y, and uses y to initialize the label distribution y d The initialization formula is as follows:
[0069] Initial(y d ) = softmax(onthot(y)) (2)
[0070] The optimization function I label (x) of the label characterization network is optimized by classification information I class (x), local purity optimization I clean (x) and single-peak optimization I one (x) consists of three parts, and the formula is as follows:
[0071] I label (x) = α 1 ·I class (x) + α 2 ·I one (x) + α 3 ·I label (x). (3)
[0072] Among them, α 1 , α 2 and α 3 are hyperparameters.
[0073] The fully connected network can deeply mine the deep features of the samples, and the predicted labels it outputs reflect the close relationship between the sample features and the predicted categories. Therefore, classifying and optimizing the label distribution based on the prediction results of the fully connected network can ensure that the label distribution fits the classification characteristics shown by the sample features.
[0074] Considering that both the label distribution and the predicted labels of the one-hot processing are high-dimensional vectors, in the embodiments of this case, the KL divergence function is used to measure the distance between the two high-dimensional vectors. The smaller the function value, the closer the distance between the label distribution and the predicted labels, which means that the label distribution incorporates more predicted classification information and better meets the optimization goal of label characterization. Based on the KL divergence function, the classification information optimization function I class (x) is defined as follows:
[0075]
[0076]
[0077] In the training set D, in addition to the label-noisy data, there are also a large number of pure samples with correct labels. Reasonably utilizing these pure samples helps to construct a more accurate label distribution. Assume that the prediction fluctuations of pure samples during training are smaller than those of noisy samples. A purity function f simile (x) can be set. By calculating the similarity between the original label vector and the label distribution vector of the sample, the possibility that the original label of the sample is correct can be judged. In order to quantify the similarity between these two vectors, in the embodiments of this case, based on the cross-entropy function, the formula of the purity function f simile (x) can be expressed as follows:
[0078]
[0079] where y i,j and respectively represent the j-th component of the original label vector y i of x i and the label distribution vector output in this round. According to the property of the cross-entropy function, the closer y i is to , the smaller the value of f simile (x), indicating that the probability that the label y i of x i is correct is greater. In order to screen out the pure samples, a sorting module P(D) of f simile (x i ) can be introduced. P(D) can be expressed as
[0080] Form a new set by taking the samples ranked in the top δ in P(D), denoted as P(D)[δ]. Then, the samples in P(D)[δ] are regarded as pure samples. During the optimization process, it should be ensured that the label distribution of these samples is as close as possible to the original label, that is, the value of f simile (x i ) is as small as possible. Therefore, using the set P(D)[δ] and the values of f simile (x i ) of the samples in it, a local purity optimization function I clean (x) is designed, and the specific formula is as follows:
[0081]
[0082] In the label distribution, there may be multiple non-zero label components, but only one of them is the true label. This means that there is exactly one maximum component in the label distribution, and there should be an obvious numerical difference between this maximum component and other non-maximum components to avoid misjudgment by the network during the learning process. To maintain this characteristic of the label distribution, in the embodiments of this case, the self-cross entropy function is introduced and used as the single-peak optimization function I one (x) of the label characterization network. The specific formula can be expressed as follows:
[0083]
[0084] The reason for choosing the self-cross entropy function as the single-peak optimization function is that it has good properties: the self-cross entropy value of a probability distribution vector with a uniform distribution is the largest, while the self-cross entropy value of a one-hot vector is the smallest. Figure 4 The curve showing the change of the self-cross entropy value of the probability distribution vector with the distribution is presented. It can be seen that when the values of the vector components are equal, the self-cross entropy function value is the largest; as the proportion of the peak component in the whole increases, the self-cross entropy function value gradually decreases; when the vector becomes a one-hot vector, the self-cross entropy function value reaches the minimum. Therefore, using the self-cross entropy function as the optimization function can effectively help the label distribution vector maintain the characteristic of a single peak.
[0085] After being processed by the label characterization network, the sample label y in the original traffic training data set i is characterized as a label distribution The data set after label characterization is denoted as
[0086] The feature characterization process integrates the label characterization result and constructs a noise-resistant representation of the sample in the feature characterization space. In the embodiments of this case, neural network mapping and loss function constraints are used to strengthen the class boundary in the decision space. Specifically, by increasing the distance between different-class samples and reducing the distance between same-class samples, the class distinction becomes more significant, so as to effectively learn classification features and improve the classification accuracy. Feature characterization is carried out in a supervised scenario, and the enhanced label information is used to further strengthen the noise-resistant ability of the features. A double constraint of an information layer and a classification layer is introduced into the feature characterization network, and the feature characterization network is denoted as The general situation of each constraint of the feature characterization method is as Figure 5 shown.
[0087] Feature characterization is essentially a dimensionality reduction mapping of sample features, and this process is necessarily accompanied by information loss. The information layer constraint aims to maximize the retention of the original feature information and reduce the information loss during the mapping process. Information loss can be considered from two dimensions: representation loss and coding loss. The representation loss directly measures the reduction of information entropy before and after representation, while the coding loss calculates the information difference between the encoding input and the decoding output from the encoder-decoder structure of the autoencoder. The specific definitions and calculation formulas are as follows:
[0088] (1) Representation loss constraint: This constraint aims to minimize the information entropy difference before and after representation to ensure that key features are not lost.
[0089] The representation loss constraint can be expressed as
[0090]
[0091] where represents the feature representation of sample x. According to (8), the formula for calculating the representation loss of each sample x l in the input data set D i can be obtained:
[0092]
[0093] where is the th component of. Based on the above formula, the representation loss constraint function can be defined as:
[0094]
[0095] (2) Coding loss constraint: This constraint considers the encoder-decoder structure of the autoencoder and quantifies the information loss during the coding process by calculating the information entropy difference between the decoding output and the coding input. Let the decoder corresponding to the encoder be However, is not 's strict inverse, that is, (where I is the identity mapping network). Therefore, the coding loss is essentially the information loss caused by the non-equivalence between and I. In the embodiments of this case, the KL divergence is used to quantify this loss, and the coding loss constraint can be expressed as:
[0096]
[0097] where represents 's KL divergence loss with I. In actual calculation, the KL divergence between the decoding output and the coding input x i can be used instead. Accordingly, the coding loss constraint function The formula is:
[0098]
[0099] Among them, represents the j-th component of. Combining the representation loss constraint and the encoding loss constraint function, the final expression of the information layer constraint function is:
[0100]
[0101] The classification layer constraint aims to fuse the label representations to achieve noise-resistant reconstruction of features in the low-dimensional space. Its core lies in strengthening the classification characteristics of features, making the features exhibit distinct class attributes, thereby reducing the sample misclassification rate. For this purpose, in the embodiments of this case, based on the assumption that similar samples are close to each other and dissimilar samples are far from each other, the classification characteristics of samples are transformed into spatial distribution characteristics, and the classification boundary between different classes of samples is strengthened by increasing the distribution difference between similar and dissimilar samples in the feature space.
[0102] Denote the mean distance between similar samples in the feature space as d inside , and the mean distance between dissimilar samples as d outside . The classification layer constraint aims to reduce d inside while increasing d outside to highlight the class distinctiveness. Further, the actual goal of this constraint is to maximize the distance difference between external and internal classes while minimizing the distance between internal samples, thereby enhancing the discrimination ability of the model and ensuring the accuracy and robustness of classification. Therefore, the optimization goal of the classification layer constraint can be clearly expressed as:
[0103]
[0104] To obtain the specific classification layer constraint function, the sample distance calculation formula in the feature space can be described as follows:
[0105] Take two samples x from the data set p and x q , and are the feature representation results of x p and x q respectively. Using n′ to represent the dimension of the representation vector, then and The distance calculation formula between them is defined as follows:
[0106]
[0107] According to the label representation data set Perform the division of category clusters, and group the samples with the same highest component in the label representation into the same cluster. Given that the dataset contains c categories of samples, denote the clusters obtained by the division as {cluser 1 ,cluser 2 ,...,cluser c}, and assume that each cluster contains n 1 ,n 2 ,...,n c samples. The feature representation process does not change the cluster membership of the samples. Therefore, the cluster division before and after the feature representation remains consistent. For any sample x p ∈cluser k , the average distance from its feature representation to the other feature representations within the cluster can be expressed as:
[0108]
[0109] Using the above formula, calculate the average distance from the feature representation of each sample in cluser k to the other sample feature representations within the cluster, and store these average distance values in the set , that is Then, find the sample with the minimum average distance in , and denote it as the in-cluster confident sample k of cluser Represent the position of the cluster cluser in the feature space by the position of k , then the distance p between the feature representation of the sample x and other clusters is defined as follows:
[0110]
[0111] Thus, the formula for calculating the average distance outside the cluster of the feature representation of x p can be obtained:
[0112]
[0113] By calculating the average distance within the cluster and the average distance outside the cluster of each sample feature representation, the overall average distance within the cluster d inside and the average distance outside the cluster d outside can be obtained, and their calculation formulas are as follows:
[0114]
[0115] According to (14), denote the relative quantity for the classification layer constraint as ω, then the calculation formula for ω is:
[0116]
[0117] Among them, Normal() is a normalization function used to ensure the convergence of the fraction. The specific calculation formula of the classification layer constraint function is as follows:
[0118]
[0119] Combined with the information layer constraint and the distance layer constraint Set hyperparameters β 1 and β 2 , and the overall constraint function of the feature representation network is obtained as follows:
[0120]
[0121] Furthermore, based on the above-mentioned noisy network traffic representation method, an embodiment of the present invention also provides a network traffic recognition method, including:
[0122] Using the network traffic data feature representation obtained based on the above-mentioned noisy network traffic representation method to train a network traffic recognition model to obtain a network traffic recognition target model;
[0123] Input the network traffic data to be recognized into the network traffic recognition target model, and use the network traffic recognition target model to obtain the type of the network traffic data to be recognized.
[0124] Output the network traffic data feature representation and connect it to the network traffic recognition model to realize the recognition and classification of the representation data set. In the embodiment of this case, the Kmeans model can be selected as the classifier in the network traffic recognition model, and the data set is divided into different clusters based on the relative distance of samples in the feature representation space. Then, the Kuhn-Munkres algorithm is used to correspond the clusters in the feature representation space with the corresponding labels, and finally the encrypted traffic recognition task is completed.
[0125] Furthermore, based on the above-mentioned noisy network traffic representation method, an embodiment of the present invention also provides a network traffic recognition system, including: a model training module and a traffic recognition module, where
[0126] The model training module is used to train a network traffic recognition model using the network traffic data feature representation obtained based on the above method to obtain a network traffic recognition target model;
[0127] The traffic recognition module is used to input the network traffic data to be recognized into the network traffic recognition target model and use the network traffic recognition target model to obtain the type of the network traffic data to be recognized.
[0128] To verify the effectiveness of the solution in this case, the following further explains with experimental data:
[0129] The DuRe algorithm of the solution in this case obtains the noise-resistant representation of traffic in the label noise scenario by enhancing the label and feature information, ensuring that the traffic can convey accurate classification information even in the presence of label noise. The performance of the DuRe algorithm is analyzed below through experimental data, mainly focusing on the following three core issues: (1) the effectiveness of label representation; (2) the effectiveness of feature representation; (3) the superiority of the DuRe method compared with other methods.
[0130] To evaluate the solution in this case, two traffic datasets, MALICIOUS_TLS and CICDS2017, are used in the experiment. The specific information is outlined as follows:
[0131] (1) MALICIOUS_TLS dataset: Constructed to verify the performance of MCRe, it contains 22 types of malicious and benign traffic encrypted using the TLS protocol. The traffic features are highly similar, making it difficult to identify. The training and testing environments provided by this dataset are close to the real network, which helps to verify the actual application effect of the method.
[0132] (2) CICDS2017 dataset: An open traffic dataset that contains malicious traffic and benign traffic collected by simulating 12 types of real network attacks. The CICDS2017 dataset has a rich variety of traffic types and is close to real traffic scenarios, and is widely used in the field of intrusion traffic detection.
[0133] Both MALICIOUS_TLS and CICDS2017 are datasets with accurate labels. To simulate the label noise environment, noise injection needs to be performed on the datasets. Before injection, the datasets are divided into a training set and a testing set in a ratio of 7:3. The testing set maintains the original labels, and noise is only injected into the training set. The injection ratio is determined by the preset noise rate. Noise injection is completed by randomly selecting samples in the training set and replacing the original labels with noise labels.
[0134] According to the differences in traffic annotation strategies, noise can be divided into symmetric noise and asymmetric noise. For these two types of noise, two noise label generation algorithms are designed to inject different types of noise into the datasets:
[0135] (1) Symmetric noise generation: In this scenario, the sample labels are mislabeled as any other category with the same probability. A random number generator is used to generate a random label different from the original label for the selected samples to simulate the symmetric noise in the real annotation process.
[0136] (2) Asymmetric noise generation: In this scenario, the probability of mislabeling different class samples is different. By generating the same wrong labels for the selected samples, that is, mislabeling malicious traffic as benign traffic, to simulate the situation where annotators in the real scenario cannot identify malicious traffic.
[0137] The dataset is divided into a training set and a test set according to the ratio of 7:3. A certain proportion of noise is injected into the training set, and the test set is not processed. Each experiment conducts 100 iterations of training, with 20 rounds in the first stage and 80 rounds in the second stage. After each round of training, the training results are tested on the test set.
[0138] 1. Effectiveness of feature representation
[0139] The design goal of feature representation is to strengthen the classification boundary between different class samples, aiming to reduce the risk of sample misclassification, and further reduce the adverse impact of noisy labels on the sample classification accuracy. To evaluate the effectiveness of the feature representation module in strengthening the classification boundary, this part conducts a verification experiment on the CICDS2017 dataset. Noise with proportions of 20% and 40% is respectively injected into the dataset to construct two noise environments: symmetric noise and asymmetric noise. 100 rounds of iterative training are independently executed for the feature representation module. During this period, the average distance outside the cluster, the average distance inside the cluster, and the relative magnitude of the classification layer constraint are systematically recorded in each round of training, and a curve graph showing the change of the representation distance and the relative magnitude with the number of iterative rounds is drawn, as Figure 6 shown.
[0140] Observation Figure 6 It can be seen that during the training process, the average distance outside the cluster increases significantly, demonstrating the effectiveness of the feature representation module in expanding the classification boundary between different class samples. In addition, at the initial stage of training, the three parameters under investigation all show a certain fluctuating upward trend; as the number of training rounds increases, the model gradually converges, and the values of each parameter fall back and tend to be stable accordingly, finally constructing a relatively stable classification boundary.
[0141] 2. Comparative experiments and analysis
[0142] Compare the traffic recognition model trained by the DuRe algorithm of this case with the models trained by other methods as follows:
[0143] INCV, FINE, and ULDC: Clean the noisy dataset and use the cleaned dataset to train a supervised classifier to identify the traffic types in the validation set;
[0144] Co-Teaching: Train two networks with different initial states simultaneously. The two networks will cooperate with each other and select pure samples from each other's outputs to reduce the negative impact of noisy samples on the training process.
[0145] Co-Learning: By constraining the consistency of the dual-head encoder, it ensures that they can produce consistent outputs when processing the same input, thus avoiding misfitting noisy labels to the network model.
[0146] Comparison of Classification Accuracy Results on the MALICIOUS_TLS Dataset
[0147]
[0148] Comparison of Classification Accuracy Results on the CICDS2017 Dataset
[0149]
[0150] Tables 1 and 2 show the classification accuracy performance of DuRe in the two datasets of MALICIOUS_TLS and CICDS2017 under different noise rates and two noise environments of symmetric and asymmetric. According to the experimental results, the following conclusions can be obtained:
[0151] 1). At a low noise level (noise rate equal to 20%), the performance of the DuRe algorithm of the proposed solution in this case is extremely excellent, only slightly lower than the best-performing method, with a tiny gap. On the MALICIOUS_TLS dataset, whether it is symmetric noise or asymmetric noise, the accuracy of DuRe follows closely behind Co-Learning, with a gap of no more than 1.2%. Similarly, on the CICDS2017 dataset, the performance of DuRe under symmetric and asymmetric noise is also only slightly inferior to the best-performing method, with the gap remaining within a very small range. This result shows that the DuRe algorithm of the proposed solution in this case can still maintain a high classification accuracy in a low-noise environment, and it does not overly sacrifice the performance in a low-noise environment in exchange for robustness to a high-noise environment.
[0152] 2). At a high noise level (noise rate greater than 20%), the performance of the DuRe algorithm of the proposed solution in this case shows obvious advantages. As the noise level increases, the accuracy of other methods generally drops significantly, while DuRe maintains relatively stable performance. Especially when the noise level reaches the extreme (such as 80%), the accuracy of DuRe is much higher than that of other methods. For example, in the case of 80% asymmetric noise on the MALICIOUS_TLS dataset, the accuracy of DuRe is nearly 20 percentage points higher than that of the second-best-performing method; in the case of 80% symmetric noise on the CICDS2017 dataset, the accuracy of DuRe is also significantly better than that of other methods. This result fully demonstrates the effectiveness of the DuRe algorithm of the proposed solution in this case in dealing with high-noise data, and its design strategy can significantly improve the robustness of the classifier in a complex noise environment.
[0153] 3). The DuRe method exhibits excellent robustness in the label noise scenario. As can be seen from the tabular data, as the noise ratio increases, the performance of DuRe decreases relatively slightly. This indicates that the DuRe algorithm of the present solution has a strong anti-noise ability and can maintain a high classification accuracy under noise interference. This characteristic stems from the fact that the DuRe algorithm of the present solution uses multiple constraints to optimize the representation function when dealing with label noise and adopts an appropriate classifier to mitigate the impact of noisy labels on the classification performance. Therefore, the DuRe algorithm of the present solution has significant advantages when dealing with malicious traffic datasets containing a large amount of noise.
[0154] In summary, the DuRe algorithm of the present solution performs stably at low noise levels, only slightly lower than the optimal method; while at high noise levels, it demonstrates obvious performance advantages. It shows excellent robustness to label noise and can maintain a high classification accuracy in a complex noise environment. These characteristics make the DuRe algorithm of the present solution have broad application prospects when dealing with malicious traffic datasets containing a large amount of noise.
[0155] 3. Effectiveness of label representation
[0156] There are three optimization functions in the label representation: classification information optimization I class , local purity optimization I clean and single peak optimization I one . To verify the effectiveness of these three optimizations, I class , I clean and I one are respectively eliminated during training, thereby obtaining the models DuRe / I class , DuRe / I clean and DuRe / I one .
[0157] Table 3 Ablation analysis on the MALICIOUS_TLS dataset
[0158]
[0159]
[0160] Table 4 Ablation analysis on the CICDS2017 dataset
[0161]
[0162] Tables 3 and 4 show DuRe and DuRe / I class , DuRe / I clean and DuRe / I oneIn the two datasets of MALICIOUS_TLS and CICDS2017, the classification accuracy performances under different noise rates and in both symmetric and asymmetric noise environments. According to the experimental results, the following conclusions can be drawn:
[0163] 1). All these three optimization conditions contribute to the enhancement of traffic label information. Among all the models, the complete DuRe achieved the best results in all scenarios and tasks. The three proposed optimization conditions aim to optimize the label distribution from three aspects: classification information, local purity, and single peak, so as to help obtain better feature representation results and mitigate the impact of noisy labels on classification.
[0164] 2). Among all the optimization conditions, the classification information optimization helps DuRe enhance label information the most. As the proportion of noisy labels increases, the performance of all three DuRe variants shows a significant decline to varying degrees. Among them, DuRe / I class has the largest decline. In the scenario with a noise rate of 80% in the MALICIOUS_TLS dataset, the classification accuracy is only about 65%. This result meets the design expectation because the neural network prediction results have the greatest impact on the optimization of label distribution.
[0165] The above data shows that even under the extreme condition where the noise rate of the training data is as high as 80%, the proposed solution in this case can still maintain an identification accuracy of over 85%, thus fully demonstrating its good application potential and value in the field of network traffic intrusion detection.
[0166] Unless otherwise specifically stated, the relative steps, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of the present invention.
[0167] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0168] The units and method steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those of ordinary skill in the art can use different methods to implement the described functions for each specific application, but such implementation is not considered to exceed the scope of the present invention.
[0169] Those of ordinary skill in the art can understand that all or part of the steps in the above method can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium, such as: read-only memory, magnetic disk or optical disc, etc. Optionally, all or part of the steps of the above embodiments can also be implemented using one or more integrated circuits. Correspondingly, each module / unit in the above embodiments can be implemented in the form of hardware or in the form of a software functional module. The present invention is not limited to any specific form of combination of hardware and software.
[0170] Finally, it should be noted that: the above embodiments are only specific embodiments of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions recorded in the foregoing embodiments or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A noisy network traffic representation method based on a dual representation network, characterized in that: Include: Acquire network traffic sample data, where the network traffic sample data is network traffic data with original sample labels, and the original sample labels include noise labels; The dual representation network is trained using network traffic sample data, so that the dual representation network mines the network traffic data feature representation from two dimensions: label feature and traffic feature. The dual representation network includes a label representation network for learning the label distribution in the original sample label and outputting the label representation, and a feature representation network for mining the anti-noise traffic feature in the original traffic feature of the network in combination with the label representation and outputting the feature representation.
2. The noisy network traffic representation method based on a dual representation network according to claim 1 is characterized in that: The label representation network includes: a fully connected layer for updating the sample label distribution, and a softmax layer for ensuring that the label distribution conforms to the distribution characteristics of the probability distribution vector.
3. The noisy network traffic representation method based on a dual representation network according to claim 1 is characterized in that: The feature characterization network uses a multi-layer fully connected automatic encoder to perform noise-resistant feature characterization mapping on the features of the samples in the feature characterization space.
4. The noisy network traffic representation method based on a dual representation network according to claim 1 is characterized in that: The dual representation network is trained using network traffic sample data, including: The KL divergence function is used to classify and optimize the label distribution, the cross entropy is used to set the quantified similarity between the original label vector and the label distribution vector, and the self-cross entropy function is used to keep the label distribution at a single peak, so as to set the label representation network optimization function based on the KL divergence function, cross entropy and self-cross entropy; At the information layer, the feature representation loss is set by minimizing the difference in information entropy before and after representation, and the feature encoding loss is set according to the difference in information entropy between the decoded output and the encoded input; at the classification layer, the label representation is fused to transform the feature distribution of network traffic data into a spatial distribution, and the classification boundaries between feature categories are strengthened by increasing the difference in the distribution of similar and heterogeneous features in the feature space, so as to set the feature representation network optimization function through information layer constraints and classification layer constraints; For the original sample labels, based on the label representation network optimization function and through a cyclic iterative process, the label representation network learns the label distribution of the network traffic data and outputs the label distribution as the label representation; For network traffic data, based on the feature representation network optimization function and through an iterative process combined with label representation, the network traffic data features are represented by anti-noise features and the traffic feature representation is output.
5. The noisy network traffic representation method based on a dual representation network according to claim 4 is characterized in that: The label representation network optimization function is expressed as: label (x) = α1·I class (x)+α2·I one (x)+α3·I label (x), where α1, α2, and α3 are hyperparameters, x represents the label representation network input, and I label (x) represents the label representation network optimization function, I class (x) represents the label distribution classification optimization function based on the KL divergence function, I clean (x) represents the optimization function that quantifies the similarity between the original label and the label distribution based on the cross entropy setting and performs local label purification. one (x) represents a single peak optimization function that keeps the label distribution single peak based on the self-cross entropy setting.
6. The noisy network traffic representation method based on a dual representation network according to claim 4 is characterized in that: The feature representation network optimization function is expressed as: Among them, β1 and β2 are hyperparameters, x is the feature representation network input, represents the feature representation network optimization function, Represents information layer constraints, represents the classification layer constraint, and represents the loss constraint function, represents the encoding loss constraint function, ω represents the relative magnitude of the classification layer constraint, and Normal() is the standardization function, d inside represents the average distance within the traffic feature cluster, d outside Represents the average distance outside the traffic feature cluster.
7. A network traffic identification method, characterized in that: Include: Using the network traffic data feature representation obtained based on the method of claim 1 to train a network traffic identification model to obtain a network traffic identification target model; The network traffic data to be identified is input into the network traffic identification target model, and the network traffic data type to be identified is obtained using the network traffic identification target model.
8. A network traffic identification system, characterized in that: Contains: model training module and traffic identification module, among which, A model training module, used to train a network traffic identification model using the network traffic data feature representation obtained based on the method of claim 1 to obtain a network traffic identification target model; The traffic identification module is used to input the network traffic data to be identified into the network traffic identification target model, and use the network traffic identification target model to obtain the type of network traffic data to be identified.
9. An electronic device, characterized in that: include: at least one processor, and a memory coupled to the at least one processor; The memory stores a computer program, and the computer program can be executed by the at least one processor to implement the method according to any one of claims 1 to 6 or the method according to claim 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed, the method according to any one of claims 1 to 6 or the method according to claim 7 can be implemented.