Cryptographic malicious traffic recognition method and system facing characteristics of multi-class imbalance data

CN116436678BActive Publication Date: 2026-09-11SHANDONG COMP SCI CENTNAT SUPERCOMP CENT IN JINAN +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310436976.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-18
Publication Date
2026-09-11
Estimated Expiration
2043-04-18

AI Technical Summary

Technical Problem

但由于该算法自身结构的限制,其需要满足Lipschitz约束

Benefits of technology

[0021] 1. This invention establishes a self-attention module in the malicious traffic sample generation model to help the model effectively model complex dependencies with low computational cost, so as to generate more refined and reliable malicious traffic samples. When identifying the malicious traffic samples, it can improve the accuracy of malicious traffic identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116436678B_ABST
    Figure CN116436678B_ABST
Patent Text Reader

Abstract

The application discloses a malicious traffic recognition method and system for multi-class unbalanced data characteristics, and performs preprocessing on network malicious traffic data to obtain a feature image of malicious traffic; a malicious traffic sample is obtained through the feature image of the malicious traffic and a trained malicious traffic sample generation model; a self-attention module is added to both a generator and a discriminator of the malicious traffic sample generation model; the self-attention module takes a feature map output by a convolution layer as input to obtain an attention map of the feature map input into the self-attention module, determines a self-attention feature map according to the attention map, and obtains a feature map output by the self-attention module by weighted summation of the self-attention feature map and the feature map input into the self-attention module; the feature map output by the self-attention module is input into a next convolution layer; and the malicious traffic sample is recognized to obtain a malicious traffic recognition result. The accuracy of malicious traffic recognition is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of traffic identification technology, and in particular to a method and system for identifying encrypted malicious traffic that addresses the characteristics of various types of imbalanced data. Background Technology

[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.

[0003] In recent years, with the continuous expansion of internet applications, internet network security is facing major challenges such as cyberattacks, data breaches, and information theft. Malicious traffic classification technology has gradually become a mainstay in the field of network security. Due to the widespread application of dynamic port technology and the explosive growth of traffic, traditional port-based and payload-based traffic classification methods are no longer able to accurately classify malicious traffic in the current network environment. Encrypted malicious traffic classification methods based on traditional machine learning and deep learning have become mainstream research directions, but a crucial prerequisite for them to achieve good classification performance is a relatively balanced dataset. As malicious traffic is a typical minority class, generating reliable and realistic malicious traffic samples to construct a relatively balanced dataset is an urgent need.

[0004] In class imbalance problems, researchers typically study both the data level and the classification algorithm. For example, they may use generative adversarial networks (GANs) to synthesize malicious traffic to represent the minority class, thereby enhancing the original dataset, or introduce cost-sensitive strategies to improve the ability of existing classification algorithms to learn features of the minority class.

[0005] Current methods for generating traffic data based on GANs primarily use the conversion of data packets into grayscale images as input. However, this method requires padding or truncation of data packets during the conversion process, significantly increasing the risk of information loss and redundancy during data preprocessing. Furthermore, due to the numerous shortcomings of the original GAN, researchers have proposed various improved GAN models, with WGAN being a successful variant. Its fundamental improvement over the original GAN ​​lies in modifying the loss function. This model introduces Wasserstein distance to replace JS and KL divergence, solving the gradient vanishing problem caused by the discontinuity of JS divergence when two distributions do not intersect. Simultaneously, the mode collapse problem of GANs is also adequately addressed in WGAN. However, due to the inherent structural limitations of the algorithm, it needs to satisfy Lipschitz constraints. The weight pruning method used in traditional WGAN is not ideal. Weight pruning leads to extreme weight distributions, making WGAN optimization difficult and resulting in poor sample generation reliability. Moreover, because WGAN is unsupervised and generates random samples, it is not suitable for multi-class imbalanced scenarios in the current network environment. Meanwhile, current improvements to GANs still suffer from drawbacks such as the inability to accurately apply complex geometric constraints to the global image structure, poor ability to model long-distance, multi-level dependencies across image regions, high computational costs, and low algorithm efficiency. Summary of the Invention

[0006] To address the aforementioned problems, this invention proposes an encrypted malicious traffic identification method and system tailored to the characteristics of various types of imbalanced data, achieving accurate identification of malicious traffic.

[0007] To achieve the above objectives, the present invention adopts the following technical solution:

[0008] Firstly, a method for identifying encrypted malicious traffic tailored to the characteristics of various types of imbalanced data is proposed, including:

[0009] Obtaining malicious network traffic data;

[0010] Preprocess network malicious traffic data to obtain characteristic images of malicious traffic;

[0011] Malicious traffic samples are obtained by using the feature images of malicious traffic and a trained malicious traffic sample generation model. The malicious traffic sample generation model is constructed using the CWGAN-GP model. Self-attention modules are added to both the generator and the discriminator. The self-attention module takes the feature map output by the convolutional layer as input, obtains the attention map of the feature map input to the self-attention module, determines the self-attention feature map based on the attention map, and performs a weighted summation of the self-attention feature map and the feature map input to the self-attention module to obtain the feature map output by the self-attention module. The feature map output by the self-attention module is then input into the next convolutional layer.

[0012] Malicious traffic identification results are obtained by using malicious traffic samples and a trained malicious traffic identification model.

[0013] Secondly, an encrypted malicious traffic identification system tailored to the characteristics of various types of imbalanced data is proposed, including:

[0014] The traffic data acquisition module is used to acquire malicious network traffic data;

[0015] The preprocessing module is used to preprocess malicious network traffic data to obtain characteristic images of the malicious traffic;

[0016] The traffic sample generation module is used to obtain malicious traffic samples through the feature image of malicious traffic and the trained malicious traffic sample generation model. The malicious traffic sample generation model is constructed using the CWGAN-GP model. Self-attention modules are added to both the generator and the discriminator. The self-attention module takes the feature map output by the convolutional layer as input, obtains the attention map of the feature map input to the self-attention module, determines the self-attention feature map based on the attention map, and performs a weighted summation of the self-attention feature map and the feature map input to the self-attention module to obtain the feature map output by the self-attention module. The feature map output by the self-attention module is then input into the next convolutional layer.

[0017] The traffic identification module is used to obtain malicious traffic identification results through malicious traffic samples and a trained malicious traffic identification model.

[0018] Thirdly, an electronic device is proposed, including a memory and a processor, as well as computer instructions stored in the memory and running on the processor. When the processor executes the computer instructions, the steps described in the encrypted malicious traffic identification method for multiple types of imbalanced data characteristics are completed.

[0019] Fourthly, a computer-readable storage medium is proposed for storing computer instructions, which, when executed by a processor, complete the steps described in the encrypted malicious traffic identification method for multiple types of imbalanced data characteristics.

[0020] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0021] 1. This invention establishes a self-attention module in the malicious traffic sample generation model to help the model effectively model complex dependencies with low computational cost, so as to generate more refined and reliable malicious traffic samples. When identifying the malicious traffic samples, it can improve the accuracy of malicious traffic identification.

[0022] 2. This invention mainly focuses on class imbalance, generating data for minority class samples represented by malicious traffic to construct a relatively balanced dataset and improve the accuracy of the classification model.

[0023] 3. This invention does not require uniform data packet length for the preprocessing of malicious traffic, thereby avoiding information loss and redundancy during the preprocessing process and further ensuring the accuracy of malicious traffic identification.

[0024] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0025] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments of this application and their descriptions are used to explain this application and do not constitute an undue limitation of this application.

[0026] Figure 1 This is a framework diagram of the method disclosed in Example 1;

[0027] Figure 2 This is a diagram of the preprocessing procedure disclosed in Example 1;

[0028] Figure 3 This is a schematic diagram of the malicious traffic sample generation model disclosed in Example 1;

[0029] Figure 4 This is a schematic diagram of the self-attention module disclosed in Example 1. Detailed Implementation

[0030] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0031] It should be noted that the following detailed descriptions are illustrative and intended to provide further explanation of this application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.

[0032] Example 1

[0033] This embodiment discloses a method for identifying encrypted malicious traffic based on the characteristics of various types of imbalanced data, such as... Figure 1 As shown, it includes:

[0034] S1: Obtain malicious network traffic data.

[0035] The acquired malicious network traffic data is in PCAP format.

[0036] S2: Preprocess malicious network traffic data to obtain a characteristic image of the malicious traffic.

[0037] Currently, GAN-based traffic data generation methods primarily use the conversion of data packets into grayscale images as input. However, these methods require padding or truncation of data packets during the conversion process to standardize packet length, leading to information loss and redundancy issues in traditional data preprocessing. In this embodiment, to prevent information loss and redundancy when preprocessing malicious traffic data, the preprocessing process for malicious network traffic data is as follows: Figure 2 As shown, after performing preprocessing operations such as traffic segmentation and diversion on the malicious traffic data in PCAP format obtained by S1, a feature image in a unified format is obtained.

[0038] Specifically, the preprocessing process for malicious network traffic data includes:

[0039] S21: Perform traffic segmentation on malicious network traffic data to obtain multiple session files.

[0040] Malicious traffic data in PCAP format is segmented into multiple session files based on the 5-tuple.

[0041] The quintuple includes the source IP address, source port, destination IP address, destination port, and transport layer protocol.

[0042] S22: Perform traffic scrubbing on all session files to obtain the traffic-scrubbing session files.

[0043] The session file obtained by S21 is filtered using tools such as Wireshark to remove duplicate and useless packets, such as TCP retransmission packets, to obtain the filtered session file.

[0044] Remove useless information from the header fields of the filtered session file to prevent the model from overfitting to such information.

[0045] Useless information includes noise, data link layer information, and IP addresses.

[0046] S23: Using a Markov model, generate a feature image of the session file after traffic scrubbing, which is a feature image of malicious traffic.

[0047] Viewing the traffic data in a session as a byte stream, with each byte having a different state transition probability, this transmission process exhibits the characteristics of a Markov chain. Therefore, using a Markov model can effectively characterize the spatiotemporal features of session payload.

[0048] The encrypted traffic data in the session file is read in binary format. Then, the binary data is converted into a Markov matrix, with each byte as a value. The specific calculation process is as follows:

[0049] v = {B1, B2, ..., B} N}, B i ∈(0, 1, 2, ..., 255), i∈(1, 2, ...., N) (1)

[0050] Where vector v is the encoded state vector, N is the byte length, and B i Let v be the value of the i-th byte. Treat v as a Markov chain and compute the values ​​of two adjacent random states B. i+1 and B i The transition probability, i.e., B i+1 Appeared in B i The subsequent probability is denoted as the transition probability distribution P(B). i+1 |B i ):

[0051]

[0052] In the above formula: For B i+1 Appeared in B i The probability after that; P(B) i ) indicates that B is in v. i The probability of occurrence; P(B) i+1 B i ) is a joint probability representation of B in v. i and B i+1 The probability of them occurring simultaneously; This means that in v, any state in the first i states is followed by B. i The sum of conditional probabilities, B. i B i+1 ∈(0, 1, 2, ..., 255). Then, by statistically analyzing the transition probabilities of individual session traffic data, the transition probability matrix M of this Markov process can be obtained:

[0053]

[0054] Each value in the Markov matrix of the traffic data represents the transmission probability, or transition probability, of a byte, rather than the actual value of the byte. Since sessions of the same type of traffic have similar transmission probabilities, Markov images of the same type of traffic are homologous in terms of transition probabilities. They represent the distribution characteristics of the network session field. Finally, each value in the matrix is ​​converted into a pixel in the feature image to form the final feature image.

[0055] This embodiment avoids the information loss and redundancy caused by truncating or padding data packets to a uniform length when converting them to grayscale images in traditional preprocessing schemes. When using this feature image to generate malicious samples, more sample features can be retained, thereby improving the accuracy of malicious traffic identification.

[0056] S3: Malicious traffic samples are obtained by using the feature images of malicious traffic and the trained malicious traffic sample generation model. The malicious traffic sample generation model is constructed using the CWGAN-GP model. Self-attention modules are added to both the generator and the discriminator. The self-attention module takes the feature map output by the convolutional layer as input, obtains the attention map of the feature map input to the self-attention module, determines the self-attention feature map based on the attention map, and performs a weighted summation of the self-attention feature map and the feature map input to the self-attention module to obtain the feature map output by the self-attention module. The feature map output by the self-attention module is then input into the next convolutional layer.

[0057] Traffic data generation methods based on GANs, improved using WGAN, require strict adherence to Lipschitz constraints during training to ensure model effectiveness due to the limitations of their algorithmic structure. Current WGAN weight pruning methods truncate or scale weights during training to limit their range. While this method effectively satisfies the Lipschitz constraints, it can lead to weight extrema, difficulty in model optimization, and still fails to effectively avoid gradient vanishing problems. Furthermore, as an unsupervised model, WGAN-based traffic data generation methods are only applicable to binary imbalance scenarios and cannot be applied to multi-class imbalance scenarios.

[0058] Current traffic data generation methods based on GANs and their variants generally rely on convolution to model long-range, multi-level dependencies. Since convolutional kernels can only process information within local regions of an image, multiple convolutional layers are needed to handle long-range dependencies. However, adding convolutional layers to model long-range relationships faces challenges such as inapplicability to small models, difficulty in finding a set of optimal parameters to precisely coordinate multiple convolutional layers and capture these complex dependencies, and high computational costs. Self-attention mechanisms, on the other hand, demonstrate a good balance between the ability to model long-range dependencies and computational and statistical efficiency. Furthermore, the calculation of weights and attention vectors in self-attention mechanisms requires very little computation.

[0059] Therefore, this embodiment constructs a malicious traffic sample generation model by using the feature image of malicious traffic as input and malicious traffic samples as output through the improved WGAN-based model CWGAN-GP. Self-attention modules are added to both the generator and discriminator of CWGAN-GP to help the model effectively model complex dependencies with low computational cost, so as to generate more refined and reliable malicious traffic samples.

[0060] like Figure 3 As shown, CWGAN-GP includes a generator and a discriminator. The generator takes the feature image of malicious traffic as input, concatenates the feature image of malicious traffic with application class labels to form a noise vector, and feeds it into a fully connected layer for dimensionality transformation. Then, it is input into a transposed convolutional layer with ReLU activation function for convolution operation. After convolution, to enhance the model's ability to associate distant pixels with model feature content, the feature image is input into the attention module to generate an attention feature map. Finally, the attention feature map is input into a transposed convolutional layer with a single convolution kernel, stride of 1, and sigmoid activation function to output generated samples.

[0061] The generated samples output by the generator are input to the discriminator, which also receives feature images of real malicious samples and their corresponding traffic tags. After convolutional operations using a LeakyReLU activation function, the feature images are fed into a self-attention module. By calculating the similarity between different locations in the feature map, the self-attention mechanism assigns a weight to each location, making the discriminator focus more on locations useful for determining authenticity, thus improving its accuracy in judging authenticity and generating images. Finally, the attention feature images are input to a fully connected layer to calculate the authenticity of the input samples and the probability that the samples belong to a certain category.

[0062] CWGAN-GP is an improved model based on WGAN, addressing its main shortcomings. CWGAN-GP effectively solves the problem of extreme weight distribution caused by weight clipping by introducing gradient penalty instead of weight clipping, making the discriminator more effectively satisfy the Lipschitz constraint. The gradient penalty is implemented by establishing a L2 norm between the obtained discriminator gradient and a user-defined constant K. First, random numbers are drawn from the range of 0-1, and then summed proportionally between each pair of real and fake samples. Finally, the Lipschitz constraint is achieved by sampling each batch of data. This follows the distribution P of the real data. r and the generated data distribution P g Linear sampling between The formula is shown below:

[0063]

[0064] Where x represents the distribution of the real data Pr The samples obtained from sampling Indicates the distribution of generated data P g The sample obtained by sampling. λ is a random variable uniformly distributed on [0, 1]. Represents a space between x and The intermediate value between x and λ is the data sample obtained by linear interpolation between the actual data sample and the data sample generated by the generator. λ controls x and λ. Both in The weight it occupies in the middle.

[0065] Meanwhile, to establish a supervised data generation process and make the model applicable to various imbalanced scenarios in modern network environments, the CWGAN-GP model adds extra information *c* as conditional constraints to both the discriminator and generator. In this embodiment, the category labels of the traffic data are input into the model along with the feature image samples of the traffic data as extra information to form conditional constraints. In summary, the objective function of the CWGAN-GP model is:

[0066]

[0067] Where E represents the expected value, and c represents the conditional information. Indicates the distribution of real data P r The expected value of sample x obtained by random sampling. P represents the sample distribution generated from generator G. g Samples obtained by random sampling Expected value. Indicates from the sampling distribution Samples obtained by random sampling The expected value of the sample x is given by the discriminator D under condition c. This indicates that, given condition c, the sample The probability that the discriminator D identifies it as a generated sample. Indicates the distribution of real data P r The expected value of the probability that a sample x obtained by random sampling is identified as a real sample by the discriminator D. Indicates the distribution of real data P r The expected value of the probability that a sample x obtained by random sampling is identified as a real sample by the discriminator D. Indicates sample The gradient in the discriminator, This is the gradient penalty term. For each sample in a batch, the gradient penalty value can be considered as a random variable. The expected value is used to represent the average of this random variable and serves as the gradient penalty term for the entire batch. Here, ε is a hyperparameter used to control the degree of penalty.

[0068] The loss functions for the discriminator and generator are shown below:

[0069]

[0070]

[0071] Where L(D) is the discriminator loss function and L(G) is the generator loss function.

[0072] To obtain global geometric features of images and improve the quality of generated malicious traffic sample images, a self-attention module is added to both the generator and discriminator of the CWGAN-GP model. The self-attention module complements the convolutional structure of the generator and discriminator, enabling the model to quickly and accurately locate key regions in the image. This helps to model long-distance, multi-level dependencies across image regions, and the module has low computational cost.

[0073] The generator comprises an input layer, a Dense layer, multiple convolutional layers, a self-attention module, and an output layer. The feature image of malicious traffic is input into the input layer, which concatenates the feature image with multiple labels to form a noise vector. This noise vector is then sent to the Dense layer for sampling, outputting the sampled image. The sampled image enters the first convolutional layer, and its output enters the second convolutional layer. The feature map output from the second convolutional layer enters the self-attention module, and its output enters the third convolutional layer. The output from the third convolutional layer enters the fourth convolutional layer, and its output enters the output layer, which outputs the generator's generated samples. All convolutional layers in the generator are transposed convolutional layers.

[0074] The discriminator consists of an input layer, convolutional layers, a self-attention module, a flatten layer, and a dense layer. The generated samples from the generator are fed into the first convolutional layer via the input layer. The output of the first convolutional layer is fed into the second convolutional layer. The output of the second convolutional layer is fed into the self-attention module. The feature map output by the self-attention module is fed into the third convolutional layer. The output of the third convolutional layer is fed into the fourth convolutional layer. The output of the fourth convolutional layer is fed into the flatten layer. The output of the flatten layer is fed into the dense layer. The dense layer calculates the authenticity of the generated samples and the probability that the generated samples belong to a certain category, and outputs the results.

[0075] In its implementation, the generator structure includes an input layer, a Dense layer, a transposed convolutional layer, a self-attention module, and an output layer. First, the input layer concatenates 100-dimensional random noise z with 10 application class labels to form a noise vector, which is then sampled by the Dense layer, resulting in an output dimension of (b, 50176), where b is the batch size. After reshaping, the output dimension of the Dense layer becomes (b, 16, 16, 256). Next, two transposed convolutional layers with 5×5 kernels and a stride of 2 are passed, containing 128 and 64 kernels respectively, resulting in output dimensions of (b, 64, 64, 128) and (b, 128, 128, 64). A self-attention module is added after the fourth layer to enhance the model's ability to correlate distant pixels with model features, improving the generation quality. This module ultimately outputs a feature map of (b, 128, 128, 64). The feature map obtained through the self-attention module is passed through a transposed convolutional layer with a stride of 1 and 32 kernels, resulting in an output dimension of (b, 128, 128, 32). This (b, 128, 128, 32) feature map then enters the final transposed convolutional layer with a stride of 1 and a single kernel, producing an output dimension of (b, 256, 256, 1). The sigmoid function is used as the activation function, and the output layer then produces a new sample generated by the generator.

[0076] The discriminator structure includes an input layer, convolutional layers, a self-attention module, a flattened layer, and a dense layer. Layer 1 is the input layer, with the input image data having dimensions (b, 256, 256, 1), where b represents the batch size of the input data, and 256x256x1 represents the image size and number of channels. Layers 2 and 3 are convolutional layers with a kernel size of 5x5, 32 and 64 kernels respectively, and a stride of 2. These two convolutional layers reduce the image size and number of channels through feature extraction, outputting feature maps with dimensions of (b, 128, 128, 32) and (b, 64, 64, 64). The LeakyReLU function is used as the activation function for each layer to introduce non-linearity. The self-attention module in layer 4 outputs a feature map of size (b, 64, 64, 64). The feature map is fed into layer 5, and the output of layer 5 is fed into layer 6. Layers 5 and 6 are still convolutional layers with a kernel size of 5x5, and the number of kernels are 128 and 256 respectively, with a stride of 2 for both. These two convolutional layers further reduce the image size and number of channels through feature extraction, and the dimensions of the output feature maps are (b, 32, 32, 128) and (b, 16, 16, 256). Finally, the output of layer 6 is flattened into a one-dimensional vector by layer 7 (Flatten layer), and then fed into the Dense layer to calculate the authenticity of the sample and the probability that the sample belongs to a certain class.

[0077] The self-attention module proposed in this embodiment takes the feature map output by a certain convolutional layer as input, transforms the input feature map into three feature spaces, and obtains three feature space maps. One of the feature space maps is transposed and multiplied with another feature space map to obtain the correlation matrix for generating the attention map. The attention map is obtained based on the correlation matrix and the remaining feature space map. The self-attention feature map is obtained based on the attention map and the remaining feature space map. The weight of the self-attention feature map is set as the transition parameter, the weight of the feature map input to the self-attention module is 1, and the self-attention feature map and the feature map input to the self-attention module are weighted and summed to obtain the feature map output by the self-attention module.

[0078] The structure of the self-attention module is as follows: Figure 4 As shown, the process of obtaining the feature map output by the self-attention module is as follows:

[0079] First, the previous convolutional layer t∈R C×N The feature image t is transformed into three feature spaces f, g, and v. That is, f(t) = W f t, g(t) = W g t, v(t) = W v t represents three feature spaces obtained by multiplying the feature image t by different weight matrices, where f(t) and g(t) are used to calculate the correlation matrix b for generating the attention map. ij v(t) and b ij Calculate and generate attention map B j,i The specific formula for generating the attention map is as follows:

[0080] b ij =f(t) i ) T g(t j ), t∈R C×N (8)

[0081]

[0082] Where C is the number of channels and N is the number of feature positions. B j,i Attention map.

[0083] After calculating the attention of all elements in the feature image, an attention map is output. This map is then normalized using the Softmax function to obtain a normalized attention map. Finally, a matrix operation is performed between the normalized attention map and the feature space v(t). The output of the attention layer is d = (d1, d2, ..., dt). j , ..., d N )∈R C×N The expression d for the self-attention feature map is obtained. j As shown in equation (10):

[0084]

[0085] In generating self-attention feature map d j Then, the output of the attention layer is multiplied by the parameters and added back to the input feature map t. The final output of the attention module is o. i As shown in equation (11), where δ is the transition parameter, initially set to 0, allowing the model to learn from domain information and gradually assign weights to other distant feature details:

[0086] o i =δd i +t i (11)

[0087] Among them, o i For the feature map output by the self-attention module, δ is a learnable scalar that is initialized to 0. Introducing a learnable δ allows the network to initially rely on cues within the local neighborhood and then gradually learn to assign more weights to non-local evidence.

[0088] This embodiment employs the TTUR training strategy to train the malicious traffic sample generation model. TTUR is a training strategy for GANs designed to address the imbalance between generator and discriminator training. Its core idea is that the generator and discriminator use different learning rates and update frequencies, thus allowing them to update at different speeds during training.

[0089] Specifically, the TTUR strategy involves using a smaller learning rate and a larger update frequency for the generator, while using a larger learning rate and a smaller update frequency for the discriminator. This is done to make the generator more stable during updates and to keep pace with the discriminator's learning speed. In this embodiment, the learning rate ratio for the generator and discriminator is set to 3:1. For example, if the discriminator's learning rate is 0.0002, then the generator's learning rate is set to 0.000067. Furthermore, the discriminator's update frequency is set to three times that of the generator to ensure that the discriminator can learn the data distribution more quickly.

[0090] This strategy not only ensures that the generator and discriminator remain in balance during training, enabling the model to better learn the data distribution, but also shortens the training time and reduces training costs.

[0091] This embodiment adds an extra input layer to the discriminator and generator, using traffic labels as conditional constraints to supervise the training process of the generative model, making it suitable for multi-class imbalanced environments in the current network environment. A self-attention module is introduced into CWGAN-GP to help the generative model model long-distance, multi-level dependencies across image regions, adding higher weights to key features and more accurately imposing complex geometric constraints on the global image structure, making the generated malicious traffic samples more usable. At the same time, the computational cost of the attention vector is very low, and introducing the self-attention module does not significantly increase the model's computational cost. The TTUR training strategy is adopted to improve model stability and training speed while reducing training costs.

[0092] S4: Obtain the malicious traffic identification results through malicious traffic samples and a trained malicious traffic identification model.

[0093] Currently, Convolutional Neural Networks (CNNs) are a popular deep learning traffic classification algorithm. CNNs are suitable for processing Euclidean structured data such as images, excel at capturing local spatial correlations without human intervention, and can improve generalization ability by reducing the number of trainable parameters to avoid overfitting.

[0094] The malicious traffic identification model in this embodiment takes malicious traffic samples as input and malicious traffic identification results as output. It is constructed through a convolutional neural network. Each convolutional layer of the convolutional neural network includes a set of filters. The filters assign weights to the feature blocks in the feature map of the input convolutional layer and add bias vectors to the weighted features.

[0095] Specifically, the malicious traffic identification model disclosed in this embodiment is obtained by constructing a two-dimensional convolutional neural network (2D-CNN).

[0096] The main advantage of 2D-CNN is its ability to capture potential spatial continuity in malicious traffic sample images. Since the input malicious traffic sample images allow for the description of potential data patterns appearing at adjacent pixels, 2D-CNN allows for the utilization of specific continuity behaviors embedded in the feature space grid considered within the input region of the training data. Therefore, this embodiment introduces local filters and weight sharing on top of 2D-CNN.

[0097] Each convolutional layer contains a set of filters to process small local parts of the input. For example, given an image x, the k-th feature map at position (i, j) in the l-th convolutional layer is formed by the weight matrix of the k-th filter in the l-th convolutional layer. and bias vector The activation function σ() is used to determine the activation, and the specific formula is as follows:

[0098]

[0099] in, It is the input block centered at position (i, j) in the l-th layer, and * denotes the convolution function. Each possible position (i, j) shares the same input block. This reduces model complexity and makes it easier to train.

[0100] The malicious traffic identification model constructed in this embodiment omits the pooling layer, which has little impact on model performance, and includes 3 convolutional layers, 2 dropout layers, and 3 fully connected layers in the network structure.

[0101] In this embodiment, the malicious traffic sample obtained in S3 is input into the trained malicious traffic identification model, and the malicious traffic identification result is output.

[0102] The specific process of obtaining a trained malicious traffic identification model is as follows:

[0103] Acquire large amounts of known types of malicious traffic;

[0104] Preprocess known categories of malicious traffic according to step S2 to obtain feature images of known categories of malicious traffic;

[0105] According to S3, malicious traffic samples of known categories are obtained by using feature images of known categories of malicious traffic and a trained malicious traffic sample generation model.

[0106] A dataset is constructed using malicious traffic samples of known categories. The malicious traffic identification model is then trained using this dataset. Once training is complete, a trained malicious traffic identification model is obtained.

[0107] The dataset processed by the generative model is used to train a 2D-CNN model. The higher layers of this model use a wider range of filters, adapted to lower-resolution inputs to handle more complex input components. After hidden layer operations, the output layer uses softmax as the activation function to handle classification tasks and accurately identify malicious traffic.

[0108] The encrypted malicious traffic identification method disclosed in this embodiment establishes a self-attention module in the malicious traffic sample generation model to help the model effectively model complex dependencies with low computational cost, so as to generate more refined and reliable malicious traffic samples. When identifying the malicious traffic samples, it can improve the accuracy of malicious traffic identification. For the preprocessing of malicious traffic, there is no need to unify the data packet length, thereby avoiding information loss and information redundancy problems during the preprocessing process, further ensuring the accuracy of malicious traffic identification.

[0109] Example 2

[0110] This embodiment discloses an encrypted malicious traffic identification system for various types of imbalanced data characteristics, including:

[0111] The traffic data acquisition module is used to acquire malicious network traffic data;

[0112] The preprocessing module is used to preprocess malicious network traffic data to obtain characteristic images of the malicious traffic;

[0113] The traffic sample generation module is used to obtain malicious traffic samples through the feature image of malicious traffic and the trained malicious traffic sample generation model. The malicious traffic sample generation model is constructed using the CWGAN-GP model. Self-attention modules are added to both the generator and the discriminator. The self-attention module takes the feature map output by the convolutional layer as input, obtains the attention map of the feature map input to the self-attention module, determines the self-attention feature map based on the attention map, and performs a weighted summation of the self-attention feature map and the feature map input to the self-attention module to obtain the feature map output by the self-attention module. The feature map output by the self-attention module is then input into the next convolutional layer.

[0114] The traffic identification module is used to obtain malicious traffic identification results through malicious traffic samples and a trained malicious traffic identification model.

[0115] Example 3

[0116] In this embodiment, an electronic device is disclosed, including a memory and a processor, as well as computer instructions stored in the memory and running on the processor. When the processor executes the computer instructions, it completes the steps described in the encrypted malicious traffic identification method for multiple types of imbalanced data characteristics disclosed in Embodiment 1.

[0117] Example 4

[0118] In this embodiment, a computer-readable storage medium is disclosed for storing computer instructions, which, when executed by a processor, complete the steps described in the encrypted malicious traffic identification method for multiple types of imbalanced data characteristics disclosed in Embodiment 1.

[0119] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A method for identifying encrypted malicious traffic based on the characteristics of various types of imbalanced data, characterized in that, include: Obtaining malicious network traffic data; Preprocess network malicious traffic data to obtain characteristic images of malicious traffic; The process of preprocessing malicious network traffic data is as follows: Segment malicious network traffic data to obtain multiple session files; Malicious traffic data in PCAP format is segmented into multiple session files based on the 5-tuple; the 5-tuple includes the source IP address, source port, destination IP address, destination port, and transport layer protocol; Perform traffic scrubbing on all session files to obtain the traffic-scrubbed session files; The obtained session files are filtered to remove duplicate and useless packets, resulting in a filtered session file. Remove useless information from the header fields of the filtered session file to prevent the model from overfitting to such information; Useless information includes noise, data link layer information, and IP addresses; Using a Markov model, feature images of session files after traffic scrubbing are generated, which are feature images of malicious traffic. Malicious traffic samples are obtained by using the feature images of malicious traffic and a trained malicious traffic sample generation model. The malicious traffic sample generation model is constructed using the CWGAN-GP model. Self-attention modules are added to both the generator and the discriminator. The self-attention module takes the feature map output by the convolutional layer as input, obtains the attention map of the feature map input to the self-attention module, determines the self-attention feature map based on the attention map, and performs a weighted summation of the self-attention feature map and the feature map input to the self-attention module to obtain the feature map output by the self-attention module. The feature map output by the self-attention module is then input into the next convolutional layer. The TTUR training strategy was used to train the malicious traffic sample generation model; The malicious traffic identification model takes malicious traffic samples as input and malicious traffic identification results as output, and is constructed through a convolutional neural network. Each convolutional layer in a convolutional neural network includes a set of filters. These filters assign weights to feature blocks in the feature map of the input convolutional layer and add bias vectors to the weighted features. Malicious traffic identification results are obtained by using malicious traffic samples and a trained malicious traffic identification model.

2. The encrypted malicious traffic identification method for multiple types of imbalanced data as described in claim 1, characterized in that, The weights of the self-attention feature map are set as transition parameters, and the weights of the feature maps input to the self-attention module are set to 1. The self-attention feature map and the feature map input to the self-attention module are weighted and summed to obtain the feature map output by the self-attention module.

3. The encrypted malicious traffic identification method for multiple types of imbalanced data as described in claim 1, characterized in that, The self-attention module takes the feature map output by a convolutional layer as input, transforms the input feature map into three feature spaces, and obtains three feature space maps; it transposes one of the feature space maps and multiplies it with another feature space map to obtain the correlation matrix for generating the attention map; and it obtains the attention map based on the correlation matrix and the remaining feature space map.

4. An encrypted malicious traffic identification system for various types of imbalanced data, characterized in that: include: The traffic data acquisition module is used to acquire malicious network traffic data; The preprocessing module is used to preprocess malicious network traffic data to obtain characteristic images of the malicious traffic; The process of preprocessing malicious network traffic data is as follows: Segment malicious network traffic data to obtain multiple session files; Malicious traffic data in PCAP format is segmented into multiple session files based on the 5-tuple; the 5-tuple includes the source IP address, source port, destination IP address, destination port, and transport layer protocol; Perform traffic scrubbing on all session files to obtain the traffic-scrubbed session files; The obtained session files are filtered to remove duplicate and useless packets, resulting in a filtered session file. Remove useless information from the header fields of the filtered session file to prevent the model from overfitting to such information; Useless information includes noise, data link layer information, and IP addresses; Using a Markov model, feature images of session files after traffic scrubbing are generated, which are feature images of malicious traffic. The traffic sample generation module is used to obtain malicious traffic samples through the feature image of malicious traffic and the trained malicious traffic sample generation model. The malicious traffic sample generation model is constructed using the CWGAN-GP model. Self-attention modules are added to both the generator and the discriminator. The self-attention module takes the feature map output by the convolutional layer as input, obtains the attention map of the feature map input to the self-attention module, determines the self-attention feature map based on the attention map, and performs a weighted summation of the self-attention feature map and the feature map input to the self-attention module to obtain the feature map output by the self-attention module. The feature map output by the self-attention module is then input into the next convolutional layer. The TTUR training strategy was used to train the malicious traffic sample generation model; The malicious traffic identification model takes malicious traffic samples as input and malicious traffic identification results as output, and is constructed through a convolutional neural network. Each convolutional layer in a convolutional neural network includes a set of filters. These filters assign weights to feature blocks in the feature map of the input convolutional layer and add bias vectors to the weighted features. The traffic identification module is used to obtain malicious traffic identification results through malicious traffic samples and a trained malicious traffic identification model.

5. An electronic device, characterized in that, It includes a memory and a processor, as well as computer instructions stored in the memory and running on the processor, which, when executed by the processor, complete the steps of the encrypted malicious traffic identification method for the characteristics of multiple types of imbalanced data as described in any one of claims 1-3.

6. A computer-readable storage medium, characterized in that, Used to store computer instructions, which, when executed by a processor, complete the steps of the encrypted malicious traffic identification method for the characteristics of multiple types of imbalanced data as described in any one of claims 1-3.