Network traffic class balancing method, system, and storage medium based on wgan-gp and optimization thereof

CN122764633APending Publication Date: 2026-09-15SHANXI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610977380.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-02
Publication Date
2026-09-15

AI Technical Summary

Technical Problem

[0007]针对现有WGAN-GP模型在网络流量类别平衡任务中存在的真实少数类样本历史特征利用不足、生成样本质量不高以及超参数人工调节效率低的问题,本发明提供了一种基于WGAN-GP及其优化的网络流量类别平衡方法、系统及存储介质

Benefits of technology

(1)本发明通过生成少数类网络流量样本,对原始训练集中样本数量不足的少数类类别进行补充,缓解了网络流量数据集中类别不平衡的问题,使异常检测模型能够学习到更加充分的少数类攻击特征;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122764633A_ABST
    Figure CN122764633A_ABST
Patent Text Reader

Abstract

The application discloses a network traffic category balancing method based on WGAN-GP and optimization thereof, and belongs to the technical field of network security. The method carries out pretreatment on network traffic data and determines minority classes; a WGAN-GP-B-M model composed of a generator module, a discriminator module, a memory module and a Bayesian optimization module is constructed and trained; the memory module stores the real minority class sample features extracted by the discriminator module and is dynamically updated, the generator module fuses random noise and historical features searched to generate minority class samples, and the Bayesian optimization module searches for optimal hyperparameters; the model is trained for each minority class to generate samples, and a balanced training set is constructed by combining the original training set; and an anomaly detection model is trained by using the balanced training set and unknown traffic is classified and detected. The application improves the authenticity and diversity of minority class samples and improves the recognition ability of minority class attack traffic.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of network security technology, specifically relating to a network traffic category balancing method, system, and storage medium based on WGAN-GP and its optimization. More particularly, it relates to a method that extracts high-dimensional features of real minority class network traffic samples through a discriminator intermediate layer, and utilizes a memory module for similarity retrieval, momentum updating, and Top-order matching. A method for generating minority class network traffic samples and constructing a class-balanced training set by invoking historical features and combining them with Bayesian optimization search for training hyperparameters. Background Technology

[0002] With the development of the Internet, the Internet of Things, cloud computing, and big data technologies, the scale of traffic data generated in network systems is constantly expanding, and network attacks are also showing a trend towards greater complexity, diversity, and concealment. Network traffic anomaly detection is an important technical means in network security protection. Its basic goal is to identify abnormal traffic that deviates from normal network behavior patterns by analyzing the behavioral characteristics of network traffic data, thereby discovering potential network attacks.

[0003] In real-world network environments, network traffic data often exhibits a significant class imbalance. The number of normal traffic samples is typically far greater than that of abnormal traffic samples. Furthermore, within abnormal traffic, there are substantial differences in sample numbers between different attack types. The limited number of minority attack traffic samples makes it easier for anomaly detection models to bias towards the majority class during training, hindering their ability to fully learn the characteristic patterns of minority abnormal traffic. This results in lower identification rates and higher false negative rates for minority attack traffic.

[0004] Existing class balancing methods mainly include random oversampling, undersampling, mixed sampling, SMOTE-like methods, and data augmentation methods based on generative adversarial networks. Random oversampling balances the dataset by replicating minority class samples, but it is prone to causing model overfitting; undersampling balances the dataset by reducing the number of majority class samples, but it may lose effective information from the majority class samples; SMOTE-like methods generate minority class samples through sample interpolation, but when the class boundaries are complex or the samples are noisy, they are prone to generating samples that do not conform to the true distribution.

[0005] While conventional generative adversarial networks (GANs) can generate new samples, they are prone to problems such as vanishing gradients, training instability, and pattern collapse during training. WGAN-GP improves the stability of GAN training through Wasserstein loss and gradient penalty terms, but its generator relies heavily on random noise when generating samples, failing to adequately utilize historical features of real minority class samples, and there is still room for improvement in the diversity and distribution consistency of generated samples. Furthermore, hyperparameters such as generator learning rate, discriminator learning rate, gradient penalty coefficient, and memory regularization weights have a significant impact on model training performance, and manual parameter tuning is inefficient and unstable. In addition, existing memory-enhanced generative models typically use static memory storage, making it difficult to dynamically update the historical features of minority class network traffic samples based on the training process. When the memory contains a large number of redundant or outdated features, the generator is susceptible to interference from invalid historical information, leading to a deviation between the generated samples and the real minority class network traffic distribution. Therefore, constructing a memory mechanism that can dynamically update, retrieve, and utilize high-dimensional features of real minority class samples is a crucial technical challenge for improving the quality of minority class network traffic sample generation.

[0006] The aforementioned issues limit the effectiveness of existing generative models in network traffic category balancing tasks. Summary of the Invention

[0007] To address the problems of insufficient utilization of historical features of real minority class samples, low quality of generated samples, and low efficiency of manual adjustment of hyperparameters in existing WGAN-GP models for network traffic class balancing tasks, this invention provides a network traffic class balancing method, system, and storage medium based on WGAN-GP and its optimization.

[0008] The method of this invention constructs a network traffic class balance model WGAN-GP-BM using a generator module, a discriminator module, a memory module, and a Bayesian optimization module. The memory module stores the high-dimensional features of real minority class samples extracted from the intermediate layer of the discriminator module, enabling the generator module to generate minority class network traffic samples by combining random noise vectors and historical memory features. The Bayesian optimization module searches for the optimal hyperparameter combination during model training, thereby improving the quality of generated samples and the stability of model training.

[0009] To achieve the above objectives, the present invention employs the following technical solutions: The network traffic class balancing method based on WGAN-GP and its optimizations includes the following steps: The first step is network traffic data preprocessing. Feature extraction is performed on the network traffic dataset, or pre-extracted network traffic features are read. The network traffic features are then cleaned, missing values ​​are handled, features are encoded, and normalized to obtain network traffic feature data for model training. The number of samples in each category is counted based on the sample labels to determine the minority network traffic categories.

[0010] The second step: Construction and training of the class-balanced model WGAN-GP-BM.

[0011] Model construction: A WGAN-GP-BM network traffic class balancing model is built, consisting of a generator module, a discriminator module, a memory module, and a Bayesian optimization module.

[0012] The generator module receives random noise vectors and historical memory features provided by the memory module to generate minority class network traffic samples. The discriminator module distinguishes between real minority class network traffic samples and generated minority class network traffic samples, and outputs a sample discrimination score. The memory module stores high-dimensional features of real minority class network traffic samples extracted by the intermediate layer of the discriminator module and provides historical memory features to the generator module. The Bayesian optimization module searches for the optimal combination of hyperparameters during model training. The hyperparameters include the learning rate of the generator module, the learning rate of the discriminator module, the gradient penalty coefficient, the momentum coefficient of the memory module, and the weight of the memory regularization term.

[0013] Model training: The training rounds, memory capacity, gradient penalty coefficient, memory regularization term weights, and Bayesian optimization search space are set in advance. Real minority class network traffic samples and random noise vectors are input into the model for adversarial training.

[0014] Training process: Real minority class network traffic samples and generated minority class network traffic samples are input into the discriminator module for feature extraction and discrimination. The high-dimensional features output from the intermediate layer of the discriminator module are written into the memory module, and the memory is updated according to the feature similarity and momentum update strategy. Random noise vectors are input into the generator module, which retrieves them from the memory module. A related historical memory feature ( (A preset positive integer) is used to fuse historical memory features with random noise features through a multi-head attention mechanism to generate minority class network traffic samples; the discriminator module is used to distinguish between real minority class samples and generated minority class samples respectively, and the discriminator module parameters are updated by the discriminator Wasserstein loss and gradient penalty term, and the generator module parameters are updated by the generator Wasserstein loss and memory regularization term; at the same time, the hyperparameter combination is updated by the Bayesian optimization module according to the validation loss. When the model reaches the preset number of training rounds or meets the convergence condition, training stops and the optimal generator module is saved.

[0015] The third step: Generating minority class network traffic samples and constructing a class-balanced dataset.

[0016] For each minority network traffic category, extract the real minority network traffic sample set for that category, and train the corresponding WGAN-GP-BM network traffic category balancing model for that category. Use the trained optimal generator module for that category to generate the corresponding number of minority network traffic samples, and merge the generated minority network traffic samples with the original network traffic training set to obtain the category-balanced network traffic training set.

[0017] Step 4: Generate a sample quality and detection performance evaluation.

[0018] The distributional and numerical differences between the generated minority class network traffic samples and the real minority class network traffic samples are evaluated using KL divergence and mean absolute error. The network traffic anomaly detection performance after balancing the categories is evaluated using one or more of accuracy, precision, recall, and F1 score.

[0019] Step 5: Network traffic anomaly detection.

[0020] The network traffic training set after category balance is input into the network traffic anomaly detection model for training, and the trained network traffic anomaly detection model is used to classify and detect unknown network traffic samples.

[0021] Furthermore, the specific operation of the first step is as follows: Step 1.1: Data Feature Acquisition: Analyze the collected network traffic data to obtain relevant feature information of the network traffic samples. At the same time, determine the true labels of normal traffic and abnormal traffic based on the dataset labels or attack traffic labeling information, which will serve as the input for Step 1.2.

[0022] Step 1.2: Data cleaning: The network traffic feature information obtained in Step 1.1 is cleaned by deleting invalid samples, deleting abnormal format samples, and handling missing values ​​to obtain cleaned network traffic feature data.

[0023] Step 1.3: Feature Encoding and Normalization: Encode the non-numerical features in the network traffic data and normalize or standardize the numerical features to bring the network traffic features of different dimensions into a uniform numerical range.

[0024] Step 1.4: Minority Classification: Based on the network traffic sample labels, count the number of samples in each category. Determine the categories with a sample number lower than the preset proportion threshold or significantly less than the number of samples in the majority category as minority network traffic categories, and extract the minority network traffic sample set as the input for the second step.

[0025] Furthermore, the specific operation of the second step is as follows: ① Data Sample Loading: A data loading module is built using the deep learning framework PyTorch to package the minority class network traffic samples obtained in the first step into a batch of a preset size, which serves as the input to the generator and discriminator modules. During training, a batch of real minority class network traffic samples is sampled from the real minority class network traffic sample set, and at the same time, random noise vectors of the same batch size are sampled from a standard Gaussian distribution, which serve as the input to the generator module.

[0026] ② Generator Module Construction: The generator module generates minority class network traffic samples based on random noise vectors and historical memory features provided by the memory module. First, the random noise vector is input into the first fully connected layer to obtain the first latent feature representation. Then, batch normalization and LeakyReLU activation are sequentially applied to the first latent feature representation to obtain the first latent query features. Subsequently, the memory module retrieves relevant features from the memory database. The system uses historical memory features and a multi-head attention mechanism to fuse these features with a first latent query feature representation. The fused features are then input into a second fully connected layer to obtain a second latent feature representation, which is further processed by batch normalization and LeakyReLU activation. Finally, the second latent feature representation is input into a third fully connected layer, mapping it to the same feature dimensions as the target minority class real network traffic sample, and outputting the generated minority class network traffic sample through a Tanh activation function.

[0027] ③ Discriminator Module Construction: The discriminator module distinguishes between real minority class network traffic samples and generated minority class network traffic samples, and outputs a sample discrimination score. First, real minority class network traffic samples and generated minority class network traffic samples are input into the first fully connected layer to obtain the first discriminant feature; then, spectral normalization and LeakyReLU activation are applied to the first discriminant feature. Subsequently, the first discriminant feature is input into the second fully connected layer to obtain the second discriminant feature, and spectral normalization and LeakyReLU activation are applied again after the second fully connected layer. The second discriminant feature serves two purposes: firstly, it is used as a high-dimensional feature input to the memory module for feature storage and updating in the memory bank; secondly, it is input into the third fully connected layer, outputting a sample discrimination score with dimension 1. The discriminator module calculates the total loss of the discriminator module based on the real minority class sample discrimination score, the generated minority class sample discrimination score, and the gradient penalty term, and updates the discriminator module parameters through backpropagation.

[0028] ④ Memory module construction: The memory module includes feature extraction, feature storage, momentum update, diversity-aware replacement, feature retrieval, and feature fusion.

[0029] First, the second discriminant feature output from the second fully connected layer of the discriminator module is used as the high-dimensional feature of the real minority class network traffic samples and written into the memory. The memory is managed using a circular buffer structure, and its capacity is set to a preset value. When the memory is not full, the newly extracted high-dimensional feature is directly written into the memory; when the memory is full, the cosine similarity between the newly extracted feature and the existing features in the memory is calculated, and the memory is dynamically updated according to the momentum update strategy and the diversity-aware replacement strategy to maintain the diversity of feature distribution in the memory.

[0030] The momentum update strategy is as follows:

[0031] The cosine similarity is:

[0032] The diversity-aware replacement strategy is as follows: when the memory bank is full, if the maximum cosine similarity between the new feature and the existing features in the memory bank is greater than a preset threshold, then the new feature is replaced by momentum fusion with the most similar historical feature; otherwise, the new feature is written into the memory bank and replaces the earliest stored historical feature in the memory bank.

[0033] When the generator module generates samples, the memory module retrieves the most relevant features from the memory bank based on an approximate nearest neighbor search method. One characteristic of historical memory: Furthermore, a multi-head attention mechanism is used to fuse the retrieved historical memory features with the random noise features of the generator module.

[0034] In the formula, This represents the high-dimensional features extracted by the intermediate layer of the discriminator module; This represents the intermediate feature extraction mapping of the discriminator module; This represents the true minority class samples; This represents the memory after step t; This represents the memory after the (t-1)th step; Indicates the momentum coefficient; This represents the i-th historical memory feature stored in the memory bank; This indicates the Top- retrieved results. Historical memory feature matrix; This represents the query features provided by the generator and is the output of the intermediate layer of the generator module. , , These represent the query matrix, key matrix, and value matrix in the multi-head attention mechanism, respectively. Let W be a random noise vector; W represents a learnable matrix. Indicates the attention head dimension; Indicates transpose; This represents the fused feature vector. This represents the approximate nearest neighbor search algorithm. The Softmax function represents the normalized exponential function, which is used to convert the correlation scores between multiple historical features and the current query feature into attention weights.

[0035] ⑤ Bayesian optimization module construction: The Bayesian optimization module is used to search for the optimal combination of the learning rate of the generator module, the learning rate of the discriminator module, the gradient penalty coefficient, the momentum coefficient of the memory module, and the weight of the memory regularization term during model training.

[0036] First, a hyperparameter search space is defined, including the learning rate of the generator module, the learning rate of the discriminator module, the gradient penalty coefficient, the momentum coefficient of the memory module, and the weight of the memory regularization term. Then, a probabilistic mapping relationship between the hyperparameter combination and the model validation loss is established using a Gaussian process surrogate model. Next, the next set of candidate hyperparameter combinations is selected based on the acquisition function. Subsequently, the WGAN-GP-BM network traffic class balance model is trained using the candidate hyperparameter combinations, and the corresponding validation loss is calculated. Finally, the Gaussian process surrogate model is updated based on the validation loss. The above process is repeated until the maximum number of optimization iterations is reached or the preset convergence condition is met, and the optimal hyperparameter combination is output.

[0037] The Bayesian optimization module operates by combining outer hyperparameter search with inner model training. During outer optimization, the module proposes candidate hyperparameter combinations based on a Gaussian process surrogate model and a data acquisition function. During inner training, the WGAN-GP-BM network traffic class balance model is trained based on these candidate hyperparameter combinations, and the Gaussian process surrogate model is updated according to the validation loss or generated sample quality evaluation results obtained after training. This process of candidate hyperparameter generation, model training, validation loss calculation, and surrogate model update is repeated until a preset number of optimization iterations is reached or the convergence condition is met, at which point the optimal hyperparameter combination is output.

[0038] The Gaussian process surrogate model in the Bayesian optimization process is:

[0039] The mean function is:

[0040] The kernel function of the Gaussian process surrogate model is the Matérn 5 / 2 kernel function:

[0041]

[0042]

[0043] The acquisition function is the acquisition function that is expected to be improved:

[0044]

[0045]

[0046]

[0047]

[0048] In the formula, This represents the current combination of hyperparameters to be evaluated. This represents another set of hyperparameter combinations. This represents the validation loss or objective function value corresponding to the hyperparameter combination. Represents a Gaussian process. Represents the mean function, This represents a kernel function used to measure the correlation between two sets of hyperparameters. Indicates the learning rate. Represents the gradient penalty coefficient. Indicates the weight of the memory module. This indicates scaling the Mahalanobis distance. Let Λ denote the signal variance, Λ denote the anisotropic length scaling matrix, and diag denote the diagonal matrix. This indicates the correlation length of each dimension. This indicates a desire to improve the acquisition function. This represents the weighting coefficients that change with iteration. and Let represent the desired improvements for local search and global search, respectively. This represents the smallest observed value of the objective function; This represents the mean of the Gaussian process's predictions for the current hyperparameter combination; This indicates the uncertainty in the Gaussian process's prediction of this location; Indicates the exploration intensity coefficient; This indicates the current iteration number of the Bayesian optimization; This represents the maximum number of iterations for Bayesian optimization.

[0049] The validation loss is measured by the KL divergence or mean absolute error between the minority class network traffic samples generated by the generator module on the validation set and the real minority class network traffic samples on the validation set.

[0050] ⑥ Model Training: First, initialize the generator module, discriminator module, memory module, optimizer, and Bayesian optimization module. Then, sample a batch of real minority class network traffic samples from the real minority class network traffic sample set, and sample random noise vectors of the same batch size from a standard Gaussian distribution. Input the random noise vectors into the generator module to obtain generated minority class network traffic samples. Input the real minority class network traffic samples and generated minority class network traffic samples into the discriminator module respectively, and calculate the discrimination scores of the real minority class samples and the generated minority class samples. Construct interpolation samples based on the real minority class samples and generated minority class samples, and calculate the gradient penalty term. Calculate the total loss of the discriminator module based on the discriminator's Wasserstein loss and the gradient penalty term, and update the discriminator module parameters. Extract high-dimensional features of the real minority class samples using the intermediate layers of the discriminator module, and update the memory according to the rules of the memory module. Then, retrieve the data from the memory. The relevant historical memory features are fused with the latent features of the generator module. The total loss of the generator module is calculated based on the feedback results of the discriminator module on the generation of minority class samples and the memory regularization term. The parameters of the generator module are then updated. The above training process is repeated until the model reaches the preset training rounds or meets the preset convergence condition. The optimal generator module is then saved.

[0051] The calculation formula for the interpolated sample is: ,in, Indicates the interpolated sample; This represents a sample of real minority network traffic. ε represents the minority class network traffic samples generated by the generator module; ε represents the random coefficients sampled from the preset distribution.

[0052] The total loss function of the discriminator module is:

[0053] in, This represents the total loss of the discriminator module; For the discriminator module; This represents the true minority class sample distribution; The distribution of minority class samples generated by the generator module; For the interpolated sample distribution; This represents the true minority class samples; This represents the minority class samples generated by the generator module; For interpolation samples; This is the gradient penalty coefficient; For expectation calculation; This represents the L2 norm.

[0054] It should be noted that the first half of the total loss function of the discriminator... The Wasserstein loss for the discriminator is used to approximate the Wasserstein distance between the generated and real distributions by reflecting the difference in scores between the discriminator module for real and generated minority class samples. Compared to traditional GANs using JS divergence, the Wasserstein distance exhibits better continuity and smoothness, providing a relatively stable gradient signal to the generator module even with low overlap between the generated and real distributions. This alleviates the gradient vanishing and mode collapse problems that commonly occur during the training of traditional GANs. The gradient penalty term... This is used to constrain the discriminator module to satisfy the 1-Lipschitz condition in the interpolation space between real and generated samples, making the discriminator gradient norm at the interpolated samples close to 1, thereby improving the stability of the adversarial training process.

[0055] The total loss function of the generator module is:

[0056] in, This represents the total loss of the generator module; For generator modules; For the discriminator module; It is a random noise vector; The noise distribution is random. To generate a minority class sample distribution; The distribution of historical features in the memory module; To remember the weights of the regularization terms; This is the KL divergence constraint term.

[0057] It should be noted that the first half of the formula The generator module uses the Wasserstein loss to minimize this term, enabling the discriminator module to assign a more accurate discrimination score to the generated minority class network traffic samples compared to the real samples. This, in turn, helps the generated samples gradually approximate the distribution of the real minority class samples. (The latter part of the formula...) The memory regularization term is used to ensure that the distribution of generated samples is consistent with the memory distribution represented by the high-dimensional features of real minority class samples stored in the memory bank, thereby guiding the generator module to make full use of historical real feature information and improve the quality, authenticity and distribution consistency of generated samples.

[0058] Furthermore, the specific operation of the third step is as follows: Step 3.1: Count the number of samples of each category in the original network traffic training set and determine the target number of samples for category balance.

[0059] Step 3.2: Based on the target number of samples for each minority network traffic category and the original number of samples for each minority network traffic category, determine the number of samples that need to be generated for each minority network traffic category.

[0060] Step 3.3: For each minority network traffic category, extract the real minority network traffic sample set corresponding to that category, and use the real minority network traffic sample set to train the corresponding WGAN-GP-BM network traffic category balance model to obtain the optimal generator module corresponding to that category; use the optimal generator module to generate the corresponding number of minority network traffic samples with corresponding category labels.

[0061] It should be noted that the network traffic characteristics of different attack types are completely different (e.g., DDoS attacks are characterized by a surge in traffic, brute-force attacks by an abnormal frequency of login failures, and port scanning by accessing a large number of different ports), resulting in significant differences in the data distribution across categories. Therefore, in step 3.3, an independent WGAN-GP-BM network traffic category balancing model is trained for each minority network traffic category. Each model has independent initialization parameters, an independent memory, and an independent Bayesian optimization process to ensure that the generated samples for each category accurately reflect the characteristic distribution of the real samples for that category. In a preferred embodiment, the training of WGAN-GP-BM network traffic category balancing models for multiple minority network traffic categories is performed in parallel to improve overall processing efficiency.

[0062] Step 3.4: Add the generated minority class network traffic samples to the original network traffic training set, and randomly shuffle the merged training set to obtain a class-balanced network traffic training set.

[0063] Furthermore, the calculation formulas for generating the sample quality evaluation index and the network traffic anomaly detection performance evaluation index in the fourth step are as follows:

[0064]

[0065]

[0066]

[0067]

[0068]

[0069] in, Represents the true sample distribution With the distribution of generated samples Between Divergence; Represents the true sample distribution In the Probability values ​​over an interval or category; Represents the distribution of generated samples In the Probability values ​​over an interval or category; Indicates the mean absolute error; Representing the first real sample One eigenvalue; Indicates the number of generated samples One eigenvalue; Indicates the number of samples or features; Indicates accuracy; This indicates the number of abnormal traffic instances that were correctly detected as abnormal. This indicates the number of normal traffic flows that were incorrectly detected as abnormal. This indicates the number of abnormal traffic flows that were incorrectly detected as normal. This indicates the number of normal traffic flows that were correctly detected as normal. Indicates accuracy; Indicates recall rate; This represents the F1 score.

[0070] It should be noted that the KL divergence and mean absolute error (MAE) in the fourth step are not only used to evaluate the quality of the final generated samples, but also, during the Bayesian optimization process, the validation loss is measured by the KL divergence or mean absolute error between the minority class network traffic samples generated by the generator module on the validation set under the current hyperparameter combination and the real minority class network traffic samples on the validation set. The Bayesian optimization module uses this validation loss as the objective function to guide the hyperparameter search process.

[0071] Furthermore, the specific operation of the fifth step is as follows: Step 5.1: Input the class-balanced network traffic training set into the network traffic anomaly detection model for training.

[0072] Step 5.2: Evaluate the performance of the trained network traffic anomaly detection model using a validation set or test set independent of the generated samples.

[0073] Step 5.3: Input the network traffic sample to be detected into the trained network traffic anomaly detection model to obtain the predicted category of the network traffic sample.

[0074] The network traffic anomaly detection model employs a pre-defined multi-classifier architecture, including but not limited to any one of the following: tree-based ensemble models (such as random forests and XGBoost), kernel-based classification models (such as support vector machines), or deep neural network-based classification models (such as multilayer perceptrons (MLP), convolutional neural networks (CNN), and recurrent neural networks (LSTM). This invention does not limit the specific type of anomaly detection model; this model serves only as a downstream task carrier for verifying the effectiveness of the class balancing method of this invention.

[0075] In addition, the present invention also provides a network traffic category balancing system based on WGAN-GP and its optimization, including: a data preprocessing module, a model building and training module, a sample generation and dataset construction module, an evaluation module, and an anomaly detection module.

[0076] The data preprocessing module is used to preprocess the network traffic dataset to obtain network traffic feature data, and to count the number of samples in each category based on the sample labels to determine the minority network traffic categories. The model building and training module is used to construct and train a WGAN-GP-BM network traffic class balance model consisting of a generator module, a discriminator module, a memory module, and a Bayesian optimization module. The generator module receives random noise vectors and historical memory features provided by the memory module to generate minority class network traffic samples. The discriminator module distinguishes between real minority class network traffic samples and generated minority class network traffic samples, and outputs a sample discrimination score. The memory module stores high-dimensional features of real minority class network traffic samples extracted from the intermediate layers of the discriminator module and provides historical memory features to the generator module. The Bayesian optimization module searches for the optimal combination of hyperparameters during training, including the generator module learning rate, discriminator module learning rate, gradient penalty coefficient, memory module momentum coefficient, and memory regularization term weights. Among them, the sample generation and dataset construction module is used to train the corresponding WGAN-GP-BM network traffic class balancing model for each minority network traffic class, generate a corresponding number of minority network traffic samples using the trained optimal generator module, and merge the generated minority network traffic samples with the original network traffic training set to obtain the class-balanced network traffic training set. The evaluation module is used to evaluate the distribution and numerical differences between the generated minority network traffic samples and the real minority network traffic samples using KL divergence and mean absolute error. It evaluates the network traffic anomaly detection performance after balancing the evaluation categories using one or more of accuracy, precision, recall and F1 score. The anomaly detection module is used to input the class-balanced network traffic training set into the network traffic anomaly detection model for training, and to use the trained network traffic anomaly detection model to classify and detect unknown network traffic samples.

[0077] Finally, the present invention also provides a non-transitory computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned network traffic category balancing method based on WGAN-GP and its optimization.

[0078] The beneficial effects of this invention are: (1) This invention supplements the minority class categories in the original training set that are insufficient in number by generating minority class network traffic samples, thereby alleviating the problem of class imbalance in the network traffic dataset and enabling the anomaly detection model to learn more sufficient minority class attack features. (2) Based on WGAN-GP, this invention introduces a memory module to store the high-dimensional features of real minority class samples extracted by the intermediate layer of the discriminator into the memory bank, and dynamically maintains the memory features through momentum update and similarity retrieval mechanisms, so that the generator can combine random noise and historical memory features to generate minority class samples that are closer to the real distribution. (3) This invention utilizes Wasserstein loss and gradient penalty terms to improve the stability of the generative adversarial training process and reduce the risk of gradient vanishing, gradient exploding and mode collapse during the training of ordinary generative adversarial networks; (4) This invention introduces hyperparameters such as the learning rate of the Bayesian optimization module search generator, the learning rate of the discriminator, the gradient penalty coefficient, the momentum coefficient of the memory module, and the weight of the memory regularization term, which reduces the cost of manual parameter tuning and improves the efficiency of model training and the stability of parameter selection. (5) The class-balanced training set generated by the present invention can be used to train the network traffic anomaly detection model, improve the model's ability to identify minority classes of abnormal traffic, and thus improve the overall performance of network traffic anomaly detection. Attached Figure Description

[0079] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0080] Figure 1 is an overall flowchart of the network traffic category balancing method based on WGAN-GP and its optimization in this invention; Figure 2 is a diagram of the overall architecture of the WGAN-GP-BM category balance model in this invention; Figure 3 is a structural diagram of the generator module in this invention; Figure 4 is a structural diagram of the discriminator module in this invention; Figure 5 is a flowchart of the feature storage, momentum update and Top-K retrieval process of the memory module in this invention; Figure 6 is a flowchart of the hyperparameter search process of the Bayesian optimization module in this invention; Figure 7 is a flowchart of the class-balanced dataset construction and anomaly detection process in this invention. Detailed Implementation

[0081] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0082] like Figure 1 As shown, this invention first preprocesses network traffic data to obtain network traffic feature data; then, it counts the number of samples in each category based on the category labels to determine the minority network traffic category; next, it constructs a WGAN-GP-BM network traffic category balancing model and performs adversarial training through a generator module, a discriminator module, a memory module, and a Bayesian optimization module; after training, it uses the optimal generator module to generate minority network traffic samples and adds the generated samples to the original training set to construct a category-balanced network traffic training set; finally, it uses the category-balanced network traffic training set to train a network traffic anomaly detection model and performs classification and detection on unknown network traffic samples.

[0083] The network traffic class balancing method based on WGAN-GP and its optimization, which relates to this invention, mainly includes five steps, each of which is described in detail below. For ease of explanation, the real minority class samples involved in this invention refer to the real minority class network traffic categories, and the generated minority class samples refer to the generated minority class network traffic categories.

[0084] Step 1: Network Traffic Data Preprocessing

[0085] Step 1.1: Data Feature Acquisition. The collected pcap format network traffic data is parsed to extract relevant feature information from the network traffic samples, including but not limited to: source IP address, destination IP address, source port, destination port, protocol type, packet length, number of packets, flow duration, and flag information. Alternatively, network traffic features already extracted from publicly available network traffic datasets can be directly read. Simultaneously, the true labels for normal and abnormal traffic are determined based on dataset labels or attack traffic annotation information.

[0086] Step 1.2: Data Cleaning. The network traffic feature information obtained in Step 1.1 is cleaned by removing invalid samples (e.g., samples with all-zero features), removing samples with abnormal formats (e.g., samples with incomplete feature dimensions), and handling missing values ​​(e.g., filling with the mean or median).

[0087] Step 1.3: Feature Encoding and Normalization. Non-numerical features (such as protocol type, service type, etc.) in network traffic data are encoded using one-hot encoding or label encoding. Numerical features are normalized or standardized, such as using Min-Max normalization to scale feature values ​​to the [0,1] interval, or using Z-Score normalization to make the features conform to a standard normal distribution, ensuring that network traffic features of different dimensions are within a uniform numerical range.

[0088] Step 1.4: Minority Classification. Based on the network traffic sample labels, count the number of samples in each category. Classes with a sample number lower than a preset threshold (e.g., 10% of the majority class sample number) or significantly less than the majority class sample number are identified as minority network traffic categories. Extract the minority network traffic sample set as the input for the second step.

[0089] Step 2: Construction and training of the class-balanced model WGAN-GP-BM

[0090] like Figure 2 As shown, the WGAN-GP-BM network traffic class balancing model includes a generator module, a discriminator module, a memory module, and a Bayesian optimization module. The generator module generates minority class network traffic samples; the discriminator module determines whether the input sample is a real minority class sample or a generated minority class sample and outputs a discrimination score; the memory module stores the high-dimensional features of real minority class samples extracted by the intermediate layers of the discriminator module and provides the historical memory features to the generator module; the Bayesian optimization module searches for the optimal hyperparameter combination during model training.

[0091] The construction of the class-balanced model WGAN-GP-BM includes the following steps: ① Data Sample Loading. A data loading module is built using the deep learning framework PyTorch to package the minority class network traffic samples obtained in the first step into batches of a preset size. In one specific implementation, the batch size is set to 256. During training, a batch of real minority class network traffic samples is sampled from the real minority class network traffic sample set, and at the same time, random noise vectors of the same batch size are sampled from a standard Gaussian distribution as input to the generator module.

[0092] ② Generator module construction. For example... Figure 3 As shown, the generator module generates minority class network traffic samples based on a random noise vector and historical memory features provided by the memory module. In one specific implementation, the generator module is input to a 100-dimensional random noise vector z, which follows a standard Gaussian distribution. First, the 100-dimensional random noise vector is input to the first fully connected layer to obtain a 256-dimensional first latent feature representation; then, batch normalization and LeakyReLU activation (with a negative slope set to 0.2) are sequentially applied to the first latent feature representation to obtain the first latent query features. Subsequently, the memory module retrieves relevant data from the memory bank related to the current generation task. A historical memory feature, in a specific implementation =5, and the historical memory features are fused with the first latent query feature representation through a multi-head attention mechanism. The number of heads in the multi-head attention mechanism is set to 4, and the dimension of each attention head is set to 64. The fused features are input into the second fully connected layer to obtain a 512-dimensional second latent feature representation, which is then subjected to batch normalization and LeakyReLU activation. Finally, the second latent feature representation is input into the third fully connected layer, mapping it to the same feature dimension d as the real minority class network traffic sample, and the output value is compressed to the [-1,1] interval using the Tanh activation function, outputting the generated minority class network traffic sample. .

[0093] It's important to note that the core purpose of introducing a multi-head attention mechanism to fuse memory features in the generator module is to enable the generator module, when generating minority class network traffic samples, to selectively focus on historical real features in the memory bank that are relevant to the current generation task, based on the potential query features mapped from the current random noise, rather than relying solely on random noise for sample generation. The multi-head attention mechanism, through parallel computation by multiple attention heads, can capture the correlation between historical memory features and the current generation task from different feature subspaces, thereby improving the consistency between the generated samples and the distribution of real minority class samples, and enhancing the diversity of the generated samples.

[0094] ③ Discriminator module construction. For example... Figure 4 As shown, the discriminator module distinguishes between real minority class network traffic samples and generated minority class network traffic samples, and outputs a sample discrimination score. In one specific implementation, firstly, real minority class network traffic samples and generated minority class network traffic samples are respectively input into the first fully connected layer to obtain a 512-dimensional first discriminant feature; then, the first discriminant feature is subjected to spectral normalization and LeakyReLU activation (with a negative slope set to 0.2). The role of spectral normalization is to constrain the Lipschitz constant of the discriminator module, which, together with the gradient penalty term, ensures the stability of WGAN-GP training. Subsequently, the first discriminant feature is input into the second fully connected layer to obtain a 256-dimensional second discriminant feature, and spectral normalization and LeakyReLU activation are performed again after the second fully connected layer. The second discriminant feature serves two purposes: firstly, it is used as a high-dimensional feature input to the memory module in the intermediate layer of the discriminator module for feature storage and updating in the memory bank; secondly, it is input into the third fully connected layer, outputting a sample discrimination score with a dimension of 1 (i.e., the sample's realism score, with a higher value indicating a closer resemblance to the real sample). The discriminator module calculates the discriminator module loss based on the real minority class sample discrimination score, the generated minority class sample discrimination score, and the gradient penalty term, and updates the discriminator module parameters through backpropagation.

[0095] ④ Memory module construction. For example... Figure 5 As shown, the memory module includes five core sub-modules: feature extraction, feature storage, momentum update, diversity-aware replacement, feature retrieval, and feature fusion.

[0096] Feature extraction: The 256-dimensional second discriminant feature output from the second fully connected layer of the discriminator module is used as a high-dimensional feature representation of the real minority class network traffic samples. This feature, located in the intermediate layer of the discriminator module, possesses stronger abstract expressive power and discriminative ability compared to the original input features. In the formula, This represents the intermediate feature extraction mapping of the discriminator module.

[0097] Feature storage: storing the extracted high-dimensional features Write to the memory. The memory is managed using a circular buffer structure, and its capacity is set to a preset size of 10,000. The advantage of the circular buffer structure is that its memory usage is fixed, eliminating the need for dynamic memory allocation.

[0098] Momentum Update: When the memory is not full, the newly extracted high-dimensional features are directly written into the memory; when the memory is full, the cosine similarity between the newly extracted features and the existing features in the memory is calculated, and the feature is updated according to the momentum update strategy. The formula for calculating the momentum update strategy is:

[0099] In the formula, This represents the memory after step t; This represents the memory after the (t-1)th step; This represents the momentum coefficient. Wherein, the momentum coefficient... Used to control the fusion ratio between historical memory features and newly extracted features. The larger the value, the smoother the update of memory features and the more historical information is retained; The smaller the value, the more sensitive the memory feature is to the response of the most recent samples. In one specific implementation, =0.9. The core advantage of the momentum update strategy lies in its ability to smoothly integrate historical memory features with newly extracted features, avoiding drastic changes in memory features due to abnormal fluctuations in a single batch of samples.

[0100] Diversity-aware replacement: When the memory is full and new features need to be written, a diversity-aware replacement strategy is used to determine which historical feature to replace. The specific rule is: if the maximum cosine similarity between the new feature and existing features in the memory is greater than a preset similarity threshold... (In one specific implementation, If the similarity is 0.8, it indicates that the new feature is highly similar to an existing feature in the database. In this case, the new feature is momentum-fused with the most similar historical feature to replace the most similar historical feature. Otherwise, it indicates that the new feature has low similarity to existing historical features in the database and high feature difference. The new feature is written into the database and replaces the oldest stored historical feature in the database to dynamically update the database content and maintain the diversity of feature distribution in the database. The core innovation of this diversity-aware replacement strategy is that it avoids storing a large number of redundant features in the database (when the new feature is highly similar to an existing feature) while ensuring that the database can cover diverse feature patterns (when the new feature is significantly different from an existing feature), thus achieving a balance between the stability and diversity of feature storage.

[0101] The formula for calculating cosine similarity is:

[0102] In the formula, This represents the high-dimensional features extracted by the intermediate layer of the discriminator module; This represents the i-th historical memory feature stored in the memory bank.

[0103] Feature retrieval and fusion: When the generator module generates samples, the memory module retrieves features from the memory that match the currently queried features based on an approximate nearest neighbor search method (such as using the Annoy algorithm to build an index). Most relevant A historical memory feature, in a specific implementation =5. The retrieved results The historical memory feature matrix is ​​represented as follows:

[0104] In the formula, This indicates the Top- retrieved results. Historical memory feature matrix; This indicates the query features provided by the generator.

[0105] Subsequently, the retrieved historical memory features are fused with the random noise features of the generator module using a multi-head attention mechanism:

[0106]

[0107] in, , , Let represent the query matrix, key matrix, and value matrix in the multi-head attention mechanism, respectively; W represents the learnable matrix. Indicates transpose; This represents the fused feature vector. The number of attention heads is 4, and the dimension of each attention head is... =64. The multi-head attention mechanism, by computing multiple attention heads in parallel, can capture the correlation between memory features and the current generation task from different feature subspaces, enabling the generator module to focus on feature information at different levels when fusing memory features.

[0108] ⑤ Construction of the Bayesian optimization module. For example... Figure 6 As shown, the Bayesian optimization module is used to search for the optimal combination of the generator module learning rate, discriminator module learning rate, gradient penalty coefficient, memory module momentum coefficient, and memory regularization term weights during model training.

[0109] It is important to note that the Bayesian optimization module operates by combining outer-layer hyperparameter search with inner-layer model training and evaluation. During the outer-layer optimization, the Bayesian optimization module proposes candidate hyperparameter combinations based on the Gaussian process surrogate model and the acquisition function. During the inner-layer training and evaluation, the WGAN-GP-BM network traffic class balance model is trained based on these candidate hyperparameter combinations, and the Gaussian process surrogate model is updated according to the validation loss or comprehensive evaluation index obtained during training. Subsequently, the Bayesian optimization module continues to propose the next set of candidate hyperparameter combinations based on the updated surrogate model, repeating the above process until the preset number of optimization iterations is reached or the convergence condition is met, outputting the optimal hyperparameter combination.

[0110] In one specific implementation, a hyperparameter search space is first defined, the hyperparameter search space comprising:

[0111] Then, a probabilistic mapping relationship between hyperparameter combinations and model validation loss is established using a Gaussian process surrogate model. Next, a candidate hyperparameter combination is selected based on the acquisition function. Subsequently, the WGAN-GP-BM network traffic class balance model is trained using the candidate hyperparameter combinations, and the corresponding validation loss is calculated. Finally, the Gaussian process surrogate model is updated based on the validation loss. This process is repeated until the maximum number of optimization iterations is reached or the preset convergence condition is met, at which point the optimal hyperparameter combination is output for subsequent model training.

[0112] The validation loss is measured by the KL divergence or mean absolute error (MAE) between the minority class network traffic samples generated by the generator module on the validation set under the current hyperparameter combination and the real minority class network traffic samples on the validation set. The validation set is an independent set of samples separated from the original real minority class samples that do not participate in the training process of the WGAN-GP-BM model.

[0113] The Gaussian process surrogate model in the Bayesian optimization process is:

[0114] The mean function is:

[0115] The kernel function is the Matérn 5 / 2 kernel function:

[0116]

[0117]

[0118] The acquisition function is the acquisition function that is expected to be improved:

[0119]

[0120]

[0121]

[0122]

[0123] In the formula, This represents the current combination of hyperparameters to be evaluated. This represents another set of hyperparameter combinations. This represents the validation loss or objective function value corresponding to the hyperparameter combination. Represents a Gaussian process. Represents the mean function, This represents a kernel function used to measure the correlation between two sets of hyperparameters. Indicates the learning rate. Represents the gradient penalty coefficient. Indicates the weight of the memory module. This indicates scaling the Mahalanobis distance. Let Λ denote the signal variance, Λ denote the anisotropic length scaling matrix, and diag denote the diagonal matrix. This indicates the correlation length of each dimension. This indicates a desire to improve the acquisition function. This represents the weighting coefficients that change with iteration. and Let represent the desired improvements for local search and global search, respectively. This represents the smallest observed value of the objective function; This represents the mean of the Gaussian process's predictions for the current hyperparameter combination; This indicates the uncertainty in the Gaussian process's prediction of this location; Indicates the exploration intensity coefficient; This indicates the current iteration number of the Bayesian optimization; This represents the maximum number of iterations for Bayesian optimization.

[0124] The core mechanism of the Gaussian process surrogate model in establishing the probabilistic mapping relationship between hyperparameter combinations and validation loss is as follows: First, the similarity between any two sets of explored hyperparameters is calculated using the Matérn 5 / 2 kernel function (the closer the distance, the higher the similarity). Then, based on the true validation loss values ​​of all explored points, a high-dimensional joint normal distribution is constructed. When it is necessary to predict the validation loss of a new set of candidate hyperparameter combinations, the predicted mean (i.e., the estimated value of the validation loss) and predicted variance (i.e., the uncertainty of the estimate) of that point are calculated using the conditional probability formula. After each real model training iteration and obtaining the accurate validation loss, this real value is fed back to the Gaussian process to update the joint distribution, thereby making subsequent predictions more accurate. The core advantage of this mechanism is that it can fully utilize the expensive model training results each time, finding the optimal hyperparameter combination with the fewest training iterations.

[0125] ⑥ Model Training. First, initialize the generator module, discriminator module, memory module, optimizer, and Bayesian optimization module. The optimizer uses the Adam adaptive moment estimation optimizer. The generator and discriminator modules are each configured with independent Adam optimizers, and their learning rates are determined as hyperparameters by the Bayesian optimization module. Then, a batch of real minority class network traffic samples is sampled from the real minority class network traffic sample set, and random noise vectors of the same batch size are sampled from a standard Gaussian distribution. The random noise vectors are input into the generator module to obtain generated minority class network traffic samples. The real minority class network traffic samples and generated minority class network traffic samples are input into the discriminator module respectively to calculate the discrimination scores of the real minority class samples and the generated minority class samples. Interpolation samples are constructed based on the real minority class samples and generated minority class samples, and gradient penalty terms are calculated. The discriminator module loss is calculated based on the discriminator Wasserstein loss and gradient penalty terms, and the discriminator module parameters are updated. High-dimensional features of the real minority class samples are extracted using the intermediate layer of the discriminator module, and the memory is updated according to the rules of the memory module. Subsequently, the features are retrieved from the memory. In one specific implementation, a related historical memory feature, =5, fuse it with the latent features of the generator module, and calculate the generator module loss based on the feedback results of the discriminator module on the generation of minority class samples and the memory regularization term, and update the generator module parameters; repeat the above training process until the model reaches the preset training epoch (e.g., Epoch=1000) or meets the preset convergence condition (e.g., the generator module loss no longer decreases for 50 consecutive epochs), and save the optimal generator module.

[0126] It is important to note that the Bayesian optimization module and model training process can be viewed as a two-layer optimization structure: the outer layer is the Bayesian optimization module's search and evaluation of the hyperparameter space, and the inner layer is the adversarial training of the generator and discriminator modules under the current candidate hyperparameter combinations. The outer Bayesian optimization module updates the Gaussian process surrogate model based on the validation loss or comprehensive evaluation index obtained from the inner training layer, and proposes new candidate hyperparameter combinations based on the updated surrogate model; the inner model training process updates the network parameters of the generator and discriminator modules based on the current candidate hyperparameter combinations. This process is repeated until the optimal hyperparameter combination is obtained, and the final model training is completed using the optimal hyperparameter combination.

[0127] The calculation formula for the interpolated sample is:

[0128] in, Indicates the interpolated sample; This represents a sample of real minority network traffic. ε represents the minority class network traffic samples generated by the generator module; ε represents the random coefficients sampled from the preset distribution.

[0129] The total loss function of the discriminator module is:

[0130] in, This represents the total loss of the discriminator module; For the discriminator module; This represents the true minority class sample distribution; The distribution of minority class samples generated by the generator module; For the interpolated sample distribution; This represents the true minority class samples; This represents the minority class samples generated by the generator module; For interpolation samples; This is the gradient penalty coefficient; For expectation calculation; This represents the L2 norm.

[0131] It should be noted that the first half of the formula The Wasserstein loss for the discriminator is used to approximate the Wasserstein distance between the real and generated distributions by measuring the difference in scores between real and generated samples from the discriminator module. Compared to traditional GANs using JS divergence or KL divergence, the Wasserstein distance exhibits better smoothness and can alleviate gradient vanishing and mode collapse problems to some extent. Gradient penalty term. This constraint ensures that the discriminator module satisfies the 1-Lipschitz condition in the interpolation space between real and generated samples, making the discriminator gradient norm at the interpolated samples close to 1, thereby improving the stability of the adversarial training process. The optimization objective of the discriminator module is to minimize... That is, while maximizing the difference between the discrimination scores of real samples and generated samples, the gradient penalty constraint is satisfied.

[0132] The total loss function of the generator module is:

[0133] in, This represents the total loss of the generator module; For generator modules; For the discriminator module; It is a random noise vector; The noise distribution is random. To generate a minority class sample distribution; The distribution of historical features in the memory module; To remember the weights of the regularization terms; This is the KL divergence constraint term.

[0134] It should be noted that the first half of the formula The generator module uses the Wasserstein loss to minimize this term, enabling the discriminator module to assign a more accurate discrimination score to the generated minority class network traffic samples compared to the real samples. This, in turn, helps the generated samples gradually approximate the distribution of the real minority class samples. (The latter part of the formula...) For memorizing the regularization terms, where, This represents the KL divergence between the generated sample distribution and the memory distribution represented by historical features in the memory module. The technical role of the memory regularization term is to constrain the statistical distribution of the samples generated by the generator module to be consistent with the high-dimensional feature distribution of the real minority class samples stored in the memory bank, thereby guiding the generator module to fully utilize historical real feature information and reducing the risk of generated sample distribution deviation or getting trapped in local optima. After introducing the memory regularization term, the generator module is simultaneously constrained in two ways during the optimization process: on the one hand, by minimizing the Wasserstein loss, it improves the authenticity score of the generated samples for the discriminator module; on the other hand, by minimizing the KL divergence, it enhances the consistency between the generated sample distribution and the memory distribution. Memory regularization term weights. In one specific implementation, the Bayesian optimization module automatically searches and determines the method. The search scope is .when When the value is large, the generator module tends to generate samples that are consistent with the distribution of the memorized features; when When the value is small, the generator module is less constrained by the memory distribution. Bayesian optimization is used to automatically search for a better value. The value can strike a balance between the quality of generated samples, distribution consistency, and sample diversity.

[0135] Step 3: Generating Minority Class Network Traffic Samples and Constructing a Class Balanced Dataset

[0136] like Figure 7 As shown, the trained optimal generator module generates a specified number of network traffic samples for minority network traffic categories. The generated minority network traffic samples are then merged with the original network traffic training set to obtain a class-balanced network traffic training set.

[0137] Step 3.1: Count the number of samples in each category in the original network traffic training set to determine the target number of samples for category balancing. Usually, the number of samples in the majority class (such as normal traffic) is used as the target number of samples for category balancing to ensure that the number of samples in all categories is equal to that in the majority class after balancing.

[0138] Step 3.2: Based on the target number of samples for category balancing and the original number of samples for each minority class of network traffic, determine the number of samples to be generated for each minority class. For example, if the majority class has 10,000 samples and a minority class has 200 original samples, then 9,800 samples need to be generated for that minority class. For categories whose sample size exceeds the target balancing number, a random sampling method is used to control their sample size.

[0139] Step 3.3: For each minority network traffic category, extract the real minority network traffic sample set corresponding to that category, and use the real minority network traffic sample set to train the corresponding WGAN-GP-BM network traffic category balance model to obtain the optimal generator module corresponding to that category; use the optimal generator module to generate the corresponding number of minority network traffic samples with corresponding category labels.

[0140] It is important to note that the network traffic characteristics of different attack types are completely different (e.g., DDoS attacks are characterized by a surge in traffic, brute-force attacks by an abnormal frequency of login failures, and port scanning by accessing a large number of different ports), resulting in significant differences in data distribution across categories. Therefore, in step 3.3, an independent WGAN-GP-BM network traffic category balancing model is trained for each minority class of network traffic. Each model has independent initialization parameters, an independent memory, and an independent Bayesian optimization process. In other words, if there are L minority classes, the model building and training process needs to be executed L times to obtain L independent optimal generator modules. This strategy ensures that the generated samples for each category accurately reflect the true feature distribution of that category, avoiding interference between features of different categories that could lead to a decrease in the quality of the generated samples.

[0141] In a preferred embodiment, the training of the WGAN-GP-BM network traffic category balancing model for multiple minority network traffic categories is performed in parallel computing (e.g., using multi-GPU parallel training) to significantly improve overall processing efficiency.

[0142] Step 3.4: Add the generated minority class network traffic samples to the original network traffic training set, and randomly shuffle the merged training set to obtain a class-balanced network traffic training set. The purpose of random shuffling is to avoid gradient oscillation during anomaly detection model training caused by arranging the training set in class order, and to ensure that normal samples and attack samples are evenly interspersed during training.

[0143] Step 4: Generate sample quality and detection performance evaluation

[0144] The distributional and numerical differences between the generated minority class network traffic samples and the real minority class network traffic samples are evaluated using KL divergence and mean absolute error. The network traffic anomaly detection performance after balancing the categories is evaluated using one or more of accuracy, precision, recall, and F1 score.

[0145] It should be noted that the KL divergence and mean absolute error (MAE) in the fourth step have a dual function: firstly, they serve as evaluation metrics for the quality of the final generated samples, verifying the technical effectiveness of the method in sample generation; secondly, during the Bayesian optimization process, the validation loss is measured by the KL divergence or mean absolute error between the minority class network traffic samples generated by the generator module on the validation set under the current hyperparameter combination and the actual minority class network traffic samples on the validation set. The Bayesian optimization module uses this validation loss as the objective function to guide the hyperparameter search process.

[0146] The formulas for calculating KL divergence, mean absolute error, accuracy, precision, recall, and F1 score are as follows:

[0147]

[0148]

[0149]

[0150]

[0151]

[0152] in, Represents the true sample distribution With the distribution of generated samples Between Divergence; Represents the true sample distribution In the Probability values ​​over an interval or category; Represents the distribution of generated samples In the Probability values ​​over an interval or category; Indicates the mean absolute error; Representing the first real sample One eigenvalue; Indicates the number of generated samples One eigenvalue; Indicates the number of samples or features; Indicates accuracy; This indicates the number of abnormal traffic instances that were correctly detected as abnormal. This indicates the number of normal traffic flows that were incorrectly detected as abnormal. This indicates the number of abnormal traffic flows that were incorrectly detected as normal. This indicates the number of normal traffic flows that were correctly detected as normal. Indicates accuracy; Indicates recall rate; This represents the F1 score.

[0153] Step 5: Network Traffic Anomaly Detection

[0154] The network traffic training set after category balance is input into the network traffic anomaly detection model for training, and the trained network traffic anomaly detection model is used to classify and detect unknown network traffic samples.

[0155] Step 5.1: Input the class-balanced network traffic training set into the network traffic anomaly detection model for training. The network traffic anomaly detection model employs a pre-defined multi-classifier architecture, including but not limited to any one of the following: tree-based ensemble models (such as Random Forest, XGBoost), kernel-based classification models (such as Support Vector Machine (SVM), or deep neural network-based classification models (such as Multilayer Perceptron (MLP), Convolutional Neural Network (CNN), Recurrent Neural Network (LSTM)). This invention does not limit the specific type of anomaly detection model; this model serves only as a downstream task carrier for verifying the effectiveness of the class balancing method of this invention.

[0156] Step 5.2: Evaluate the performance of the trained network traffic anomaly detection model using a validation set or test set independent of the generated samples. It is important to note that this validation / test set should be an independent set of samples reserved from the original real network traffic samples that has not participated in any of the processes in steps one through three (including not participating in the training and sample generation of the WGAN-GP-BM network traffic category balancing model) to ensure the objectivity of the evaluation results.

[0157] Step 5.3: Input the network traffic sample to be detected into the trained network traffic anomaly detection model to obtain the predicted category of the network traffic sample (such as "normal" or a specific attack type label).

[0158] Finally, to verify the effectiveness of the proposed method, experiments were conducted using publicly available network traffic datasets. In the experiments, the anomaly detection model was first trained using the original imbalanced training set to obtain benchmark detection results. Then, the proposed WGAN-GP-BM network traffic class balancing model was used to generate minority class network traffic samples, and a class-balanced training set was constructed. Finally, the anomaly detection model was retrained using the class-balanced training set (keeping the structure and hyperparameters of the anomaly detection model completely consistent with the benchmark experiment to eliminate the influence of model structure differences). KL divergence, mean absolute error, accuracy, precision, recall, and F1 score were used to evaluate the quality of the generated samples and the anomaly detection performance. Through the above-mentioned controlled variable method of experimental evaluation, the technical effectiveness of this invention in terms of generated sample quality, class balancing effect, and minority class anomaly traffic identification ability can be accurately verified. If the experimental group significantly outperforms the control group in terms of recall, F1 score, and other indicators, it proves that the class balancing method of this invention can effectively improve the anomaly detection model's ability to identify minority class attack traffic.

[0159] In addition, the present invention also provides a network traffic category balancing system based on WGAN-GP and its optimization, including: a data preprocessing module, a model building and training module, a sample generation and dataset construction module, an evaluation module, and an anomaly detection module.

[0160] The data preprocessing module is used to preprocess the network traffic dataset to obtain network traffic feature data, and to count the number of samples in each category based on the sample labels to determine the minority network traffic categories.

[0161] The model building and training module is used to construct and train a WGAN-GP-BM network traffic class balance model consisting of a generator module, a discriminator module, a memory module, and a Bayesian optimization module. The generator module receives random noise vectors and historical memory features provided by the memory module to generate minority class network traffic samples. The discriminator module distinguishes between real minority class network traffic samples and generated minority class network traffic samples, and outputs a sample discrimination score. The memory module stores high-dimensional features of real minority class network traffic samples extracted from the intermediate layers of the discriminator module and provides historical memory features to the generator module. The Bayesian optimization module searches for the optimal combination of hyperparameters during training, including the generator module learning rate, the discriminator module learning rate, the gradient penalty coefficient, the memory module momentum coefficient, and the memory regularization term weight.

[0162] The sample generation and dataset construction module is used to train the corresponding WGAN-GP-BM network traffic class balancing model for each minority network traffic class, generate a corresponding number of minority network traffic samples using the trained optimal generator module, and merge the generated minority network traffic samples with the original network traffic training set to obtain the class-balanced network traffic training set.

[0163] The evaluation module is used to evaluate the distribution and numerical differences between the generated minority class network traffic samples and the real minority class network traffic samples using KL divergence and mean absolute error. It evaluates the network traffic anomaly detection performance by using one or more of the following: accuracy, precision, recall, and F1 score, after balancing the evaluation categories.

[0164] The anomaly detection module is used to input the class-balanced network traffic training set into the network traffic anomaly detection model for training, and to use the trained network traffic anomaly detection model to classify and detect unknown network traffic samples.

[0165] Finally, the present invention also provides a non-transitory computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned network traffic category balancing method based on WGAN-GP and its optimization.

[0166] The foregoing has shown and described the main features and advantages of the present invention. It will be apparent to those skilled in the art that the specific embodiments of the present invention are not limited to the details of the exemplary embodiments described above. Furthermore, without departing from the spirit or essential characteristics of the present invention, the inventive concept and design ideas of the present invention can be implemented in other specific forms, and these should be equivalently included within the protection scope disclosed in the technical solutions of the present invention. Therefore, the embodiments should be considered exemplary and non-limiting in all respects. The scope of the present invention is defined by the appended claims rather than the foregoing description, and thus all changes falling within the meaning and scope of the equivalents of the claims are intended to be included within the present invention.

[0167] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

[0168] Contents not described in detail in this specification are prior art known to those skilled in the art. Although illustrative specific embodiments of the invention have been described above to facilitate understanding by those skilled in the art, it should be understood that the invention is not limited to the scope of these specific embodiments. For those skilled in the art, various modifications will be obvious as long as they fall within the spirit and scope of the invention as defined and established by the appended claims, and all inventions utilizing the inventive concept are protected.

Claims

1. A network traffic category balancing method based on WGAN-GP and its optimization, characterized in that, Includes the following steps: S1: Preprocess the network traffic dataset to obtain network traffic feature data, and count the number of samples in each category based on the sample labels to determine the minority network traffic categories; S2: Construct and train a WGAN-GP-BM network traffic class balance model consisting of a generator module, a discriminator module, a memory module, and a Bayesian optimization module; The generator module receives a random noise vector and historical memory features provided by the memory module to generate minority class samples; the discriminator module distinguishes between real minority class samples and generated minority class samples, and outputs a sample discrimination score; the memory module stores the high-dimensional features of real minority class samples extracted by the intermediate layer of the discriminator module, and provides historical memory features to the generator module. The Bayesian optimization module searches for the optimal combination of hyperparameters during the training process. The hyperparameters include the learning rate of the generator module, the learning rate of the discriminator module, the gradient penalty coefficient, the momentum coefficient of the memory module, and the weight of the memory regularization term. During training, real minority class samples and generated minority class samples are input into the discriminator module for feature extraction and discrimination, respectively. The high-dimensional features output from the intermediate layer of the discriminator module are written into the memory module, and the memory is dynamically updated based on feature similarity and momentum update strategies. Random noise vectors are input into the generator module, which retrieves Top-Number samples from the memory module. The relevant historical memory features are fused with random noise features through a multi-head attention mechanism to generate minority class samples; the parameters of the discriminator module are updated by the discriminator Wasserstein loss and gradient penalty term, and the parameters of the generator module are updated by the generator Wasserstein loss and memory regularization term. At the same time, the Bayesian optimization module updates the hyperparameter combination according to the validation loss until the convergence condition is met, and the optimal generator module is saved. S3: For each minority class, train the corresponding WGAN-GP-BM network traffic class balance model, use the trained optimal generator module to generate the corresponding number of minority class samples, and merge the generated minority class samples with the original training set to obtain the class-balanced network traffic training set. S4: The distribution and numerical differences between the generated minority class samples and the real minority class samples are evaluated using KL divergence and mean absolute error. The network traffic anomaly detection performance after balancing the categories is evaluated using one or more of accuracy, precision, recall and F1 score. S5: Input the network traffic training set after category balance into the network traffic anomaly detection model for training, and use the trained network traffic anomaly detection model to classify and detect unknown network traffic samples.

2. The network traffic category balancing method based on WGAN-GP and its optimization as described in claim 1, characterized in that, The generator module is constructed as follows: a random noise vector is input into the first fully connected layer to obtain the first latent feature; batch normalization and nonlinear activation are performed on the first latent feature to obtain the first latent query feature; Retrieve Top- from the memory module The historical memory features are fused with the first potential query features through a multi-head attention mechanism; the fused features are transformed by the second fully connected layer and then mapped to the same feature dimensions as the target minority class real samples through the third fully connected layer, and the generated minority class samples are output. The discriminator module is constructed as follows: real minority class samples and generated minority class samples are respectively input into the first fully connected layer and the second fully connected layer for feature extraction, and spectral normalization and nonlinear activation are added after the first fully connected layer and the second fully connected layer, respectively; the high-dimensional features output by the second fully connected layer are used as intermediate layer features and input into the memory module, and the high-dimensional features are input into the third fully connected layer to output the sample discrimination score; The memory module is updated by extracting high-dimensional features of real minority class samples using the intermediate layer of the discriminator module. The high-dimensional features are stored in the memory bank of the memory module. When the memory bank is not full, they are written directly. When the memory bank is full, the cosine similarity between the new features and the existing features in the memory bank is calculated, and dynamic updates are performed according to the momentum update strategy and the diversity perception replacement strategy. The momentum update strategy is as follows: The cosine similarity is: The diversity-aware replacement strategy is as follows: when the memory bank is full, if the maximum cosine similarity between the new feature and the existing features in the memory bank is greater than a preset threshold, then the new feature is replaced by the most similar historical feature after momentum fusion; otherwise, the new feature is written into the memory bank and replaces the earliest stored historical feature in the memory bank. The generator module retrieves from the memory module. When considering relevant historical memory features, an approximate nearest neighbor search method is used. The features are then fused with the latent features of the generator module through a multi-head attention mechanism. The feature fusion method of the multi-head attention mechanism is as follows: In the formula, This represents the high-dimensional features extracted by the intermediate layer of the discriminator module; This represents the intermediate feature extraction mapping of the discriminator module; This represents the true minority class samples; This represents the memory after step t; This represents the memory after the (t-1)th step; Indicates the momentum coefficient; This represents the i-th historical memory feature stored in the memory bank; Indicates the Top- retrieved results Historical memory feature matrix; This represents the query features provided by the generator and is the output of the intermediate layer of the generator module. , , These represent the query matrix, key matrix, and value matrix in the multi-head attention mechanism, respectively. Let W be a random noise vector; W represents a learnable matrix. Indicates the attention head dimension; Indicates transpose; This represents the fused feature vector; This represents the approximate nearest neighbor search algorithm; the Softmax function represents the normalized exponential function, which is used to convert the correlation scores between multiple historical features and the current query feature into attention weights.

3. The network traffic category balancing method based on WGAN-GP and its optimization as described in claim 2, characterized in that, The Top- In =5; the multi-head attention mechanism has 4 heads, and each attention head has a dimension of 64; the random noise vector is 100-dimensional; the momentum coefficient =0.9; the preset similarity threshold is 0.

8.

4. The network traffic category balancing method based on WGAN-GP and its optimization as described in claim 1, characterized in that, The Bayesian optimization module uses a Gaussian process surrogate model to establish a probabilistic mapping relationship between hyperparameter combinations and validation loss. It selects candidate hyperparameter combinations based on the acquisition function, iteratively updates the Gaussian process surrogate model based on the validation loss, and outputs the optimal hyperparameter combination. The Gaussian process proxy model is as follows: The acquisition function is the acquisition function that is expected to be improved: In the formula, Indicates a combination of hyperparameters. This represents the validation loss or objective function value corresponding to the hyperparameter combination. Represents a Gaussian process. Represents the mean function, Represents the kernel function. This indicates a desire to improve the acquisition function. This represents the weight coefficients that change with iteration. and These represent the desired improvements for local and global searches, respectively.

5. The network traffic category balancing method based on WGAN-GP and its optimization according to claim 4, characterized in that, The Bayesian optimization module operates as follows: Step 2.1: Define the hyperparameter search space; Step 2.2: Establish the probabilistic mapping relationship between hyperparameter combinations and validation loss using a Gaussian process surrogate model; Step 2.3: Select candidate hyperparameter combinations based on the acquisition function; Step 2.4: Perform a complete training cycle using the candidate hyperparameter combination and calculate the corresponding validation loss; the validation loss is measured by the KL divergence or mean absolute error between the minority class samples generated by the generator module on the validation set and the real minority class samples on the validation set. Step 2.5: Update the Gaussian process surrogate model based on the verification loss; Step 2.6: Repeat steps 2.3 to 2.5 until the maximum number of optimization attempts is reached or the convergence condition is met, and output the optimal hyperparameter combination; Furthermore, the specific process of performing a complete training of the WGAN-GP-BM network traffic class balancing model is as follows: (10) Initialize the generator module, the discriminator module, the memory module, the optimizer, and the Bayesian optimization module; (20) Sample a batch of real minority class samples from the real minority class sample set, and sample random noise vectors of the same batch size from the standard Gaussian distribution; (30) Input the random noise vector into the generator module to obtain the generated minority class samples; (40) Input the real minority class samples and the generated minority class samples into the discriminator module respectively, calculate the discrimination score, construct the interpolation sample and calculate the gradient penalty term, calculate the total loss of the discriminator module based on the discriminator Wasserstein loss and the gradient penalty term and update the discriminator module; (50) Extract high-dimensional features of real minority class samples using the intermediate layer of the discriminator module and update the memory bank of the memory module; (60) Retrieve from the memory module The historical memory features are fused with the latent features of the generator module. The total loss of the generator module is calculated based on the feedback and memory regularization terms of the discriminator module, and the generator module is updated. (70): Repeat steps (20) to (60) until the preset training rounds are reached or the convergence condition is met, and save the optimal generator module.

6. The network traffic category balancing method based on WGAN-GP and its optimization as described in claim 5, characterized in that, The total loss function of the discriminator module is: in, This represents the total loss of the discriminator module; For the discriminator module; This represents the true minority class sample distribution; The distribution of minority class samples generated by the generator module; For the interpolated sample distribution; This represents the true minority class samples; This represents the minority class samples generated by the generator module; For interpolation samples; This is the gradient penalty coefficient; For expectation calculation; Describing the L2 norm, These are the random coefficients sampled from a preset distribution; The total loss function of the generator module is: in, This represents the total loss of the generator module; For generator modules; For the discriminator module; It is a random noise vector; The noise distribution is random. To generate a minority class sample distribution; The distribution of historical features in the memory module; To remember the weights of the regularization terms; This is the KL divergence constraint term.

7. The network traffic category balancing method based on WGAN-GP and its optimization as described in claim 1, characterized in that, S3 includes: Step 3.1: Count the number of samples of each category in the original network traffic training set to determine the target number of samples for category balance; Step 3.2: Based on the target number of samples for each category and the original number of samples for each minority network traffic category, determine the number of samples that need to be generated for each minority network traffic category; Step 3.3: For each minority network traffic category, extract the real minority network traffic sample set corresponding to that category, train the corresponding WGAN-GP-BM network traffic category balancing model, obtain the optimal generator module for that category, and use the optimal generator module to generate a corresponding number of minority network traffic samples with category labels; wherein, the training of WGAN-GP-BM network traffic category balancing models for multiple minority network traffic categories is performed in parallel computing. Step 3.4: Add the generated minority class network traffic samples to the original network traffic training set and shuffle them randomly to obtain a class-balanced network traffic training set.

8. The network traffic category balancing method based on WGAN-GP and its optimization as described in claim 1, characterized in that, S5 includes: Step 5.1: Input the class-balanced network traffic training set into the network traffic anomaly detection model for training; Step 5.2: Evaluate the performance of the trained network traffic anomaly detection model using a validation set or test set that is independent of the generated minority class network traffic samples; Step 5.3: Input the network traffic sample to be detected into the trained network traffic anomaly detection model to obtain the predicted category of the network traffic sample.

9. A network traffic category balancing system based on WGAN-GP and its optimization, characterized in that, include: The data preprocessing module is used to preprocess the network traffic dataset to obtain network traffic feature data, and to count the number of samples in each category based on the sample labels to determine the minority network traffic categories. The model building and training module is used to build and train a WGAN-GP-BM network traffic class balance model consisting of a generator module, a discriminator module, a memory module, and a Bayesian optimization module. The generator module receives a random noise vector and historical memory features provided by the memory module to generate minority class network traffic samples. The discriminator module distinguishes between real minority class network traffic samples and generated minority class network traffic samples and outputs a sample discrimination score. The memory module stores high-dimensional features of real minority class network traffic samples extracted by the intermediate layer of the discriminator module and provides historical memory features to the generator module. The Bayesian optimization module is used to search for the optimal combination of hyperparameters during the training process. The hyperparameters include the learning rate of the generator module, the learning rate of the discriminator module, the gradient penalty coefficient, the momentum coefficient of the memory module, and the weight of the memory regularization term. The sample generation and dataset construction module is used to train the corresponding WGAN-GP-BM network traffic class balancing model for each minority network traffic class, generate a corresponding number of minority network traffic samples using the trained optimal generator module, and merge the generated minority network traffic samples with the original network traffic training set to obtain the class-balanced network traffic training set. The evaluation module is used to evaluate the distribution and numerical differences between the generated minority class network traffic samples and the real minority class network traffic samples using KL divergence and mean absolute error. It evaluates the network traffic anomaly detection performance after balancing the categories using one or more of accuracy, precision, recall and F1 score. The anomaly detection module is used to input the class-balanced network traffic training set into the network traffic anomaly detection model for training, and to use the trained network traffic anomaly detection model to classify and detect unknown network traffic samples.

10. A non-transitory computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the network traffic category balancing method based on WGAN-GP and its optimization as described in any one of claims 1 to 9.