Intrusion detection method for data imbalance problem
By improving the auxiliary classifier generation adversarial network model and Bayesian optimization algorithm, a diverse minority class samples are generated, which solves the problem of data imbalance in traditional intrusion detection systems and improves the detection effect.
Patent Information
- Application Number
- CN202510908255.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2025-08-15
AI Technical Summary
When traditional intrusion detection systems face data imbalance, especially the detection rate of a few attack categories is low, and the existing methods generate samples with insufficient realism and insufficient diversity, resulting in poor detection results.
By improving the auxiliary classifier to generate an adversarial network model, replacing the convolutional layer with a fully connected layer, optimizing hyperparameters with Bayesian optimization algorithm, generating diverse and realistic minority class samples, and using a random forest detection model for training, balancing the data set to improve detection performance.
It effectively improves the detection effect of a few types of samples, reduces the rate of missed and false alarms, and improves the overall performance of intrusion detection.
Smart Images

Figure CN120498882A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of network security, and in particular to an intrusion detection method for data imbalance problem. Background Art
[0002] Against the backdrop of the rapid development of information technology, the widespread adoption of internet-connected devices has greatly improved communication efficiency and data transmission capabilities, bringing unprecedented convenience to humanity. However, this trend has also been accompanied by the frequent occurrence of internet security incidents, highlighting numerous potential cybersecurity risks and making cyberspace security a global concern.
[0003] Traditional intrusion detection systems rely on rule- and signature-based detection methods, which are effective in identifying known threats but lack the ability to respond to new, unknown attacks. The application of machine learning methods to intrusion detection has significantly improved their ability to identify unknown attacks. However, as network data becomes increasingly complex and multidimensional, traditional machine learning methods often perform poorly. In recent years, deep learning methods have also been applied to intrusion detection, enhancing detection of unknown attacks through automated feature extraction.
[0004] However, machine learning and deep learning methods generally face the problem of data imbalance. In real-world scenarios, there is a significant imbalance between normal and malicious traffic, as well as between different malicious traffic categories. The detection rate of a few attack categories remains low.
[0005] To address the poor detection results caused by data imbalance, researchers have proposed various methods for handling imbalanced data. Early approaches employed by researchers included random oversampling and random undersampling. However, random oversampling can lead to overfitting, while random undersampling can cause the loss of important information. As research on data imbalance deepens, methods such as synthetic minority oversampling and adaptive synthetic sampling have been proposed to address the shortcomings of these two approaches. However, these methods often suffer from insufficient realism and diversity in the generated samples. Summary of the Invention
[0006] In view of the above technical deficiencies, the purpose of the present invention is to provide an intrusion detection method for processing data imbalance problems, so as to solve the data imbalance problem during model training and improve network intrusion detection performance.
[0007] To achieve the above objectives, the technical solution adopted by the present invention is an intrusion detection method for the data imbalance problem, comprising:
[0008] Obtain network intrusion data, remove dirty data, and standardize the data;
[0009] Improve the auxiliary classifier generative adversarial network model by multiplying the noise vector with the embedded category label element by element and replacing the convolutional layers of the generator and discriminator with fully connected layers;
[0010] The Bayesian optimization algorithm is used to optimize the hyperparameters of the improved auxiliary classifier generative adversarial network model, and the model with the best parameters is used to generate minority class samples to balance the dataset.
[0011] Use the Bayesian optimization algorithm to optimize the hyperparameters of the random forest detection model and use the balanced dataset to train the optimal detection model;
[0012] The detection model is trained using a balanced data set, and the trained target intrusion detection model is used for network intrusion detection.
[0013] Preferably, the data preprocessing of the network intrusion detection data includes:
[0014] Performing missing value processing on the network intrusion data;
[0015] Eliminating dirty data from the network intrusion data;
[0016] Standardize the processed data:
[0017]
[0018] in, is the original eigenvalue, is the mean of the feature, is the standard deviation of the feature, is the normalized eigenvalue.
[0019] Preferably, the improvement of the auxiliary generative adversarial network model includes:
[0020] The category label Embedded into the same dimensional space as the noise vector, the embedding vector is obtained :
[0021]
[0022] The noise vector Multiply element-wise with the embedded category label to combine the category information with the noise:
[0023]
[0024] The convolutional layers of the generator and discriminator are replaced with fully connected layers, and the noise vector is gradually mapped to the generated sample through a series of fully connected layers and activation function LeakyReLU:
[0025]
[0026] The output is generated by the hyperbolic tangent activation function tanh :
[0027]
[0028] in, and are the weights and biases of the network.
[0029] Preferably, the Bayesian optimization algorithm performs hyperparameter tuning on the auxiliary classifier generative adversarial network model and the random forest detection model, including:
[0030] Initial sampling: Randomly sample several sets of hyperparameters and calculate the objective function value to construct the initial data set;
[0031] Constructing a proxy model: Constructing a proxy model of the objective function based on the current data set. The Gaussian process used in this invention is:
[0032]
[0033] in, is the mean function, is the covariance function;
[0034] The kernel function used in the present invention is the radial aggregate function:
[0035]
[0036] in, is the output variance, is the length scale;
[0037] Collect new samples: Use the collection function to find the optimal collection point. The collection function is used to select the next evaluation point. The method used in this paper is the expected improvement:
[0038]
[0039] in, It is the best point currently known;
[0040] Update model: Evaluate the objective function value, add new data to the dataset, and update the Gaussian process model;
[0041] Iterative optimization: Repeat the steps of collecting new samples and updating the model until a predetermined stopping condition is met.
[0042] Preferably, the specific parameters of the hyperparameter tuning include:
[0043] Auxiliary classifier generative adversarial network model: generator learning rate, discriminator learning rate, latent space vector and batch size;
[0044] Random forest detection model: the number of decision trees, the maximum depth of the decision tree, the minimum number of samples required for internal nodes to split again, and the minimum number of samples required for leaf nodes.
[0045] Preferably, the target balanced data is used to train a network intrusion detection model to obtain a target network intrusion detection model, including:
[0046] Determine the number of training times based on the amount of data in the dataset;
[0047] The network intrusion detection model is trained and verified using a balanced data set to obtain multiple verification results. When all verification results meet the preset requirements, the trained network intrusion detection model is determined to be the target network intrusion detection model.
[0048] Compared to existing technologies, this paper proposes an intrusion detection method for addressing data imbalance. This method offers the following advantages: It improves upon the traditional auxiliary classifier generative adversarial network model, enabling more efficient generation of numerical data for specific categories. It utilizes a Bayesian optimization algorithm to optimize the improved auxiliary classifier generative adversarial network model to balance the dataset, avoiding hyperparameter sensitivity and generating diverse and realistic augmented samples to balance the dataset. It also utilizes a Bayesian optimization algorithm to optimize the random forest detection model, resulting in improved intrusion detection performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 This is a flow chart of an intrusion detection method for data imbalance problem proposed by the present invention;
[0050] Figure 2 for Figure 1 The structure diagram of the auxiliary classifier generating adversarial network model;
[0051] Figure 3 for Figure 1 Flowchart of the Bayesian optimization algorithm. DETAILED DESCRIPTION
[0052] The following is a detailed description of the embodiments of the technical solution of the present invention in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention and are therefore only examples and are not intended to limit the scope of protection of the present invention.
[0053] It should be noted that, unless otherwise specified, the technical or scientific terms used in this application should have the common meanings understood by those skilled in the art to which the present invention belongs;
[0054] like Figure 1As shown in FIG, an embodiment of the present invention provides a flow chart of an intrusion detection method for data imbalance problem, the method comprising:
[0055] Step 1: Obtain network intrusion data, remove dirty data, and standardize the data;
[0056] First, since there is a small amount of dirty data in the dataset, it is removed, for example, the mean is used to replace the missing value "NAN";
[0057] Secondly, the eigenvalues are normalized to eliminate the dimensional differences between different features and ensure that each feature is on the same scale. This ensures that during the model training process, certain features do not dominate the learning process due to differences in feature value ranges.
[0058] Step 2: The present invention improves the auxiliary classifier generative adversarial network model to solve the data imbalance problem and enhance the diversity of training samples by generating numerical data;
[0059] Traditional generative adversarial network models with auxiliary classifiers are primarily used for image generation tasks. Their generators and discriminators typically employ convolutional neural network structures, which are suitable for spatially structured image data. However, malicious traffic data is numerical, structured, and high-dimensional, lacking the spatial correlation found in images. Therefore, traditional convolutional layers are not suitable for processing numerical data.
[0060] like Figure 2 As shown in the figure, this figure shows the structure of the improved auxiliary classifier generation adversarial network model, including:
[0061] The generator multiplies the noise vector by the embedded category label element by element, combines the category information with the noise, and generates numerical data that meets the data distribution and category requirements;
[0062] The present invention replaces the convolutional layers of the generator and discriminator with fully connected layers;
[0063] The specific implementation is as follows:
[0064] The category label Embedded into the same dimensional space as the noise vector, the embedding vector is obtained :
[0065]
[0066] The noise vector Multiply element-wise with the embedded category label to combine the category information with the noise:
[0067]
[0068] The convolutional layers of the generator and discriminator are replaced with fully connected layers, and the noise vector is gradually mapped to the generated sample through a series of fully connected layers and activation function LeakyReLU:
[0069]
[0070] The output is generated by the hyperbolic tangent activation function tanh :
[0071]
[0072] in, and are the weights and biases of the network.
[0073] In specific implementations, the improved ACGAN not only helps generate numerical data of specific categories, but also enhances the generator's ability to learn category features, thereby improving the quality and diversity of generated data;
[0074] Step 3: The present invention uses the Bayesian optimization algorithm to optimize the parameters of the auxiliary classifier generative adversarial network model. Reasonable hyperparameter tuning can significantly improve the generation quality and classification accuracy of the model;
[0075] BOA is a global optimization method for black-box functions with high computational costs. It efficiently explores and utilizes the solution space to find hyperparameter combinations close to the global optimal with fewer evaluations.
[0076] The core idea is to efficiently explore the solution space by building a prior probability model of the objective function and gradually updating this model. In this way, the optimal hyperparameter combination can be found while ensuring computational efficiency, thereby improving the overall performance of the model.
[0077] like Figure 3 As shown in the figure, this is the flow chart of the Bayesian optimization algorithm, including:
[0078] Initial sampling: Randomly sample several sets of hyperparameters and calculate the objective function value to construct the initial data set;
[0079] Constructing a proxy model: Constructing a proxy model of the objective function based on the current data set. The Gaussian process used in this invention is:
[0080]
[0081] in, is the mean function, is the covariance function;
[0082] The kernel function used in the present invention is the radial aggregate function:
[0083]
[0084] in, is the output variance, is the length scale;
[0085] Collect new samples: Use the collection function to find the optimal collection point. The collection function is used to select the next evaluation point. The method used in this paper is the expected improvement:
[0086]
[0087] in, It is the best point currently known;
[0088] Update model: Evaluate the objective function value, add new data to the dataset, and update the Gaussian process model;
[0089] Iterative optimization: Repeat the steps of collecting new samples and updating the model until a predetermined stopping condition is met.
[0090] In this step, the Bayesian optimization algorithm optimizes four parameters of the auxiliary classifier generative adversarial network model: the generator learning rate, the discriminator learning rate, the latent space vector, and the batch size;
[0091] Step 4: The present invention uses the Bayesian optimization algorithm to optimize the parameters of the random forest detection model; the Bayesian optimization algorithm process is the same as the above process;
[0092] In this step, the Bayesian optimization algorithm optimizes four parameters of the random forest detection model, namely: the number of decision trees, the maximum depth of the decision tree, the minimum number of samples required for internal nodes to split again, and the minimum number of samples required for leaf nodes.
[0093] Step 5: Use the balanced data set to train the detection model, and use the trained target intrusion detection model to perform network intrusion detection.
[0094] In summary, this paper proposes a novel intrusion detection method designed to effectively address data imbalance and enhance intrusion detection performance. The method improves the traditional auxiliary classifier generative adversarial network model, enabling it to effectively generate numerical data of specific categories. By combining the Bayesian optimization algorithm with the improved auxiliary classifier generative adversarial network model, the diversity and authenticity of the generated samples are enhanced. Furthermore, by combining the Bayesian optimization algorithm with the random forest detection model, the detection effect is significantly improved.
[0095] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions merely illustrate the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and modifications are intended to fall within the scope of the present invention. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. An intrusion detection method for data imbalance problem, characterized in that: include: Obtain network intrusion data, remove dirty data, and standardize the data; Improve the auxiliary classifier generative adversarial network model by multiplying the noise vector with the embedded category label element by element and replacing the convolutional layers of the generator and discriminator with fully connected layers; The Bayesian optimization algorithm is used to optimize the hyperparameters of the improved auxiliary classifier generative adversarial network model, and the model with the best parameters is used to generate minority class samples to balance the dataset. Use the Bayesian optimization algorithm to optimize the hyperparameters of the random forest detection model and use the balanced dataset to train the optimal detection model; The detection model is trained using a balanced data set, and the trained target intrusion detection model is used for network intrusion detection.
2. The intrusion detection method for data imbalance problem according to claim 1, characterized in that: The preprocessing of the acquired network intrusion data includes: Performing missing value processing on the network intrusion data; Eliminating dirty data from the network intrusion data; Standardize the processed data: in, is the original eigenvalue, is the mean of the feature, is the standard deviation of the feature, is the normalized eigenvalue.
3. The intrusion detection method for data imbalance problem according to claim 2, characterized in that: The improvement of the auxiliary generative adversarial network model includes: The category label Embedded into the same dimensional space as the noise vector, the embedding vector is obtained : The noise vector Multiply element-wise with the embedded category label to combine the category information with the noise: The convolutional layers of the generator and discriminator are replaced with fully connected layers, and the noise vector is gradually mapped to the generated sample through a series of fully connected layers and activation function LeakyReLU: The output is generated by the hyperbolic tangent activation function tanh : in, and are the weights and biases of the network.
4. The intrusion detection method for data imbalance problem according to claim 3, characterized in that: The Bayesian optimization algorithm performs hyperparameter tuning on the auxiliary classifier generative adversarial network model and the random forest detection model, including: Initial sampling: Randomly sample several sets of hyperparameters and calculate the objective function value to construct the initial data set; Constructing a proxy model: Constructing a proxy model of the objective function based on the current data set. The Gaussian process used in this invention is: in, is the mean function, is the covariance function; The kernel function used in the present invention is the radial aggregate function: in, is the output variance, is the length scale; Collect new samples: Use the collection function to find the optimal collection point. The collection function is used to select the next evaluation point. The method used in this paper is the expected improvement: in, It is the best point currently known; Update model: Evaluate the objective function value, add new data to the dataset, and update the Gaussian process model; Iterative optimization: Repeat the steps of collecting new samples and updating the model until a predetermined stopping condition is met.
5. The intrusion detection method for data imbalance problem according to claim 4, characterized in that: The specific parameters of the hyperparameter tuning include: Auxiliary classifier generative adversarial network model: generator learning rate, discriminator learning rate, latent space vector and batch size; Random forest detection model: the number of decision trees, the maximum depth of the decision tree, the minimum number of samples required for internal nodes to split again, and the minimum number of samples required for leaf nodes.
6. The intrusion detection method for data imbalance problem according to claim 5, characterized in that: The target balanced data is used to train a network intrusion detection model to obtain a target network intrusion detection model, including: Determine the number of training times based on the amount of data in the dataset; The network intrusion detection model is trained and verified using a balanced data set to obtain multiple verification results. When all verification results meet the preset requirements, the trained network intrusion detection model is determined to be the target network intrusion detection model.