Table data type adversarial sample generation method based on double-discriminator CTGAN

Through the tabular data-type adversarial sample generation method based on dual discriminator CTGAN, the problems of insufficient adversarial sample universality and feature distribution in the prior art are solved. The generated adversarial samples are suitable for a variety of network intrusion detection systems, improving the defense capability of the model.

CN120281536APending Publication Date: 2025-07-08UNIV OF ELECTRONICS SCI & TECH OF CHINA +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510432763.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The adversarial samples generated by the prior art are relatively low in versatility and cannot be applied to multiple network intrusion detection systems, and the feature distribution is not rich enough.

Method used

The tabular data-type adversarial sample generation method based on the dual discriminator CTGAN is adopted. Through the dual discriminator collaborative optimization mechanism and hyperparameter α control, the generator learns the data distribution and generates tabular data-type adversarial samples with rich features.

Benefits of technology

The generated adversarial samples can be used in a variety of network intrusion detection systems, with richer feature distribution and improved the model's defense capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120281536A_ABST
    Figure CN120281536A_ABST
Patent Text Reader

Abstract

The invention discloses a table data type adversarial sample generation method based on a double-discriminator CTGAN model. Comprising the following steps: preprocessing a data set, specifying an attack type of an adversarial sample needing to be generated, and dividing functional features and non-functional features according to the attack type; constructing and training a double-discriminator CTGAN model, introducing a new discriminator component on the basis of the CTGAN model, and training by respectively inputting a normal traffic sample and a specified attack traffic sample into two discriminators to achieve sample distribution of a collaborative optimization generator; performing adversarial sample generation by using the trained double-discriminator CTGAN model; and performing quality evaluation on the generated adversarial sample by using a network intrusion detection system based on machine learning. According to the method, the table data type adversarial sample is generated by using the double-discriminator CTGAN model, and the generated adversarial sample is suitable for various network intrusion detection system models, so that a network intrusion detection system based on machine learning can be effectively helped to improve the detection precision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network traffic intrusion detection, and specifically relates to a method for generating tabular data type adversarial samples based on a dual discriminator CTGAN. Background Art

[0002] A network intrusion detection system is a tool for monitoring network traffic, aiming to detect malicious network traffic in a timely manner and make interceptions. Currently, the mainstream network intrusion detection systems use machine learning or deep learning to build intrusion detection models, and predict whether subsequent traffic belongs to malicious traffic by learning the feature distribution rules existing in a large amount of network traffic data. However, such machine learning-based network intrusion detection systems (ML-NIDS) have some defects, that is, the quality of the training dataset directly determines the accuracy of the model discrimination boundary. With the development of network attack and defense technologies, some attackers use the low accuracy of the model discrimination boundary to fine-tune the features of malicious traffic to make the network intrusion detection system misjudge malicious traffic as normal traffic (referred to as adversarial attack). In order to improve the reliability of the model, it is of great significance to enhance the defense ability of the network intrusion detection model against adversarial attacks.

[0003] Chinese Patent CN118337526B discloses a method for generating adversarial attack samples. An improved Wasserstein GAN network designed independently is used to learn the features of benign traffic, so as to be used to disguise malicious traffic to generate adversarial attack samples, and thus use the adversarial attack samples to better test and upgrade the machine learning-based network intrusion detection system. The improved Wasserstein GAN network introduces a multi-generator structure into the Wasserstein GAN network, and adds a distortion rate to the generator loss function.

[0004] However, the above method has certain deficiencies. The adversarial samples generated by the improved Wasserstein GAN in this method are in numerical format and have low generality. For example, this method generates adversarial samples targeted at a certain ML-NIDS, and the adversarial samples generated for ML-NIDS1 cannot be input into ML-NIDS2 for detection. Therefore, how to generate a more general adversarial sample (that is, applicable to multiple ML-NIDS) has become a research hotspot. Summary of the Invention

[0005] In view of the deficiencies of the adversarial samples generated by the existing model, the present invention independently designs a method for generating tabular data type adversarial samples based on a dual discriminator CTGAN, which can generate tabular data type adversarial samples with higher generality and richer feature distributions.

[0006] The technical content of the present invention includes:

[0007] A method for generating adversarial samples of tabular data type based on a dual-discriminator CTGAN, characterized in that the method comprises the following steps:

[0008] S1. Training of the adversarial sample generation model:

[0009] S11. Obtain a tabular data type network traffic dataset, specify the types of adversarial attack samples to be generated, and extract normal traffic samples and specified attack traffic samples from the dataset.

[0010] S12. According to the specified attack types, use expert knowledge to divide the extracted attack traffic samples and normal traffic samples into functional features and non-functional features.

[0011] S13. Input the non-functional features of the normal traffic samples and the specified attack traffic samples into the dual-discriminator CTGAN for training and learning.

[0012] The dual-discriminator CTGAN model is specifically embodied in introducing a dual-discriminator collaborative optimization mechanism, changing the loss function of the generator from single discriminator feedback to dual-discriminator collaborative feedback, and the feedback weight ratio is controlled by the hyperparameter α. By adjusting the hyperparameter α, the data distribution learned by the generator G is controlled.

[0013] S2. Generation of tabular data type adversarial samples:

[0014] S21. Use the trained model to generate tabular data type adversarial samples. Input the non-functional features of the attack traffic samples into the dual-discriminator CTGAN model, and modify the non-functional features of the input attack traffic through the feature learning of the existing traffic by the dual-discriminator CTGAN (abbreviation: feature disguise).

[0015] S22. Concatenate the non-functional features of the attack traffic after feature disguise with the functional features of the original attack traffic samples. The data after concatenation is the adversarial attack sample.

[0016] S3. Verify the quality of the generated tabular data type adversarial samples:

[0017] S31. Input the adversarial attack samples into the ML-NIDS for detection. If they can be recognized as normal samples, it means that such samples can successfully escape the detection of the ML-NIDS. Then use such samples as the training set to provide adversarial training for the ML-NIDS.

[0018] S32. Use the ML-NIDS that has completed adversarial training in S31 to re-detect the attack traffic samples. If the detection rate is improved, it proves that the quality of the generated tabular data type adversarial samples meets the standard.

[0019] Advantages of the present invention:

[0020] In view of the deficiencies of the adversarial samples generated by existing models, the present invention independently designs a method for generating tabular data type adversarial samples based on a dual discriminator CTGAN. The adversarial samples generated by this method have the following two advantages: (1) Higher generality. Existing adversarial sample generation methods generate adversarial samples for a specific ML-NIDS, and the generated samples cannot be input into other ML-NIDS for detection. Specifically, the numerical adversarial samples generated for ML-NIDS1 cannot be input into ML-NIDS2 for detection. However, the adversarial samples generated by the present invention are in the same tabular data format as intrusion detection datasets such as KDD99 and NSL-KDD, and can be applied to multiple ML-NIDS. (2) The distribution of adversarial samples is richer. By adjusting the hyperparameter α, the characteristics of the existing data distribution learned by the generator G can be controlled, and then the degree of feature camouflage of the attack traffic samples can be controlled. Aggregating the samples generated by the models under different α can obtain an adversarial sample set with a richer distribution. Description of the drawings

[0021] Figure 1 It is the overall flowchart of the method for generating tabular data type adversarial samples proposed by the present invention;

[0022] Figure 2 It is the architecture diagram and usage example diagram of the tabular data type adversarial sample generation model of the dual discriminator CTGAN described in the present invention;

[0023] Figure 3 For Figure 2 Symbol description in the model architecture diagram;

[0024] Figure 4 It is the feature classification diagram in the NSL-KDD dataset;

[0025] Figure 5 It is the feature division of 4 types of attacks in the NSL-KDD dataset (tick for functional features, and correspondingly no tick for non-functional features);

[0026] Figure 6 It is the quality assessment result of the DoS tabular data type adversarial samples generated by the method of the present invention; Specific implementation manners

[0027] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Apparently, the described embodiments are only specific embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0028] Figure 2 The figure shows a schematic diagram of the training and usage process of the dual-discriminator CTGAN model provided by an embodiment of the present invention. As shown in the figure, the method includes:

[0029] S1. Training of the adversarial sample generation model:

[0030] S11. Obtain a network traffic dataset of table data types, specify the types of adversarial attack samples to be generated, and extract normal traffic samples and specified attack traffic samples from the dataset.

[0031] S12. According to the specified attack types, use expert knowledge to divide the extracted attack traffic samples and attack traffic samples into functional features and non-functional features.

[0032] S13. Input the non-functional features of the normal traffic samples and the specified attack traffic samples into the dual-discriminator CTGAN for training and learning.

[0033] The dual-discriminator CTGAN model is specifically embodied in introducing a dual-discriminator collaborative optimization mechanism, changing the loss function of the generator from being feedback by a single discriminator to being collaboratively feedback by dual discriminators. The feedback weight ratio is controlled by the hyperparameter α. By adjusting the hyperparameter α, the data distribution learned by the generator G is controlled.

[0034] S2. Generation of table data type adversarial samples;

[0035] S21. Use the trained model to generate table data type adversarial samples. Input the non-functional features of the attack traffic samples into the dual-discriminator CTGAN model, and modify the non-functional features of the input attack traffic through the feature learning of the dual-discriminator CTGAN for the features of the existing traffic (abbreviated as feature camouflage).

[0036] S22. Concatenate the non-functional features of the attack traffic after feature camouflage in S21 with the functional features of the original attack traffic samples. The data after concatenation is the adversarial attack sample.

[0037] S3. Verify the quality of the generated table data type adversarial samples:

[0038] S31. Input the adversarial attack samples into the ML-NIDS for detection. If they can be recognized as normal samples, it indicates that such samples can successfully escape the ML-NIDS detection. Then, use such samples as the training set to provide adversarial training for the ML-NIDS.

[0039] S32. Use the ML-NIDS that has completed adversarial training in S31 to re-detect the attack traffic samples. If the detection rate is improved, it proves that the quality of the generated tabular data type adversarial samples meets the standard.

[0040] The embodiments of the present invention mainly include three: Embodiment 1 introduces the training process of the tabular data type adversarial sample generation model; Embodiment 2 introduces using the trained model to complete the generation of tabular data type adversarial samples; Embodiment 3 introduces the quality evaluation of the generated tabular data type adversarial samples.

[0041] Embodiment 1

[0042] Regarding step S11, here the NSL-KDD intrusion detection traffic dataset is selected. This dataset contains 4 types of attack categories. For the convenience of explanation, it is specified here that the type of adversarial sample to be generated is DoS, that is, extract the samples with labels of DoS and normal traffic.

[0043] Regarding step S12, the features mentioned therein are divided into functional features with high semantic relevance to the attack traffic and non-functional features with low semantic relevance to the attack traffic. The division of feature fields and traffic semantics is as Figure 2 shown. Specifically, regarding the DoS attack, because its semantics is to send a large number of requests to the target system to exhaust its resources or paralyze it, so the features related to time series and expressing network traffic protocol specifications such as source IP address, destination IP address, and port number are divided into functional features, and the content and host-based feature information are divided into non-functional information. The feature division for other attack categories in the NSL-KDD is as Figure 5 shown.

[0044] Regarding the dual discriminator CTGAN model described in step S13, specifically:

[0045] As Figure 2 shown in the model training stage, the content of the dual discriminator is to introduce a new discriminator component D2 to achieve the purpose of synergistically optimizing the loss function of the generator G. G represents the generator, D1 and D2 are discriminators. D1 is used to discriminate the similarity degree of non-functional features between normal traffic samples and generated samples, and D2 is used to discriminate the similarity degree of non-functional features between attack traffic samples and generated samples. And by setting the hyperparameter α to control the feedback weight ratio of D1 and D2 to the loss function of G, Figure 2 The meanings expressed by the symbols in Figure 3 are asnff A non-functional feature representing normal traffic (normal non-functional feature), m nff A non-functional feature representing malicious traffic (malicious non-functional feature).

[0046] The loss function of the generator G is as follows:

[0047]

[0048] The loss function of the generator G consists of three parts, namely the feedback from D1 and D2 and the cross-entropy conditional loss L of the discrete columns cond , where represents the joint expectation of the noise vector z ∼ N(0, I) and the discrete conditional vector c ∼ P c , α ∈ (0, 1) controls the weights of the two discriminators, is the cross-entropy loss function of the discrete columns, used to constrain the discrete feature distribution of the generated samples.

[0049] The loss functions of the discriminators D1 and D2 are as follows:

[0050]

[0051] In the formula, x ∼ p normal and x' ∼ p mal represent the data distributions of normal samples and malicious samples respectively, and are mixed samples generated by the linear interpolation coefficient, λ gp = 10 is the coefficient of the fixed gradient penalty term, ||·|| represents the L2 norm, represents the gradient operator for the interpolated samples, L D1 is used to train the discriminator D1 to enable it to distinguish the difference between normal traffic and the generated samples, L D( is used to train the discriminator D2 to enable it to distinguish the difference between malicious traffic and the generated samples.

[0052] Example 2,

[0053] For step S21, use the model trained in Example 1 to perform feature camouflage on the DoS attack traffic samples in the test set.

[0054] For step S22, splice the non-functional features of the camouflaged DoS samples with the functional features of the original DoS samples. The specific splicing process is to replace the corresponding fields according to the feature division implemented in step S12. The data after splicing is the adversarial attack sample.

[0055] Example 3

[0056] For step S31, machine learning models such as the fully connected neural network MLP, logistic regression, and K-nearest neighbors can all be used as the ML-NIDS for detection. Here, the MLP model is selected. The generated DoS adversarial samples are input into the trained ML-NIDS model for detection, and the specified evaluation metric is the detection rate. If the detection rate of the ML-NIDS for adversarial DoS attacks is lower than the detection rate of DoS attacks in the original test set, it indicates that the feature camouflage performed by the double discriminator CTGAN is effective. As Figure 6 shown, in terms of the detection rate results for DoS attacks, the escape rate of the DoS tabular data type adversarial samples generated by the present invention is superior to other methods.

[0057] For step S32, the DoS adversarial samples obtained in S31 are used as the training set to input into the ML-NIDS for adversarial training. After the training is completed, the DoS attack traffic in the dataset is detected again. If the detection rate increases, it proves that the quality of the generated tabular data adversarial samples meets the standard.

Claims

1. A method for generating adversarial samples of tabular data based on a dual-discriminator CTGAN, characterized in that The method includes the following steps: S1. Data preprocessing and model training. Specify the types of adversarial attack samples to be generated, extract normal traffic and specified attack traffic, and train the adversarial sample generation model; S2. Use the trained model to generate tabular data type adversarial samples; S3. Check the quality of the generated tabular data type adversarial samples.

2. A method for generating tabular data adversarial samples based on a dual discriminator CTGAN according to claim 1, characterized in that, Step S1 further includes the following steps: S11. Obtain the tabular data type network traffic dataset, specify the types of adversarial attack samples to be generated, and extract normal traffic samples and specified attack traffic samples from the dataset. S12. According to the specified attack types, use expert knowledge to divide the extracted attack traffic samples and normal traffic samples into functional features and non-functional features. S13. Input the non-functional features of normal traffic samples and specified attack traffic samples into the dual discriminator CTGAN for training and learning. The dual discriminator CTGAN model is specifically embodied in introducing a dual discriminator collaborative optimization mechanism, changing the loss function of the generator from single discriminator feedback to dual discriminator collaborative feedback, and the feedback weight ratio is controlled by the hyperparameter a. By adjusting the hyperparameter α, the data distribution learned by the generator G is controlled.

3. A method for generating tabular data adversarial samples based on a dual discriminator CTGAN according to claim 1, characterized in that, Step S2 further includes the following steps: S21. Input the non-functional features of the attack traffic samples into the dual discriminator CTGAN model, and modify the non-functional features of the input attack traffic through the feature learning of the existing traffic by the dual discriminator CTGAN (abbreviated as feature camouflage). S22. Concatenate the non-functional features of the attack traffic after feature camouflage with the functional features of the original attack traffic samples, and the data after concatenation is the adversarial attack sample.

4. A method for generating tabular data adversarial samples based on a dual discriminator CTGAN according to claim 1, characterized in that, Step S3 further includes the following steps: S31. Input the adversarial attack samples into ML-NIDS for detection. If they can be recognized as normal samples, it means that such samples can successfully escape the detection of ML-NIDS. Then, use such samples as the training set to provide adversarial training for ML-NIDS. S32. Use the ML-NIDS that has completed adversarial training to re-detect the attack traffic samples. If the detection rate is improved, it proves that the quality of the generated tabular data type adversarial samples meets the standard.

Citation Information

Patent Citations

  • A method for generating adversarial attack samples

    CN118337526B