Botnet traffic detection method based on anti-convolution auto-encoder

Through the unsupervised learning framework of anti-convolution autoencoder and the sliding window dynamic threshold, the problem of scarcity and complex feature recognition of label data in botnet traffic detection is solved, and efficient and accurate botnet traffic detection is achieved.

CN120238357APending Publication Date: 2025-07-01NANJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510456431.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The prior art is difficult to detect botnet traffic efficiently and accurately, especially when label data and complex features are lacking. Traditional methods have poor detection results, and traditional machine learning and deep learning methods have limitations in botnet traffic detection.

Method used

An unsupervised learning framework based on adversarial convolutional autoencoder is adopted. Through the combination of convolutional autoencoder and generative adversarial network, an unsupervised learning framework is used to perform botnet traffic detection, and the adversarial training of encoder and discriminator is used to optimize the quality of hidden vectors, and a sliding window dynamic threshold adaptive adjustment of classification boundaries is introduced during the test phase.

Benefits of technology

It realizes efficient and accurate identification of botnet traffic in the absence of label data, improves detection accuracy and robustness, reduces false positive rates, and adapts to traffic fluctuations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120238357A_ABST
    Figure CN120238357A_ABST
Patent Text Reader

Abstract

The invention designs a botnet traffic detection method based on an adversarial convolution auto-encoder, which comprises the following steps of: firstly, cutting original network traffic into IP (Internet Protocol) with the same source and target, and then processing byte streams into a data form of 32 bytes multiplied by 32 bytes; secondly, a network model composed of a convolution encoder, a deconvolution decoder and a discriminator is constructed, the encoder is used for mapping input data to an implicit vector, and prior distribution is set as Gaussian distribution; in the training process, the quality of implicit vectors generated by an encoder is continuously optimized through confrontation of the encoder and a discriminator, finally, during testing, an initial threshold value is determined according to a mean value and a standard deviation of training data reconstruction loss, and a sliding window dynamic threshold value is introduced to adaptively adjust a classification boundary, so that the classification accuracy is improved. And detecting whether the input data belongs to normal flow or botnet flow by comparing the reconstruction loss obtained after the input data passes through the encoder and the decoder with the dynamic threshold value in the current sliding window.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network security, and specifically to a method for detecting botnet traffic based on adversarial convolutional autoencoders. Background Art

[0002] With the rapid development of the Internet and the popularization of Internet of Things (IoT) devices, the network environment has become increasingly complex, and the traffic between user devices and servers has increased significantly. In this process, network security issues have become increasingly severe, and various malicious network activities have shown more complex and hidden characteristics, especially the widespread application of botnets. The attack methods of botnets not only cause great harm to individual and enterprise users, but also threaten the overall security and stability of the Internet. Therefore, how to efficiently and accurately detect and defend against botnet traffic has become an important topic in network security research.

[0003] A botnet refers to a network composed of a large number of maliciously controlled devices. Attackers use these devices to launch various network attacks remotely, such as DDoS attacks, spam sending, data theft, etc. With the popularization of IoT devices, the scale and complexity of botnets have increased significantly, making their attack scope wide, destructive, and difficult to trace. Since the communication and operations of botnets are usually mixed in normal traffic, their concealment is extremely strong, and traditional network protection technologies often have difficulty effectively detecting such malicious traffic, especially when botnets adopt complex evasion techniques.

[0004] Regarding the detection of botnet traffic, although various methods have been proposed, such as rule-based detection, traditional machine learning methods, and deep learning techniques, these methods still face challenges when dealing with the constantly changing attack patterns of botnets. Rule-based detection methods rely on pre-set attack features and are difficult to adapt to the rapid changes in botnet attack patterns. Traditional machine learning methods, such as decision trees and support vector machines, although improving the detection flexibility to a certain extent, are still limited by the accuracy and comprehensiveness of manual feature extraction. Deep learning techniques, especially convolutional neural networks and recurrent neural networks, although able to automatically extract features and show good performance, are often limited in dealing with the actual situation of lack of labeled data (especially botnet traffic samples). In addition, these methods still have deficiencies in capturing the complex spatio-temporal features of botnet traffic.

[0005] To overcome the limitations of existing methods in botnet traffic detection, especially when dealing with the problem of lack of labeled data, the present invention proposes a botnet traffic detection method based on an adversarial convolutional autoencoder. This method utilizes an unsupervised learning framework and only requires normal traffic data during the training phase, thus solving the problem of scarce botnet traffic samples in real-world scenarios. The adversarial convolutional autoencoder combines the autoencoding structure of a convolutional neural network and the adversarial training mechanism of a generative adversarial network. Among them, the convolutional encoder is used to map the input data into a low-dimensional latent vector space, and the deconvolutional decoder attempts to reconstruct the original data from the latent vector. At the same time, a discriminator is introduced to distinguish whether the source of the latent vector is generated by the encoder or from a preset prior distribution. Through the adversarial training between the encoder and the discriminator, the quality of the latent vector generated by the encoder is continuously optimized, and at the same time, the reconstruction ability of the decoder is improved to minimize the reconstruction error. In the testing phase, an initial threshold is determined based on the mean and standard deviation of the reconstruction loss of the training data. A sliding window dynamic threshold is introduced to adaptively adjust the classification boundary. By comparing the reconstruction loss of the input data after passing through the encoder and the decoder with the dynamic threshold within the current sliding window, it is detected whether the input data belongs to normal traffic or botnet traffic. This method can not only effectively utilize limited normal traffic data for model training, but also accurately distinguish normal traffic and botnet traffic in the detection phase, providing a new solution for the efficient and accurate detection of botnet traffic. Summary of the Invention

[0006] To address the deficiencies of existing botnet traffic detection methods in dealing with the lack of labeled data, difficulty in capturing complex features, and inaccurate identification of covert attacks, the present invention proposes a botnet traffic detection method based on an adversarial convolutional autoencoder. This method utilizes an unsupervised learning framework and effectively solves the problem of scarce botnet traffic samples by combining the feature extraction ability of a convolutional autoencoder and the adversarial training mechanism of a generative adversarial network, and significantly improves the accuracy and robustness of detection. By optimizing the generation of latent vectors and the reconstruction process of data, and introducing a sliding window dynamic threshold during testing, the present invention can more accurately identify botnet traffic hidden in normal traffic.

[0007] To achieve the above object, the present invention provides the following technical solutions:

[0008] A botnet traffic detection method based on an adversarial convolutional autoencoder, comprising the following steps:

[0009] Step 1: Data preprocessing. Cut the original network traffic data into byte streams with the same source IP, destination IP, source port, destination port, and protocol to ensure the consistency and comparability of the data.

[0010] Step 2: Data normalization. Process the cut byte stream into a form of 32 bytes by 32 bytes.

[0011] Step 3: Model structure initialization. Define three main modules, including an encoder: mapping the input data to a latent vector, which is a convolutional structure; a decoder: reconstructing the input data from the latent vector, which is a transposed convolutional structure; and a discriminator: distinguishing whether the source of the latent vector is generated by the encoder or from a preset prior distribution. Finally, the prior distribution is determined to be a Gaussian distribution.

[0012] Step 4: Model training. First, perform forward propagation on each batch of data, including encoder processing, decoder reconstruction, and discriminator judgment; secondly, calculate the losses, including reconstruction loss, discriminator loss, and encoder adversarial loss; finally, update the parameters to minimize the reconstruction loss and balance the discriminator loss and encoder adversarial loss.

[0013] Step 5: Model testing. Input the test data, perform forward propagation and reconstruction, calculate the reconstruction error, and compare the reconstruction error with the dynamic threshold within the current sliding window to determine whether the input data is normal traffic or botnet traffic.

[0014] In the said Step 1, use a network flow splitting tool to cut the original traffic data in the form of five-tuples, that is, each traffic has the same source IP address, destination IP address, source port, destination port, and protocol. Then, an original traffic file can be decomposed into numerous flows f1, f2,... f with unique five-tuple information i .

[0015] In the said Step 2, for each f, perform data form conversion, intercept the first 1024 bytes, and convert the sequentially arranged 1024 bytes into a two-dimensional matrix form with 32 bytes per row and 32 columns. If f is less than 1024 bytes, pad it with 0x00. To eliminate the differences in dimension and data size, perform on each byte b min where the hexadecimal number b max = 0, b

[0016] In the said Step 3, it specifically includes the following steps:

[0017] Step 3.1: Construct an encoder, and its network structure is specifically as follows:

[0018] The first layer of convolution: the input channel is 1, the output channel is 16, the convolution kernel size is 3*3, the stride is 2, the padding is 1, and the output feature map size is 16*16*16.

[0019] ReLU activation function: introduce non-linearity to the convolution output.

[0020] Second - layer Convolution: The input channels are 16, the output channels are 32, the convolution kernel size is 3*3, the stride is 2, the padding is 1, and the output feature map size is 32*8*8.

[0021] ReLU Activation Function: Introduce non - linearity again.

[0022] Flattening Layer: Expand the three - dimensional feature map of 32*8*8 into a one - dimensional vector with 2048 features.

[0023] Fully - connected Layer: The input features are 2048, and the output is a low - dimensional hidden vector, whose dimension is set to 64 in this invention. Step 3.2: Construct the decoder, and its network structure is as follows:

[0024] Fully - connected Layer: The input features are the hidden vector dimension 64, and the input features are 32*8*8, that is, 2048. The purpose is to restore the hidden vector to the shape required for convolution.

[0025] ReLU Activation Function: Introduce non - linearity.

[0026] Unflattening Layer: Restore the one - dimensional vector 2048 back to a three - dimensional tensor: 32*8*8.

[0027] First - layer Transposed Convolution: The input channels are 32, the output channels are 16, the convolution kernel size is 3*3, the stride is 2, the padding is 1, and the output feature map size is 16*16*16.

[0028] ReLU Activation Function: Introduce non - linearity.

[0029] Second - layer Transposed Convolution: The input channels are 16, the output channels are 1, the convolution kernel size is 3*3, the stride is 2, the padding is 1, and the output feature map size is 1*32*32.

[0030] Sigmoid Activation Function: Compress the output data into the range [0, 1], which is suitable for the case where the input data is a normalized image.

[0031] Step 3.3: Construct the discriminator, and its network structure is as follows:

[0032] First - layer Fully - connected: The input features are the hidden vector dimension 64, and map the input hidden vector to a hidden layer with a dimension of 128.

[0033] ReLU Activation Function: Introduce non - linearity.

[0034] Second - layer Fully - connected: The input features are 128 - dimensional, and the output features are 1, mapping the output of the hidden layer to a single scalar representing the discrimination probability.

[0035] Sigmoid activation function: Compresses the output to the range [0, 1], representing the probability of being judged as a real latent vector. Step 3.4: Determine that the prior distribution is a Gaussian distribution, and its formula is

[0036]

[0037] In step 4, it specifically includes the following steps:

[0038] Step 4.1: Forward propagation:

[0039] For each batch of data The encoder E processes the input data x to generate a latent vector z, that is

[0040] z = E(x)

[0041] The decoder G reconstructs the latent vector z and reconstructs the output data That is

[0042]

[0043] The discriminator D randomly samples a latent vector z from the prior distribution p(z) real , and processes two inputs, the real latent vector z real (from the prior distribution), and generates a latent vector z (from the encoder E). At the same time, it outputs the discrimination result, the probability of being judged as real D(z real ), and the probability of being judged as generated D(z).

[0044] Step 4.2: Loss calculation:

[0045] Calculate the reconstruction loss, that is, calculate the error between the input data x and the reconstructed data :

[0046]

[0047] The discriminator attempts to distinguish between the real latent vector z real and the generated latent vector z, and its loss is:

[0048]

[0049] The encoder attempts to "deceive" the discriminator so that z is judged as a real latent vector, and its adversarial loss is:

[0050] L E = -E z [log D(z)]

[0051] Step 4.3: Parameter update:

[0052] Update the discriminator D to minimize the discriminator loss L D .

[0053] Update the encoder E and decoder G to minimize the reconstruction loss L rec and the encoder adversarial loss L E :

[0054] L = L rec + λL E

[0055] where λ is a weight coefficient used to balance the two losses.

[0056] Finally, perform iterative training, repeating the above steps until the network converges.

[0057] In step 5, it specifically includes the following steps:

[0058] Step 5.1: Set the dynamic threshold of the sliding window, specifically:

[0059] Determine the initial threshold according to the mean and standard deviation of the reconstruction loss of normal traffic in the training phase, that is

[0060] τ init = μ train + k·σ train

[0061] where μ train and σ train are the mean and standard deviation of the training data respectively, and k is an adjustment coefficient.

[0062] Perform local statistics calculation of the sliding window, maintain a sliding window with a length of W, and store the reconstruction loss values of the last W test samples. Calculate the local mean and standard deviation within the window in real time, that is

[0063]

[0064] Combine the initial threshold and local statistics to dynamically adjust the current threshold, that is

[0065] τ t = β·τ init +(1 - β)·(μ window + k·σ window )

[0066] where β is a weight coefficient used to balance the influence of global statistics and local statistics.

[0067] Step 5.2: Input the test data X test , including normal traffic and botnet traffic, and call the model for anomaly detection, specifically:

[0068] The test data x test passes through the encoder E to generate the hidden vector ztest , namely

[0069] z test = E(x test )

[0070] The latent vector z test passes through the decoder G to reconstruct the data namely

[0071]

[0072] Calculate the reconstruction error between the test data x test and the reconstructed data :

[0073]

[0074] For normal traffic, since the encoder generates high-quality z, the reconstruction error is small. For botnet traffic, since the model has never seen it before, the generated z does not conform to the prior distribution, resulting in a significantly larger reconstruction error.

[0075] If L rec,test > τ t , it is determined as botnet traffic. If L rec,test < τ t , it is determined as normal traffic.

[0076] An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements the described botnet traffic detection method based on an adversarial convolutional autoencoder when executing the program.

[0077] A computer-readable storage medium, on which a computer instruction is stored, and the computer instruction implements the described botnet traffic detection method based on an adversarial convolutional autoencoder when executed by a processor.

[0078] Compared with the prior art, the botnet traffic detection method based on an adversarial convolutional autoencoder proposed by the present invention has the following beneficial effects:

[0079] 1. The unsupervised framework of the present invention combines the adversarial training mechanism of a convolutional autoencoder and a generative adversarial network. The model only needs to be trained with normal traffic and can identify botnet traffic through the reconstruction error. It can learn the feature distribution of normal traffic without relying on a large amount of labeled data, avoiding the limitation of insufficient labeled data.

[0080] 2. Optimize the encoder through the discriminator adversarial loss to make the latent vector follow a Gaussian distribution, enhancing the sensitivity of the model to abnormal traffic. The latent vector of an ordinary autoencoder may overfit the training data, while adversarial training improves the generalization ability through distribution constraints.

[0081] 3. Introduce a sliding window dynamic threshold during testing to solve the misjudgment problem of the fixed threshold during traffic fluctuations. The fixed threshold of the traditional method is prone to a high false alarm rate, and the dynamic threshold significantly improves the environmental adaptability. Description of the Drawings

[0082] Figure 1 It is a schematic diagram of the process of the present invention.

[0083] Figure 2 It is a structural diagram of the model of the present invention. Detailed Implementation Manner

[0084] Next, in combination with the drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0085] Embodiment: In combination with Figure 1 , a zombie network traffic detection method based on an adversarial convolutional autoencoder, the method includes the following steps.

[0086] Step 1: Data preprocessing.

[0087] In a specific embodiment, use the open-source network flow analysis tool Splitcap to parse and cut the original network traffic data according to the conditions of the unique source and destination IP, the unique source and destination ports, and the protocol, that is, each piece of traffic data f after cutting has a unique five-tuple.

[0088] Step 2, Data normalization.

[0089] In a specific embodiment, take the size of the data form as a side length of 32 bytes, that is, intercept 1024 bytes, and then arrange these bytes into a two-dimensional matrix form, that is If the total length of f is less than 1024 bytes, it is padded with null bytes 0x00.

[0090] Since one byte is composed of 8 bits, its decimal representation has a total of 2 8 = 256 possibilities, that is, 0 - 255. To eliminate the influence of data size on the calculation, compress the value of all bytes to [0, 1], that is, byte / 255. To eliminate the influence of data size on the calculation, execute byte / 255 to compress each value to [0, 1].

[0091] Step 3: Model structure initialization.

[0092] In a specific embodiment, in combination with Figure 2, construct an initialized model according to the following network structure:

[0093] Construct an encoder, and its network structure is specifically as follows:

[0094] The first convolutional layer: The input channel is 1, the output channel is 16, the convolutional kernel size is 3*3, the stride is 2, the padding is 1, and the output feature map size is 16*16*16.

[0095] ReLU activation function: Introduce non-linearity to the convolutional output.

[0096] The second convolutional layer: The input channel is 16, the output channel is 32, the convolutional kernel size is 3*3, the stride is 2, the padding is 1, and the output feature map size is 32*8*8.

[0097] ReLU activation function: Introduce non-linearity again.

[0098] Flattening layer: Expand the three-dimensional feature map of 32*8*8 into a one-dimensional vector with 2048 features.

[0099] Fully connected layer: The input feature is 2048, and the output is a low-dimensional hidden vector, whose dimension is set to 64 in the present invention. Construct a decoder, and its network structure is as follows:

[0100] Fully connected layer: The input feature is the hidden vector dimension 64, and the input feature is 32*8*8, that is, 2048. The purpose is to restore the hidden vector to the shape required for convolution.

[0101] ReLU activation function: Introduce non-linearity.

[0102] Unflattening layer: Restore the one-dimensional vector 2048 back to a three-dimensional tensor: 32*8*8.

[0103] The first transposed convolutional layer: The input channel is 32, the output channel is 16, the convolutional kernel size is 3*3, the stride is 2, the padding is 1, and the output feature map size is 16*16*16.

[0104] ReLU activation function: Introduce non-linearity.

[0105] The second transposed convolutional layer: The input channel is 16, the output channel is 1, the convolutional kernel size is 3*3, the stride is 2, the padding is 1, and the output feature map size is 1*32*32.

[0106] Sigmoid activation function: Compress the output data into the range of [0,1], which is suitable for the case where the input data is a normalized image.

[0107] Construct a discriminator, and its network structure is as follows:

[0108] The first fully connected layer: The input feature is a latent vector with a dimension of 64, and the input latent vector is mapped to a hidden layer with a dimension of 128.

[0109] ReLU activation function: Introduce non-linearity.

[0110] The second fully connected layer: The input feature is 128-dimensional, and the output feature is 1. The output of the hidden layer is mapped to a single scalar, representing the discrimination probability.

[0111] Sigmoid activation function: Compress the output to the range [0, 1], representing the probability of discriminating as a real latent vector. Determine the prior distribution as a Gaussian distribution, and its formula is

[0112]

[0113] Step 4: Model training.

[0114] In a specific embodiment, forward propagation, loss calculation, and parameter update are performed respectively and iterated for multiple rounds.

[0115] Forward propagation:

[0116] For each batch of data The encoder E processes the input data x to generate a latent vector z, that is

[0117] z = E(x)

[0118] The decoder G reconstructs the latent vector z and reconstructs the output data That is

[0119]

[0120] The discriminator D randomly samples a latent vector z from the prior distribution p(z) real , and processes two inputs, the real latent vector z real (from the prior distribution), and the generated latent vector z (from the encoder E). At the same time, the discrimination result is output, the probability of discriminating as real D(z real ), and the probability of discriminating as generated D(z).

[0121] Loss calculation:

[0122] Calculate the reconstruction loss, that is, calculate the error between the input data x and the reconstructed data :

[0123]

[0124] The discriminator tries to distinguish between the real latent vector z real and the generated latent vector z, and its loss is:

[0125]

[0126] The encoder tries to "deceive" the discriminator so that z is judged as a real latent vector, and its adversarial loss is:

[0127] L E = -E z [log D(z)]

[0128] Parameter update:

[0129] Update the discriminator D to minimize the discriminator loss L D .

[0130] Update the encoder E and decoder G to minimize the reconstruction loss L rec and the encoder adversarial loss L E :

[0131] L = L rec + λL E

[0132] where λ is the weight coefficient used to balance the two losses.

[0133] Finally, perform iterative training and repeat the above steps until the network converges.

[0134] Step 5: Model testing.

[0135] In a specific embodiment, after inputting the test data, the model is called, and a sliding window dynamic threshold is introduced to enable the model to adaptively adjust the classification boundary and compare the reconstruction loss of the test data with the dynamic threshold.

[0136] Set the sliding window dynamic threshold, which is specifically:

[0137] Determine the initial threshold according to the mean and standard deviation of the reconstruction loss of the normal traffic in the training phase, that is

[0138] τ init = μ train + k·σ train

[0139] where μ train and σ train are the mean and standard deviation of the training data respectively, and k is the adjustment coefficient.

[0140] Perform local statistics calculation of the sliding window, maintain a sliding window with a length of W, and store the reconstruction loss values of the recent W test samples. Calculate the local mean and standard deviation within the window in real time, that is

[0141]

[0142]

[0143] Dynamically adjust the current threshold by combining the initial threshold and local statistics, that is

[0144] τ t = β·τ init +(1 - β)·(μ window + k·σ window )

[0145] where β is a weight coefficient used to balance the influence of global statistics and local statistics.

[0146] Input the test data X test , including normal traffic and botnet traffic, and call the model for anomaly detection, specifically:

[0147] The test data x test Passes through the encoder E to generate the latent vector z test , that is

[0148] z test = E(x test )

[0149] The latent vector z test Passes through the decoder G to reconstruct the data That is

[0150]

[0151] Calculate the reconstruction error between the test data x test and the reconstructed data :

[0152]

[0153] For normal traffic, since the encoder generates high-quality z, the reconstruction error is small. For botnet traffic, since the model has never seen it before, the generated z does not conform to the prior distribution, resulting in a significantly larger reconstruction error.

[0154] If L rec,test > τ t , it is determined as botnet traffic. If L rec,test < τ t , it is determined as normal traffic.

[0155] It should be noted that the above embodiments are not used to limit the protection scope of the present invention. Equivalent transformations or substitutions made on the basis of the above technical solutions fall within the protection scope of the claims of the present invention.

Claims

1. A botnet traffic detection method based on adversarial convolutional autoencoder, characterized in that: The following steps are involved: Step 1: Cut the original network traffic data into byte streams with the same source IP, destination IP, source port, destination port and protocol; Step 2: normalize the data form of the cut byte stream; Step 3: Initialize the model structure, define three main modules, convolutional encoder, deconvolutional decoder, and discriminator, and set the prior distribution to Gaussian distribution; Step 4: Perform forward propagation on each batch of training data, calculate the loss and parameter update, and iterate multiple times to train the model; Step 5: Input the test data, perform forward propagation and reconstruction, calculate the reconstruction error, and introduce the sliding window dynamic threshold as the classification condition. Compare the reconstruction error and the current dynamic threshold to determine whether the input data is normal traffic or botnet traffic.

2. The botnet traffic detection method based on adversarial convolutional autoencoder according to claim 1 is characterized in that: In step 1, the byte stream has unique five-tuple information, including source IP address, destination IP address, source port, destination port, and protocol. An original traffic file is decomposed into many flows f1, f2, ... f with unique five-tuple information. i .

3. The botnet traffic detection method based on adversarial convolutional autoencoder according to claim 1 is characterized in that: In step 2, the spatial byte stream is converted into a two-dimensional matrix form. For each flow data f, the spatial form conversion is first performed, the first 1024 bytes are intercepted, and the 1024 bytes arranged in sequence are converted into a two-dimensional matrix form with 32 bytes in each row and 32 columns in total. If f is less than 1024 bytes, it is padded with 0x00 and the values ​​of all matrix elements are compressed to [0,1].

4. The botnet traffic detection method based on adversarial convolutional autoencoder according to claim 1 is characterized in that: In step 3, a convolutional encoder is first constructed, and its structure is as follows: The first convolution layer: the input channel is 1, the output channel is 16, the convolution kernel size is 3*3, the step size is 2, the padding is 1, and the output feature map size is 16*16*16. ReLU activation function: introduces nonlinearity to the convolution output. The second convolution layer: the input channel is 16, the output channel is 32, the convolution kernel size is 3*3, the step size is 2, the padding is 1, and the output feature map size is 32*8*8. ReLU activation function: nonlinearity is introduced again. Flattening layer: Expand the 32*8*8 three-dimensional feature map into a one-dimensional vector with 2048 features. Fully connected layer: The input feature is 2048, and the output is a low-dimensional latent vector with a dimension of 64. Construct a decoder with the following network structure: Fully connected layer: The input feature is a latent vector dimension of 64, and the input feature is 32*8*8, that is, 2048. The purpose is to restore the latent vector to the shape required for convolution. ReLU activation function: introduces nonlinearity, Unflattening layer: convert the one-dimensional vector 2048 back to a three-dimensional tensor: 32*8*8, The first layer of deconvolution: the input channel is 32, the output channel is 16, the convolution kernel size is 3*3, the step size is 2, the padding is 1, and the output feature map size is 16*16*16. ReLU activation function: introduces nonlinearity, The second layer of deconvolution: the input channel is 16, the output channel is 1, the convolution kernel size is 3*3, the step size is 2, the padding is 1, and the output feature map size is 1*32*32. Sigmoid activation function: compresses the output data to the range of [0,1], which is suitable for the case where the input data is a normalized image. Construct a discriminator, whose network structure is as follows: The first layer is fully connected: the input feature is a hidden vector with a dimension of 64, and the input hidden vector is mapped to a hidden layer with a dimension of 128. ReLU activation function: introduces nonlinearity, The second layer is fully connected: the input feature is 128-dimensional, the output feature is 1, and the hidden layer output is mapped to a single scalar, indicating the discrimination probability. Sigmoid activation function: compresses the output to the range of [0,1], indicating the probability of being judged as the true latent vector, and determines the prior distribution as Gaussian distribution. Assuming that the latent vector conforms to Gaussian distribution, that is, where I is the identity matrix.

5. The botnet traffic detection method based on adversarial convolutional autoencoder according to claim 1 is characterized in that: In step 4, input the data in the form of step 2, and the model training process is as follows: Forward propagation: For each batch of data The encoder E processes the input data x and generates a latent vector z, i.e. z = E(x). The decoder G reconstructs the latent vector z and reconstructs the output data Right now The discriminator D randomly samples a latent vector z from the prior distribution p(z) real , and process two inputs, the true hidden vector z real (from the prior distribution), generates a latent vector z (from the encoder E), and outputs the discriminant result, which is the probability D(z real ), the probability of being generated is D(z), Loss calculation: First use the mean square error to calculate the reconstruction loss, that is, calculate the input data x and the reconstructed data The error is expressed as The discriminator tries to distinguish the true latent vector z real And generate the hidden vector z, the formula for calculating the discriminator adversarial loss is The encoder tries to "fool" the discriminator so that z is identified as the true latent vector. The encoder adversarial loss is calculated as Finally, the overall loss is calculated, i.e. Parameter update: Update the parameters of the discriminator D, that is Minimize the discriminator loss L D , update the parameters of encoder E, that is Update the parameters of decoder G, that is Minimize the reconstruction loss L rec and encoder adversarial loss L E , and finally perform iterative training and repeat the above steps until the network converges.

6. The botnet traffic detection method based on adversarial convolutional autoencoder according to claim 1 is characterized in that: In step 5, the sliding window dynamic threshold is introduced as the classification condition, and the test data X is input during the test. test , including normal traffic and botnet traffic, and calling model detection, which is specifically: First, the initial threshold is determined according to the mean and standard deviation of the normal traffic reconstruction loss in the training phase. The formula is τ init =μ train +k·σ train , where μ train and σ train are the mean and standard deviation of the training data respectively, k is the adjustment coefficient, and then the local statistical calculation of the sliding window is performed. A sliding window with a length of W is maintained to store the reconstruction loss values ​​of the most recent W test samples, and the local mean and standard deviation in the window are calculated in real time, that is, Finally, dynamic threshold adjustment is performed, that is, τ t =β·τ init +(1-β)·(μ window +k·σ window ), Then test the test data x test After the encoder E, the latent vector z is generated test , the latent vector z test After the decoder G, the data is reconstructed Calculate the test data x test and reconstructing data If the reconstruction error between them is greater than the dynamic threshold in the current sliding window, it is determined to be botnet traffic, otherwise it is determined to be normal traffic.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the botnet traffic detection method based on the adversarial convolutional autoencoder as described in any one of claims 1 to 6 above is implemented.

8. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the computer instruction is executed by the processor, the botnet traffic detection method based on the adversarial convolutional autoencoder as described in any one of claims 1 to 6 is implemented.