Network abnormal flow detection method and device

By combining the network anomaly traffic detection model of convolutional neural network, autoencoder and gated loop unit module, the problem of difficulty in detecting multimodal network attacks in the existing technology is solved, and the multi-dimensional processing and timing feature expression of network traffic data are realized, which improves the accuracy and robustness of detection.

CN120455114APending Publication Date: 2025-08-08孙立强
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510682974.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

Existing network anomaly traffic detection methods are difficult to effectively detect multimodal network attacks. The statistical analysis-based methods are susceptible to traffic data fluctuations and noise. The machine learning-based methods are highly dependent on feature selection algorithm design, and the deep learning-based methods are difficult to cover a variety of malicious attack methods.

Method used

The network abnormal traffic detection model consisting of a convolutional neural network module, an autoencoder, a gated cyclic unit module and a Softmax classifier is adopted. Through preprocessing, feature extraction, fusion and classification, the channel attention and spatial attention modules are used for weighting processing, and combined with the feature compression of the autoencoder and the timing feature extraction of the GRU, the multi-dimensional processing and timing feature expression of network traffic data are realized.

Benefits of technology

It improves the accuracy and robustness of network abnormal traffic detection, can effectively detect multiple network attacks, enhances the feature expression ability of network traffic data, and achieves more accurate abnormal traffic detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120455114A_ABST
    Figure CN120455114A_ABST
Patent Text Reader

Abstract

The invention discloses a network abnormal traffic detection method and device. The method comprises the following steps: performing preprocessing based on acquired network traffic data to obtain a processed feature map; inputting the processed feature map into a network abnormal flow detection model to obtain a detection result; the model training process comprises the steps of performing preprocessing based on an acquired network traffic data set to obtain an input feature map; a convolutional neural network module is utilized to perform weighting processing in a channel dimension and a space dimension to obtain a reconstructed feature map; processing by using an auto-encoder to obtain an advanced feature vector; fusing the reconstructed feature map and the advanced feature vector by using a fusion module to obtain a fused vector; performing feature extraction on the fused vectors by using a gating circulation unit module to obtain time sequence features; inputting the time sequence features into a Softmax classifier to obtain a classification result; and repeating the model training process until a preset condition is met to obtain a network abnormal flow detection model. The abnormal traffic in the network can be accurately detected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of network security, and in particular to a method and device for detecting abnormal network traffic. Background Art

[0002] With the development of information technology, information networks have become a core component of social life and work. From infrastructure to social media, from commercial transactions to government services, network applications are ubiquitous. However, the exponential growth in the scale of internet services and the number of users has also brought about significant cybersecurity challenges. Therefore, there is an urgent need to develop an effective detection system to address these issues and combat cyberattacks. Summary of the Invention

[0003] In order to overcome the limitations of existing algorithms and target various malicious network attack methods under multimodal big data, the present invention provides a method and device for detecting abnormal network traffic.

[0004] In a first aspect, an embodiment of the present invention provides a method for detecting abnormal network traffic, comprising:

[0005] Preprocessing is performed based on the acquired network traffic data to obtain a processed feature map;

[0006] Inputting the processed feature graph into a network abnormal traffic detection model to obtain a detection result;

[0007] The training process of the network abnormal traffic detection model includes:

[0008] Preprocess the acquired network traffic dataset to obtain an input feature map;

[0009] Obtaining an initial network abnormal traffic detection model, wherein the initial network abnormal traffic detection model includes a convolutional neural network module, an autoencoder, a gated recurrent unit module, a fusion module, and a Softmax classifier;

[0010] Using the convolutional neural network module to perform weighted processing on the input feature map in the channel dimension and the spatial dimension to obtain a reconstructed feature map;

[0011] Processing the input feature map using the autoencoder to obtain a high-level feature vector;

[0012] Using the fusion module to fuse the reconstructed feature map and the high-level feature vector to obtain a fused vector;

[0013] Using the gated recurrent unit module to extract features from the fused vector to obtain time series features;

[0014] Inputting the time series features into the Softmax classifier to obtain a classification result;

[0015] Repeat the above model training process until the preset conditions are met to obtain a network abnormal traffic detection model.

[0016] Optionally, the preprocessing based on the acquired network traffic data to obtain a processed feature map includes:

[0017] Get network traffic data;

[0018] Based on the data dimension of the network traffic data, the network traffic data is segmented to obtain segmented data;

[0019] Data stream extraction is performed based on the cut data to obtain a data stream vector of a fixed length and a processed feature map.

[0020] Optionally, the convolutional neural network module includes a channel attention module and a spatial attention module;

[0021] The method of using the convolutional neural network module to perform weighted processing on the input feature map in the channel dimension and the spatial dimension to obtain a reconstructed feature map includes:

[0022] Inputting the input feature map into the convolutional neural network module;

[0023] Performing weighted processing on the input feature map in the channel dimension and the spatial dimension by the channel attention module and the spatial attention module to obtain a spatial attention feature map;

[0024] The channel attention module and the spatial attention module are used to repeatedly perform weighted processing on the spatial attention feature map in the channel dimension and the spatial dimension to obtain the reconstructed feature map.

[0025] Optionally, performing weighted processing on the input feature map in the channel dimension and the spatial dimension by the channel attention module and the spatial attention module to obtain the spatial attention feature map includes:

[0026] The channel attention module is used to calculate the weight of each feature channel of the input feature map in the channel dimension, and a channel weighted feature map is obtained based on the input feature map and the weights of each feature channel;

[0027] Using the spatial attention module to perform global maximum pooling and global average pooling on the channel weighted feature map in the spatial dimension, respectively, to obtain two spatial feature maps;

[0028] Stacking the two spatial feature maps to obtain a joint feature map;

[0029] Convolution processing and Sigmoid activation operations are performed on the joint feature map in sequence to obtain a spatial attention feature map.

[0030] Optionally, the channel attention module includes a multi-layer perceptron;

[0031] The method of calculating the weights of each feature channel of the input feature map in the channel dimension by using the channel attention module, and obtaining a channel weighted feature map based on the input feature map and the weights of each feature channel, includes:

[0032] Using the channel attention module to perform global average pooling and global maximum pooling on the input feature map in the channel dimension to obtain two channel description vectors;

[0033] Inputting the two channel description vectors into the multi-layer perceptron for processing to obtain a processed result;

[0034] After adding the two processed results, a Sigmoid activation operation is performed to obtain a weight matrix for each feature channel;

[0035] Calculation is performed based on the weight matrix of each feature channel and the input feature map to obtain a channel weighted feature map.

[0036] Optionally, the autoencoder includes an encoder and a decoder;

[0037] The processing of the input feature map by the autoencoder to obtain a high-level feature vector includes:

[0038] Compressing the input feature map using the encoder to obtain compressed features;

[0039] The compressed features are decoded using the decoder to obtain decoded features as high-level feature vectors.

[0040] In a second aspect, an embodiment of the present invention provides a network abnormal traffic detection device, comprising:

[0041] A data processing module is used to perform preprocessing based on the acquired network traffic data to obtain a processed feature map;

[0042] A detection module is configured to input the processed feature graph into a network abnormal traffic detection model to obtain a detection result; wherein the training process of the network abnormal traffic detection model includes:

[0043] Preprocess the acquired network traffic dataset to obtain an input feature map;

[0044] Obtaining an initial network abnormal traffic detection model, wherein the initial network abnormal traffic detection model includes a convolutional neural network module, an autoencoder, a gated recurrent unit module, a fusion module, and a Softmax classifier;

[0045] Using the convolutional neural network module to perform weighted processing on the input feature map in the channel dimension and the spatial dimension to obtain a reconstructed feature map;

[0046] Processing the input feature map using the autoencoder to obtain a high-level feature vector;

[0047] Using the fusion module to fuse the reconstructed feature map and the high-level feature vector to obtain a fused vector;

[0048] Using the gated recurrent unit module to extract features from the fused vector to obtain time series features;

[0049] Inputting the time series features into the Softmax classifier to obtain a classification result;

[0050] Repeat the above model training process until the preset conditions are met to obtain a network abnormal traffic detection model.

[0051] In a third aspect, an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the network abnormal traffic detection method as described in the first aspect.

[0052] In a fourth aspect, an embodiment of the present invention provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method for detecting abnormal network traffic as described in the first aspect is implemented.

[0053] In a fifth aspect, an embodiment of the present invention provides a computer program product comprising instructions, which, when executed on a computer device, enables the computer device to execute the network abnormal traffic detection method as described in the first aspect.

[0054] The beneficial effects of the above technical solutions provided in the embodiments of the present invention include at least:

[0055] An embodiment of the present invention provides a method for detecting abnormal network traffic. After preprocessing network traffic data, a processed feature graph is obtained. The processed feature graph is input into a network abnormal traffic detection model to obtain a detection result. When the network abnormal traffic detection model processes the network traffic data, a convolutional neural network module is used to adjust the weights of the processed feature graph, thereby highlighting key information in the data. At the same time, an autoencoder is used to compress the features of the processed feature graph in different dimensions to obtain hidden high-level feature information in the processed feature graph. This allows deep mining, multi-dimensional processing, and effective fusion of the features of the network traffic data. The fused feature vector can more completely express the information content contained in the network traffic data. A gated recurrent unit module is used to extract features from the fused vector to obtain time series features. A Softmax classifier is used based on the time series features to obtain a classification result. This method not only mines hidden information features within the data to achieve accurate abnormal traffic detection, but also makes good use of the time series characteristics of the network traffic data, enabling the model to have temporal and spatial feature expression capabilities, thereby making the final detection result more accurate, providing a more effective method for network security protection. Compared with previous abnormal traffic detection solutions, the network abnormal traffic detection method provided by the embodiment of the present invention can handle various network attacks and malicious traffic types.

[0056] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written description and the accompanying drawings.

[0057] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0059] Figure 1 This is a flow chart of a method for detecting abnormal network traffic provided in an embodiment of the present invention;

[0060] Figure 2 A schematic diagram of the training process of the network abnormal traffic detection model provided in an embodiment of the present invention;

[0061] Figure 3 This is a schematic diagram of the structure of the channel attention module provided in an embodiment of the present invention;

[0062] Figure 4This is a schematic diagram of the structure of the spatial attention module provided in an embodiment of the present invention;

[0063] Figure 5 Schematic diagram of the structure of the autoencoder provided in an embodiment of the present invention;

[0064] Figure 6 A schematic diagram of the GRU network structure provided in an embodiment of the present invention;

[0065] Figure 7 This is a schematic diagram of the structure of a convolutional neural network module in the network abnormal traffic detection model provided in an embodiment of the present invention;

[0066] Figure 8 This is a flowchart of the overall framework of the training process of the network abnormal traffic detection model provided in an embodiment of the present invention;

[0067] Figure 9 This is a schematic diagram of the structure of a network abnormal traffic detection device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0068] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.

[0069] In the description of the present invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," "outer," "far," "near," "front," and "rear" and the like, indicating positions or relationships, are based on the positions or relationships shown in the accompanying drawings and are intended solely to facilitate and simplify the description of the present invention. They are not intended to indicate or imply that the devices or components referred to must have, be constructed, or operate in a specific orientation, and therefore should not be construed as limiting the present invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0070] In the description of the present invention, it should be noted that, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood in a broad sense. For example, they may refer to fixed, detachable, or integral connections; mechanical or electrical connections; direct or indirect connections through an intermediate medium; and internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on the specific circumstances.

[0071] The inventors have discovered that different network attacks are accompanied by large amounts of abnormal traffic, and abnormal traffic detection can effectively identify potential cybercrime. In order to cope with increasingly complex malicious attacks and changes in traffic patterns, researchers at home and abroad have proposed a variety of abnormal traffic detection methods, which are mainly divided into early abnormal traffic detection based on statistical analysis, abnormal traffic detection based on machine learning, and abnormal traffic detection based on deep learning. Among them, the abnormal traffic detection method based on statistical analysis is easily affected by fluctuations and noise in traffic data, resulting in a high false alarm rate, which limits the effectiveness of practical applications. The abnormal traffic detection method based on machine learning extracts feature information from network traffic data and uses machine learning classification algorithms to achieve abnormal traffic detection, but it is highly dependent on the design of the feature selection algorithm and the architecture and parameter selection of the model.

[0072] Currently, deep learning-based anomaly traffic detection is widely used. It uses deep learning algorithms, such as autoencoders or convolutional neural networks, to automatically learn and extract features. However, network traffic data is highly variable, and new malicious network attacks are becoming increasingly common. A single deep learning algorithm struggles to cover multimodal data and detect diverse malicious attack methods.

[0073] In order to solve the above problems, the inventors have proposed a method and device for detecting abnormal network traffic through research and development. The goal is to be able to process multimodal network data, detect potential network threats and attack behaviors in a timely manner, and maintain network security.

[0074] First, some nouns or terms that appear in the description of the embodiments of this application are subject to the following interpretations:

[0075] PCAP (Packet Capture) is a standard file format for storing network traffic data. PCAP files record raw network data packets captured by network devices (such as network cards), including complete communication content (such as protocol headers and payload data). Each packet can contain the following information: timestamp (the exact time the packet was captured), packet length (original length and actual captured length), protocol header information (such as Ethernet frame header, IP header, TCP / UDP header, etc.), and payload data (the actual content transmitted, such as HTTP requests, DNS queries, encrypted traffic, etc.).

[0076] The Attention Mechanism is a computational technique that simulates the human cognitive process of selectively focusing on important information and is widely used in the field of deep learning. Its core idea is to dynamically assign weights so that the model can focus on the most relevant parts of the input when processing it, rather than treating all information equally.

[0077] An autoencoder is an unsupervised learning model used to learn effective representations (encodings) of data. Its core idea is to compress input data into a low-dimensional representation using an encoder, and then reconstruct the original data using a decoder, thereby learning the essential characteristics of the data.

[0078] The Softmax classifier is a commonly used classification method in multi-classification problems. It is based on the Softmax function, which can map a vector to a probability distribution. The Softmax classifier is usually used in the last layer of a neural network to output the probability of each category.

[0079] The Gated Recurrent Unit (GRU) module is an improved recurrent neural network (RNN). By reducing the number of gates and simplifying the structure, it retains the long-term dependency capture capability of LSTM while improving computational efficiency. GRU controls the flow of information by introducing two gates (reset gate and update gate).

[0080] A flow is a collection of data packets with the same five-tuple (source IP, destination IP, source port, destination port, protocol).

[0081] Example 1

[0082] See Figure 1 The present embodiment proposes a method for detecting abnormal network traffic, which may include the following steps:

[0083] Step S101: pre-processing is performed based on the acquired network traffic data to obtain a processed feature map;

[0084] In the above step S101, obtaining the processed feature map specifically includes the following steps:

[0085] Step S1011: Acquire network traffic data.

[0086] In the above step S1011, a PCAP file is obtained. The PCAP file records a series of original data packets captured by a network device (such as a network card), from which complete network traffic data can be extracted. Assuming that the network traffic data has n data packets, the network traffic data can be expressed as P = {p1, p2, ..., p n Each data packet includes: source IP, destination IP, source port number, destination port number, network information used by the transport layer, data packet size s, and start time.

[0087] Step S1012: based on the data dimension of the network traffic data, the network traffic data is segmented to obtain segmented data.

[0088] In step S1012, data dimensions refer to the characteristics or attributes of data in a certain aspect, representing multiple angles or aspects of observing and describing the data. For example, a PCAP file contains 100 data packets, each of which has characteristics such as source IP address, destination IP address, protocol type, and packet size. These characteristics are then considered the dimensions of the PCAP file.

[0089] In this embodiment, SplitCap is a tool for analyzing and segmenting PCAP files. It is particularly suitable for extracting traffic of specific protocols or sessions from large packet capture files. The SplitCap tool can be used to segment network traffic data according to data dimensions for data screening. Since some fields in network traffic data do not contain content information, such as the MAC source address, destination address, protocol version, etc. of the Ethernet layer, and most packets have less than 10 and different payload sizes, the SplitCap tool is used to segment the data to ensure that the data input into the model has the same dimensions.

[0090] Step S1013: Perform data stream extraction based on the segmented data to obtain a data stream vector of a fixed length and obtain a processed feature map.

[0091] In the above step S1013, since the input end of the convolutional neural network module in the network abnormal traffic detection model usually requires fixed-dimensional input, by extracting the cut data, the variable-length traffic can be unified into a fixed length, avoiding model errors or performance degradation caused by size mismatch.

[0092] In this embodiment, by adding segmentation and extraction in the preprocessing stage to obtain the processed feature map, the quality and consistency of the processed feature map can be significantly improved, providing a cleaner and more regular input for the subsequent network abnormal traffic detection model, and ultimately improving the accuracy of anomaly detection and system robustness.

[0093] Step S102: Input the processed feature graph into the network abnormal traffic detection model to obtain the detection result.

[0094] See Figure 2 The training process of the above-mentioned network abnormal traffic detection model may specifically include the following steps:

[0095] Step S201: preprocessing is performed based on the acquired network traffic data set to obtain an input feature map;

[0096] Step S202: Obtain an initial network abnormal traffic detection model, where the initial network abnormal traffic detection model includes a convolutional neural network module, an autoencoder, a gated recurrent unit module, a fusion module, and a Softmax classifier;

[0097] Step S203: Using a convolutional neural network module, weighted processing is performed on the input feature map in the channel dimension and the spatial dimension to obtain a reconstructed feature map;

[0098] Step S204: Process the input feature map using an autoencoder to obtain a high-level feature vector;

[0099] Step S205: using a fusion module to fuse the reconstructed feature map and the high-level feature vector to obtain a fused vector;

[0100] Step S206: Use the gated recurrent unit module to extract features from the fused vector to obtain time series features;

[0101] Step S207: Input the time series features into the Softmax classifier to obtain the classification result;

[0102] Step S208: Repeat the above model training process until the preset conditions are met to obtain a network abnormal traffic detection model.

[0103] In order to explain the training process of the above network abnormal traffic detection model more clearly, each step is described in detail below.

[0104] In the above step S201 , the process of preprocessing the obtained network traffic data set can refer to the above steps S1011 to S1013 , which will not be repeated here.

[0105] Network traffic datasets can be sourced from publicly annotated datasets commonly used in the field of network security, such as KDD Cup 1999, one of the earliest classic datasets used for network intrusion detection, which includes normal traffic and various attack types (such as DoS, Probe, and R2L); CICIDS-2017, a dataset released by the Canadian Institute for Cybersecurity for network security assessment, which includes the latest attack types; and TON_IoT Dataset, a dataset designed for IoT environments that includes network traffic, operating system logs, and appliance telemetry data, making it suitable for anomaly detection in IoT scenarios. The network traffic dataset can be divided into a training set and a test set. The training set is used for model training, i.e., executing the process from steps S203 to S208 below. The test set is then used to validate the network anomaly traffic detection model.

[0106] In the above step S203, the convolutional neural network module includes a channel attention module and a spatial attention module. The specific process of obtaining the reconstructed feature map may include the following steps:

[0107] Step S2031: input the input feature map into the convolutional neural network module;

[0108] Step S2032: Perform weighted processing on the input feature map in the channel dimension and the spatial dimension through the channel attention module and the spatial attention module to obtain a spatial attention feature map;

[0109] In the above step S2032, the specific process of obtaining the spatial attention feature map may include the following steps:

[0110] Step S20321: Use the channel attention module to calculate the weights of each feature channel of the input feature map in the channel dimension, and obtain a channel weighted feature map based on the input feature map and the weights of each feature channel.

[0111] In step S20321, the input feature map is first weighted by the channel attention module to learn the dependencies between the channels of the entire input tensor. The data is weighted according to the dependencies so that each channel can be weighted in subsequent processing to process important channel information. The channel attention module uses global pooling and fully connected layers to calculate weights, and performs a dot product between the weights and the input feature map to obtain the weighted result. The channel attention module includes a multi-layer perceptron, see Figure 3 ,The specific process of obtaining the channel weighted feature map using the channel attention module can include the following steps:

[0112] Step S203211: Use the channel attention module to perform global average pooling and global maximum pooling on the input feature map in the channel dimension to obtain two channel description vectors.

[0113] Step S203212: Input the two channel description vectors into a multi-layer perceptron for processing to obtain a processed result.

[0114] In the above steps 203212, the multi-layer perceptron (MLP) (i.e. Figure 3 The shared multi-layer perceptron in is a fully connected network that can process two channel description vectors, including: passing the two channel description vectors through the hidden layer and the output layer to generate a prediction value, that is, the processed result.

[0115] Step S203213: After adding the two processed results, perform a Sigmoid activation operation to obtain a weight matrix for each feature channel.

[0116] Step S203214: Calculate based on the weight matrix of each feature channel and the input feature map to obtain a channel weighted feature map.

[0117] In the above step S203214, the calculation based on the channel weight matrix and the input feature map includes: performing a dot product between the channel weight matrix and the input feature map to obtain a channel weighted feature map, that is, Figure 3Channel attention in .

[0118] In this embodiment, the weighted calculation process of step S20321 can be expressed using the following formula (1):

[0119] M(F)=σ(MLP(GloAvgp(F))+MLP(GloMaxp(F))) (1)

[0120] In the above formula (1), F represents the input feature map; σ represents the Sigmoid activation operation, which constrains the channel weights in the interval [0, 1] through the Sigmoid function to achieve adaptive feature enhancement or suppression; GloAvgp represents the global average pooling operation on the input feature map F; GloMaxp represents the global maximum pooling operation on the input feature map F; MLP represents the multi-layer perceptron.

[0121] The channel attention module is used to perform global average pooling (GAP) on the input feature map to retain the global information of the channel dimension (such as the overall distribution of traffic) and avoid information omission. By performing global maximum pooling (GMP) on the input feature map, local significant features of the channel dimension (such as burst peaks of traffic) are obtained to highlight abnormal sensitive signals. Then, double pooling fusion (i.e., steps S203211 to S203213 mentioned above) is performed. By superimposing the results of GAP and GMP, both overall information and key local features are focused on, so that the channel weight matrix can more efficiently and accurately reflect the feature importance of the input data. In the process of obtaining the channel weight matrix, the multi-layer perceptron (MLP) nonlinear mapping is used to learn the complex relationship between channels, avoiding the limitations of linear weighting.

[0122] Step S20322, refer to Figure 4 , using the spatial attention module to weight the channel feature map (i.e. Figure 4 The input feature map A' in is subjected to global maximum pooling and global average pooling respectively to obtain two spatial feature maps.

[0123] Step S20323: stack the two spatial feature maps to obtain a joint feature map.

[0124] Step S20324: perform convolution processing and Sigmoid activation operations on the joint feature map in sequence to obtain a spatial attention feature map.

[0125] In the above step S20324, the convolution kernel is used to convolve the joint feature map to improve the receptive field and reduce the number of channels of the feature map. Finally, the extracted feature map is subjected to the Sigmoid activation function to obtain the reconstructed feature map (i.e. Figure 4 The computational process of reconstructing the feature map can be expressed as:

[0126] M(F * )=σ(f 7×7 ([GloAvgp(F * ); GloMaxp(F * )])) (2)

[0127] In the above formula (2), F * Represents the channel weighted feature map; σ represents the sigmoid activation operation; GloAvgp represents the channel weighted feature map F * Perform global average pooling operation, GloMaxp represents the channel weighted feature map F * Perform global maximum pooling operation, f 7×7 Represents a large convolution kernel of 7x7.

[0128] Step S2033: Use the channel attention module and the spatial attention module to repeatedly perform weighted processing on the spatial attention feature map in the channel dimension and the spatial dimension to obtain a reconstructed feature map.

[0129] In the above step S2033, based on the spatial attention feature map, the channel attention module and the spatial attention module are used to perform weighted processing again. The weighted processing method can refer to the above step S2032 and will not be repeated here. Of course, before executing step S2033, the spatial attention feature map can be input into the convolution layer and the maximum pooling layer in sequence for feature extraction, and then step S2033 is executed. It can retain the more significant and potentially more discriminative feature information in the feature map, so that the subsequent detection model can pay more attention to these key features, which helps to improve the accuracy of detection.

[0130] In this embodiment, through the cascade design of the channel attention module and the spatial attention module, a refined dynamic weight allocation of the network traffic feature map is achieved, which can more effectively extract the key features in the network traffic data. It has the advantages of noise resistance, adaptability and multi-dimensional focusing, and significantly improves the accuracy and robustness of anomaly detection. It is especially suitable for complex, high-dimensional and dynamically changing network environments.

[0131] In the above step S204, refer to Figure 5 , the autoencoder includes an encoder and a decoder, whose purpose is to transform the input feature map (i.e. Figure 5 The input data d) is encoded into a short representation (i.e. Figure 5 The compressed feature z in the data is extracted, and the main features in the data are then decoded by the decoder to obtain a data vector close to the input feature map (i.e. Figure 5The input data d' in the input feature map is used to mine the deep hidden information of the input feature map. Both the encoder and decoder are two-layer fully connected networks. The specific process of processing the input feature map using the pre-built autoencoder to obtain the high-level feature vector can include the following steps:

[0132] Step S2041: Use the encoder to compress the input feature map to obtain compressed features.

[0133] In the above step S2041, the encoder is set with encoder parameters, including the weight matrix in the encoder and the bias parameters in the encoder. The encoder parameters are determined during the training process of the entire network abnormal traffic detection model. Figure 5 , using the encoder to input feature maps ( Figure 5 The input data d) in the image is compressed to obtain the compressed features. The process can be expressed as:

[0134] z=f en (F)=σ(μ en F+b en ) (3)

[0135] In the above formula (3), f en Indicates the encoder encoding operation on the input feature map F; μ en represents the weight matrix in the encoder; b en represents the bias parameter in the encoder; σ represents the ReLU activation function.

[0136] Step S2042: Decode the compressed features using a decoder to obtain decoded features as high-level feature vectors.

[0137] In step S2042, the decoder is configured with decoder parameters, including a weight matrix and a bias parameter in the decoder. The decoder parameters are determined during the training of the entire network abnormal traffic detection model. The process of decoding the compressed features using the decoder can be expressed as:

[0138] d'=f de (z)=σ de (μ de z+b de ) (4)

[0139] In the above formula (4), f de Indicates the decoder decoding operation of the compressed feature z, μ de represents the weight matrix in the encoder, b de represents the bias parameter in the encoder; σ de Represents the activation function.

[0140] In the above step S205, the specific process of obtaining the fused vector may include the following steps:

[0141] The fusion module is used to concatenate the reconstructed feature map and the high-level feature vector, and the two feature vectors are expanded and combined in dimension to form a fused vector. The fused vector contains both the feature information after dynamic weight allocation by the convolutional neural network module and the hidden feature information mined by the autoencoder, providing more comprehensive feature input for subsequent time series feature extraction.

[0142] Specifically, because the reconstructed feature map is a weighted feature map, it highlights key information. The high-level feature vector represents hidden information mined by the autoencoder, which is not accessible by other modules. The fused vector created by concatenating the reconstructed feature map and the high-level feature vector can more fully represent the information contained in network traffic data.

[0143] In the above step S206, since the packet transmission of business data has a time sequence and the traffic packet data sent in a certain time and space is dynamically changing, the time characteristics of the data are emphasized. RNN (recurrent neural network) has shown good performance in processing time series. The gated recurrent unit module (GRU) used in this embodiment, as an optimized variant of the classic recurrent neural network, simplifies the complex gate structure in the previous network. GRU only contains two gates, update gate and reset gate, with fewer parameters, faster convergence, and superior performance.

[0144] See Figure 6 , is the network structure of GRU, assuming x t Represents the GRU input data, h t is the output value at time t; is the candidate memory, representing the state at time t. Reset gate r t Responsible for short-term memory, the reset gate weight is calculated through the Sigmoid activation function (σ), which means that the state of the previous moment is written into the current moment. t and candidate memory It can be expressed as:

[0145]

[0146] In the above formula (5), W r Represents the weight matrix of the reset gate; h t-1 represents the output value at time t-1; x t Represents the input data at time t; σ represents the Sigmoid activation function; Represents the weight matrix of the candidate memory, combined with the current input x tand partially selected historical information r t ·h t-1 , generating hidden state features.

[0147] Update gate z t Responsible for long-term memory, used to decide how much historical information to retain and add new information to the hidden state at the current moment. Its value is related to the output information of the previous time node and the input information of the current time node:

[0148]

[0149] In the above formula (6), W z represents the weight matrix of the update gate; h t Represents the output value at time t.

[0150] In step S207 above, the time series features are input into the Softmax classifier to obtain the classification result. The time series features are obtained by GRU, with a dimension of 2. The Softmax classifier is used to obtain a vector of the same dimension length, whose value is between 0 and 1, indicating the probability of the classification result being normal or abnormal. The formula is as follows:

[0151]

[0152] In the above formula (7), z i represents the i-th element in the time series feature; z d Represents all elements in the time series feature.

[0153] In step S208, the input feature graph includes data samples and corresponding labels, where the labels include normal network traffic and abnormal network traffic. Preset conditions include the loss function value between the classification result and the corresponding label being lower than a preset threshold or the number of training times reaching a preset upper limit. The model training process is repeated until the preset conditions are met. The specific process of obtaining the abnormal network traffic detection model may include the following steps:

[0154] S2081. Calculate the loss function value between the classification result and the corresponding label based on a preset loss function according to the classification result and the corresponding label;

[0155] In the above step S2081, the loss function value can be calculated using the following formula (7):

[0156]

[0157] In the above formula (7), d i Indicates the label corresponding to the i-th data sample; d' i represents the classification result; m represents the number of samples in the network traffic dataset.

[0158] S2082. Update the model parameters through the back propagation algorithm and repeat the iterative training process until the loss function value is lower than the preset threshold or the number of training times reaches the preset upper limit, thereby obtaining a network abnormal traffic detection model.

[0159] For example, see Figure 7 and Figure 8 , a training process of the network abnormal traffic detection model can be: obtain the PCAP file (i.e. Figure 7 PCAP packets in , from which a complete network traffic dataset can be extracted (i.e. Figure 7 The network traffic data set is preprocessed, including segmentation and data flow extraction. Only the first 10 data packets are intercepted, and each data packet is 160 bytes. For flow files with less than 10 data packets, they are padded with 0s to obtain a 40*40 grayscale image, which is the segmented data. The segmented data is passed through a convolution layer containing 32 convolution kernels and a maximum pooling layer for data flow extraction to obtain a fixed-length data flow vector and an input feature map. The input feature map is input into the convolutional neural network module for processing to obtain a reconstructed feature map. Specifically: the input feature map is input into the attention module for weighting. The input feature map is weighted in the channel dimension and spatial dimension by the channel attention module and the spatial attention module to obtain a spatial attention feature map. The spatial attention feature map is passed through a convolution layer containing 64 convolution kernels and a maximum pooling layer for feature extraction. The channel attention module and the spatial attention module are used again for attention weighting to obtain a reconstructed feature map. The reconstructed feature map is flattened and output through a fully connected layer containing 800 neurons to ensure that the feature dimension is 16000. At the same time, to prevent overfitting, the Dropout operation can be performed to randomly inactivate some neurons.

[0160] Then, an autoencoder is used to mine the hidden data information and obtain a high-level feature vector. The encoder and decoder used are both two-layer fully connected networks. The input 1600-dimensional vector is compressed to 1200 dimensions through the first encoding layer and then to 800 dimensions through the second encoding layer. It is then decoded to 1200 dimensions through the first decoding layer and to 1600 dimensions through the second decoding layer.

[0161] Finally, the reconstructed feature map and the high-level feature vector are fused to generate a fused vector. This fused vector is then fed into the gated recurrent unit module to extract temporal features from the fused vector. The temporal features are then fed into the Softmax classifier to generate a vector with values between 0 and 1, representing the probability of normal and abnormal classifications, thus generating the classification result.

[0162] During the training process of the network anomaly traffic detection model, convolutional neural network modules, autoencoders, gated recurrent unit modules, and softmax classifiers are organically combined to form an integrated whole. This system leverages the strengths of each component to solve the problem of network anomaly traffic detection and collaborates to improve detection effectiveness. The convolutional neural network module adjusts the weights of the input feature map to produce a reconstructed feature map, dynamically assigning weights to the channel and spatial dimensions of the input feature map. The autoencoder compresses the features of the input feature map along different dimensions, further extracting hidden high-level feature information from the input feature map. This allows for deep mining, multi-dimensional processing, and effective fusion of network traffic data features. The gated recurrent unit module extracts features from the fused vector to produce temporal features, fully accounting for the temporal and complexity of network traffic data. The resulting trained network anomaly traffic detection model can more accurately detect abnormal traffic in the network, providing a more effective approach for network security protection.

[0163] Therefore, in the process of detecting network traffic data using the network abnormal traffic detection model of this embodiment, the convolutional neural network module is used to dynamically assign weights to the channel dimension and spatial dimension of the processed feature map, which can pay more attention to the important information part in the data, highlight the key features and reduce the interference of irrelevant or minor features, thereby enhancing the expressiveness and discrimination of the input features; at the same time, the autoencoder compresses the features of the processed feature map in different dimensions, further extracts the abstract features in the data, captures deeper data structure information, and effectively mines the hidden high-level feature information in the data, thereby deeply mining, multi-dimensionally processing and effectively fusing the features of the network traffic data. The fused feature vector can more completely express the information content contained in the network traffic data, thereby making the final detection result more accurate. Compared with previous abnormal traffic detection schemes, the network abnormal traffic detection method provided by this embodiment can handle a variety of different types of network attacks and malicious traffic types.

[0164] Traditional anomaly detection based on autoencoders usually simply uses the autoencoder to encode and decode the input data, and determines whether the data is normal by calculating the reconstruction error. However, the method of this embodiment not only uses the autoencoder to extract high-level feature vectors, but also combines the convolutional neural network module to dynamically assign weights to the input feature map to obtain a reconstructed feature map, and fuses it with the high-level feature vector, and then performs time series feature extraction and classification. It can more comprehensively and deeply mine the feature information in the data, thereby improving the accuracy and robustness of detection.

[0165] Compared with the method based only on attention mechanism and LSTM / GRU: Although some existing methods also add attention mechanism on the basis of recurrent neural networks such as LSTM or GRU to focus on important time steps or features, in this embodiment, since the autoencoder and gated recurrent unit modules are combined in the training process of the network abnormal traffic detection model, the features are first processed and fused in multiple dimensions through the autoencoder and attention mechanism, and then the GRU is used to capture the time series features. Such an architecture is more comprehensive in feature extraction and modeling. The network abnormal traffic detection model finally trained can better cope with complex network traffic data and various abnormal situations.

[0166] Compared with traditional anomaly detection methods based on statistics or machine learning: Traditional machine learning methods such as Naive Bayes and Random Forest have limited capabilities in expressing complex functions and processing large-scale data. The method of this embodiment, with the help of various advanced technologies in deep learning, can automatically learn complex patterns and features in the data, effectively solving the problems of low computational efficiency and poor generalization ability faced by traditional methods when facing large-scale, high-dimensional network traffic data, and can handle various different types of network attacks and malicious traffic types.

[0167] The method of this embodiment requires less training and trains models quickly. Compared to previous models that often use RNNs or LSTMs, the training process of this embodiment uses a more concise GRU module, which converges faster. For practical application scenarios with diverse datasets and frequently updated datasets, the method of this embodiment is more universal, reliable, and robust.

[0168] Example 2

[0169] Based on the same inventive concept, see Figure 9 , the embodiment of the present application also proposes a network abnormal traffic detection device, including:

[0170] Data processing module 1, used for preprocessing the acquired network traffic data to obtain a processed feature map;

[0171] Detection module 2 is used to input the processed feature map into the network abnormal traffic detection model to obtain the detection result. The training process of the network abnormal traffic detection model includes:

[0172] Preprocess the acquired network traffic dataset to obtain an input feature map;

[0173] Obtain an initial network abnormal traffic detection model, which includes a convolutional neural network module, an autoencoder, a gated recurrent unit module, a fusion module, and a Softmax classifier;

[0174] Use the convolutional neural network module to perform weighted processing on the input feature map in the channel dimension and spatial dimension to obtain the reconstructed feature map;

[0175] Use the autoencoder to process the input feature map to obtain a high-level feature vector;

[0176] The reconstructed feature map and the high-level feature vector are fused using the fusion module to obtain the fused vector;

[0177] The gated recurrent unit module is used to extract features from the fused vector to obtain temporal features;

[0178] Input the time series features into the Softmax classifier to obtain the classification results;

[0179] Repeat the above model training process until the preset conditions are met to obtain a network abnormal traffic detection model.

[0180] The implementation principle and technical effects of the network abnormal traffic detection device provided in the embodiment of the present invention are similar to those of the first embodiment and will not be repeated here.

[0181] Example 3

[0182] Based on the same inventive concept, an embodiment of the present application further proposes a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the network abnormal traffic detection method as in embodiment 1 is implemented.

[0183] The computer-readable storage medium may be included in the device / apparatus described in the above embodiments, or may exist independently and not be incorporated into the device / apparatus. The computer-readable storage medium carries one or more programs, which, when executed, implement the method according to the first embodiment of the present invention.

[0184] According to an embodiment of the present invention, a computer-readable storage medium may be a non-volatile computer-readable storage medium, such as, but not limited to, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0185] Example 4

[0186] Based on the same inventive concept, an embodiment of the present application also proposes a computer device, including a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the program, the network abnormal traffic detection method as in embodiment one is implemented.

[0187] Example 5

[0188] Based on the same inventive concept, an embodiment of the present application proposes a computer program product containing instructions. When the computer program product runs on a computer device, the computer device executes the network abnormal traffic detection method in embodiment one.

[0189] The principles of solving the problems described above by the apparatus, client, medium, and related equipment in the embodiments of the present invention are similar to those of the aforementioned methods, so their implementation can refer to the implementation of the aforementioned methods, and repeated details will not be repeated.

[0190] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage and optical storage, etc.) containing computer-usable program code.

[0191] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0192] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0193] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0194] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. The present disclosure is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and variations may be made without departing from the scope of the present disclosure. The scope of the present disclosure is limited solely by the appended claims. Thus, to the extent such modifications and variations fall within the scope of the claims and their equivalents, the present disclosure is intended to include such modifications and variations.

Claims

1. A method for detecting abnormal network traffic, characterized in that: include: Preprocessing is performed based on the acquired network traffic data to obtain a processed feature map; Inputting the processed feature graph into a network abnormal traffic detection model to obtain a detection result; The training process of the network abnormal traffic detection model includes: Preprocess the acquired network traffic dataset to obtain an input feature map; Obtaining an initial network abnormal traffic detection model, wherein the initial network abnormal traffic detection model includes a convolutional neural network module, an autoencoder, a gated recurrent unit module, a fusion module, and a Softmax classifier; Using the convolutional neural network module to perform weighted processing on the input feature map in the channel dimension and the spatial dimension to obtain a reconstructed feature map; Processing the input feature map using the autoencoder to obtain a high-level feature vector; Using the fusion module to fuse the reconstructed feature map and the high-level feature vector to obtain a fused vector; Using the gated recurrent unit module to extract features from the fused vector to obtain time series features; Inputting the time series features into the Softmax classifier to obtain a classification result; Repeat the above model training process until the preset conditions are met to obtain a network abnormal traffic detection model.

2. The method for detecting abnormal network traffic according to claim 1, wherein: The preprocessing is performed based on the acquired network traffic data to obtain a processed feature map, including: Get network traffic data; Based on the data dimension of the network traffic data, the network traffic data is segmented to obtain segmented data; Data stream extraction is performed based on the cut data to obtain a data stream vector of a fixed length and a processed feature map.

3. The method for detecting abnormal network traffic according to claim 1, wherein: The convolutional neural network module includes a channel attention module and a spatial attention module; The method of using the convolutional neural network module to perform weighted processing on the input feature map in the channel dimension and the spatial dimension to obtain a reconstructed feature map includes: Inputting the input feature map into the convolutional neural network module; Performing weighted processing on the input feature map in the channel dimension and the spatial dimension by the channel attention module and the spatial attention module to obtain a spatial attention feature map; The channel attention module and the spatial attention module are used to repeatedly perform weighted processing on the spatial attention feature map in the channel dimension and the spatial dimension to obtain the reconstructed feature map.

4. The method for detecting abnormal network traffic according to claim 3, wherein: The step of performing weighted processing on the input feature map in the channel dimension and the spatial dimension by the channel attention module and the spatial attention module to obtain a spatial attention feature map includes: The channel attention module is used to calculate the weight of each feature channel of the input feature map in the channel dimension, and a channel weighted feature map is obtained based on the input feature map and the weights of each feature channel; Using the spatial attention module to perform global maximum pooling and global average pooling on the channel weighted feature map in the spatial dimension, respectively, to obtain two spatial feature maps; Stacking the two spatial feature maps to obtain a joint feature map; Convolution processing and Sigmoid activation operations are performed on the joint feature map in sequence to obtain a spatial attention feature map.

5. The method for detecting abnormal network traffic according to claim 4, wherein: The channel attention module includes a multi-layer perceptron; The method of calculating the weights of each feature channel of the input feature map in the channel dimension by using the channel attention module, and obtaining a channel weighted feature map based on the input feature map and the weights of each feature channel, includes: Using the channel attention module to perform global average pooling and global maximum pooling on the input feature map in the channel dimension to obtain two channel description vectors; Inputting the two channel description vectors into the multi-layer perceptron for processing to obtain a processed result; After adding the two processed results, a Sigmoid activation operation is performed to obtain a weight matrix for each feature channel; Calculation is performed based on the weight matrix of each feature channel and the input feature map to obtain a channel weighted feature map.

6. The method for detecting abnormal network traffic according to claim 1, wherein: The autoencoder includes an encoder and a decoder; The processing of the input feature map by the autoencoder to obtain a high-level feature vector includes: Compressing the input feature map using the encoder to obtain compressed features; The compressed features are decoded using the decoder to obtain decoded features as high-level feature vectors.

7. A network abnormal traffic detection device, characterized in that: include: A data processing module is used to perform preprocessing based on the acquired network traffic data to obtain a processed feature map; A detection module is configured to input the processed feature graph into a network abnormal traffic detection model to obtain a detection result; wherein the training process of the network abnormal traffic detection model includes: Preprocess the acquired network traffic dataset to obtain an input feature map; Obtaining an initial network abnormal traffic detection model, wherein the initial network abnormal traffic detection model includes a convolutional neural network module, an autoencoder, a gated recurrent unit module, a fusion module, and a Softmax classifier; Using the convolutional neural network module to perform weighted processing on the input feature map in the channel dimension and the spatial dimension to obtain a reconstructed feature map; Processing the input feature map using the autoencoder to obtain a high-level feature vector; Using the fusion module to fuse the reconstructed feature map and the high-level feature vector to obtain a fused vector; Using the gated recurrent unit module to extract features from the fused vector to obtain time series features; Inputting the time series features into the Softmax classifier to obtain a classification result; Repeat the above model training process until the preset conditions are met to obtain a network abnormal traffic detection model.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the network abnormal traffic detection method according to any one of claims 1 to 6 is implemented.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the network abnormal traffic detection method according to any one of claims 1 to 6 is implemented.

10. A computer program product comprising instructions, which, when executed on a computer device, enables the computer device to execute the network abnormal traffic detection method according to any one of claims 1 to 6.