Flow intrusion detection method, device, equipment, storage medium and program product
By adopting lightweight neural network structure and Fourier transform in traffic intrusion detection, the computing resource consumption problem caused by excessive complexity of deep learning models is solved, efficient and real-time traffic intrusion detection is achieved, and the detection accuracy is improved.
Patent Information
- Application Number
- CN202211352783.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-31
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-10-31
AI Technical Summary
In the prior art, when deep learning models are used for traffic intrusion detection, they are too complex and occupy a large amount of computing and memory resources, making it difficult to ensure the real-time operational resources and detection required for deployment without strong computing resources.
The traffic intrusion detection method based on the lightweight neural network structure is adopted to extract feature vectors from traffic data packets through convolutional neural networks, and compress the feature vectors through the full connection layer to reduce node connectivity and save computational amount. At the same time, the frequency domain characteristics of the data stream are extracted using Fourier transform to capture the mutual relationship between traffic data packets.
It improves computing efficiency, reduces the consumption of computing resources, ensures real-time detection, and improves the accuracy of traffic intrusion detection. The representation mode of normal traffic and attack traffic is significantly differentiated in the frequency domain.
Smart Images

Figure CN115695002B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the fields of network security and computer network management, and particularly to a traffic intrusion detection method, device, equipment, storage medium, and program product based on a neural network. Background Art
[0002] With the development of Internet technology, network traffic intrusion detection is an important means to ensure network security. For example, traditional intrusion detection methods mainly include signature-based methods, statistic-based methods, information entropy-based methods, and so on. However, the above methods often have a high false alarm rate, and the defined detection rules are prone to becoming obsolete over time and with the evolution of attacks.
[0003] Currently, some efforts have been made to use deep learning for traffic intrusion detection. However, deep learning models usually have too high a complexity, which will consume a large amount of computing and memory resources, bringing certain difficulties to practical applications. For example, without powerful computing resources as hardware support, if the amount of computation and the number of parameters of a deep learning model are large, it is impossible to ensure the computing resources required for deployment and the real-time performance of detection. Summary of the Invention
[0004] In view of the above problems, the present disclosure provides a traffic intrusion detection method, device, equipment, storage medium, and program product based on a lightweight neural network structure that improves computing efficiency.
[0005] According to a first aspect of the present disclosure, there is provided a traffic intrusion detection method, including: obtaining a plurality of traffic data packets; respectively extracting a plurality of feature vectors from the plurality of traffic data packets based on a neural network; extracting frequency domain features of a data stream by performing a Fourier transform on a data stream formed by the plurality of feature vectors; and detecting the type of the data stream according to the frequency domain features through a classifier.
[0006] According to an embodiment of the present disclosure, respectively extracting a plurality of feature vectors from the plurality of traffic data packets based on a neural network includes: respectively extracting a plurality of initial feature vectors from the plurality of traffic data packets through multiple input channels of a convolutional layer of the neural network; and for the initial feature vector of each traffic data packet among the plurality of traffic data packets, compressing the initial feature vector through a fully connected layer of the neural network to obtain a feature vector, where the dimension of the feature vector is smaller than that of the initial feature vector.
[0007] According to an embodiment of the present disclosure, a plurality of traffic data packets include M data packets, a convolutional layer of a neural network includes M input channels and N output channels, where M and N are positive integers; multiple initial feature vectors are respectively extracted from the plurality of traffic data packets through multiple groups of input channels of the convolutional layer of the neural network, including: dividing the M input channels into M groups of channels, each group of channels includes one input channel, and each input channel is provided with a convolutional kernel; performing depth convolution operations on the M data packets respectively through the M convolutional kernels of the M groups of channels to obtain M first feature vectors; performing pointwise convolution operations on the M first feature vectors, and passing through the N output channels to obtain N initial feature vectors.
[0008] According to an embodiment of the present disclosure, performing depth convolution operations on the M data packets respectively through the M convolutional kernels of the M groups of channels to obtain M first feature vectors, including: According to
[0009]
[0010] performing a depth convolution operation on the M data packets, where K d,m is the convolutional kernel of the m-th input channel among the M input channels, Q l+d-1,m is a one-dimensional convolutional feature map of length l in the m-th channel, and d is the pixel point in the one-dimensional convolutional feature map traversed by the convolutional kernel.
[0011] According to an embodiment of the present disclosure, performing pointwise convolution operations on the M first feature vectors, and passing through the N output channels to obtain N initial feature vectors, including: According to
[0012]
[0013] performing pointwise convolution operations on the M first feature vectors, where W d,m,n is the convolutional kernel of the m-th input channel among the M input channels and the n-th output channel among the N output channels, Q l+d,m is a one-dimensional convolutional feature map of length l in the m-th input channel, and d is the pixel point in the one-dimensional convolutional feature map traversed by the convolutional kernel.
[0014] According to an embodiment of the present disclosure, a pointwise convolution operation is performed on M first feature vectors, and N initial feature vectors are obtained through N output channels. It further includes: dividing M input channels and N output channels into G groups to obtain G grouped convolution channel matrices, where G is a positive integer; for each of the G grouped convolution channel matrices, splitting the grouped convolution channel matrix into a first sub-convolution channel matrix and a second sub-convolution channel matrix, the first sub-convolution channel matrix includes C output channels, and the second sub-convolution channel matrix includes C input channels, where C is a positive integer; respectively performing pointwise convolution on M first feature vectors through G first sub-convolution channel matrices to obtain G×C second feature vectors; performing channel shuffle on the G×C second feature vectors to obtain G×C third feature vectors; and respectively performing pointwise convolution on the G×C third feature vectors through G second sub-convolution channel matrices to obtain N initial feature vectors.
[0015] According to an embodiment of the present disclosure, splitting the grouped convolution channel matrix into a first sub-convolution channel matrix and a second sub-convolution channel matrix includes: According to
[0016]
[0017] Splitting the grouped convolution channel matrix into a first sub-convolution channel matrix and a second sub-convolution channel matrix into two sub-convolutions, where, is the grouped convolution channel matrix of the L-th layer in the convolutional layer, having e input channels and f output channels, is the first sub-convolution channel matrix, is the second sub-convolution channel matrix, and the (L - 1)-th layer is an intermediate layer having C output channels, where M = G×e, N = G×f, and e and f are positive integers.
[0018] According to an embodiment of the present disclosure, the initial feature vectors include an initial packet header feature vector and an initial payload feature vector; through multiple groups of input channels of the convolutional layer of the neural network, multiple initial feature vectors are respectively extracted from multiple traffic data packets, including: for each of the multiple traffic data packets, obtaining the header data and payload data of each traffic data packet; constructing a first convolutional model for the header data and a second convolutional model for the payload data, the number of input channels and output channels of the first convolutional model being less than the number of input channels and output channels of the second convolutional model; and respectively extracting the initial packet header feature vector and the initial payload feature vector from the header data and the payload data through the first convolutional model and the second convolutional model.
[0019] According to an embodiment of the present disclosure, the multiple initial feature vectors include N initial feature vectors; by passing through the fully connected layer of the neural network, the initial feature vectors are compressed to obtain feature vectors, including: dividing the N initial feature vectors into multiple byte segments, each byte segment including N eigenvalues; and respectively performing intra-group compression on the N eigenvalues of each byte segment among the multiple byte segments through the fully connected layer to obtain multiple feature vectors.
[0020] According to an embodiment of the present disclosure, by performing a Fourier transform on the data stream composed of multiple feature vectors, the frequency domain features of the data stream are extracted, including: performing a Fourier transform on the data stream to obtain frequency domain variables; and calculating the amplitude of the frequency domain variables to obtain frequency domain features.
[0021] According to an embodiment of the present disclosure, performing a Fourier transform on the data stream to obtain frequency domain variables includes:
[0022] According
[0023] y m = F seq (F h (p m ))
[0024] Calculate the frequency domain variable y m , where F h is an extraction function for extracting the mutual relationship between the internal features of multiple traffic packets, F seg is an extraction function for extracting the sequential relationship of multiple traffic packets in the data stream, and p m is the feature vector of the m-th traffic packet among the multiple traffic packets.
[0025] A second aspect of the present disclosure provides a traffic intrusion detection device, including: an acquisition module for acquiring multiple traffic packets; a first extraction module for respectively extracting multiple feature vectors from the multiple traffic packets based on a neural network; a second extraction module for extracting the frequency domain features of the data stream by performing a Fourier transform on the data stream composed of the multiple feature vectors; and a detection module for detecting the type of the data stream according to the frequency domain features through a classifier.
[0026] A third aspect of the present disclosure provides an electronic device, including: one or more processors; a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the above-mentioned traffic intrusion detection method.
[0027] A fourth aspect of the present disclosure further provides a computer-readable storage medium, on which executable instructions are stored, and when the instructions are executed by a processor, the processor is caused to execute the above-mentioned traffic intrusion detection method.
[0028] The fifth aspect of the present disclosure also provides a computer program product, including a computer program which, when executed by a processor, implements the above-mentioned traffic intrusion detection method.
[0029] Through the embodiments of the present disclosure, for each traffic data packet, an efficient lightweight convolutional neural network is used to extract features from the original bytes, and a fully connected layer is used to compress and abstract the features extracted by the convolutional layer, thereby saving computational amount by reducing node connectivity and reducing computing resources by sparsification. For all traffic data packets, the mutual relationship between traffic data packets is captured, and the Fourier transform is used to extract the frequency domain features of the data stream composed of the feature vectors of multiple data packets. The characterization patterns of normal traffic and attack traffic are significantly distinguishable in the frequency domain. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, the above content and other objects, features and advantages of the present disclosure will become clearer. In the drawings:
[0031] Figure 1 Schematically shows a flowchart of a traffic intrusion detection method according to an embodiment of the present disclosure;
[0032] Figure 2 Schematically shows a schematic diagram of a traffic intrusion detection method according to an embodiment of the present disclosure;
[0033] Figure 3 Schematically shows a flowchart of extracting feature vectors according to an embodiment of the present disclosure;
[0034] Figure 4A Schematically shows a schematic diagram of extracting initial feature vectors according to an embodiment of the present disclosure;
[0035] Figure 4B Schematically shows a schematic diagram of compressing initial feature vectors according to an embodiment of the present disclosure;
[0036] Figure 5 Schematically shows a flowchart of extracting frequency domain features according to an embodiment of the present disclosure;
[0037] Figure 6 Schematically shows the data stream processing rates of different devices according to an embodiment of the present disclosure;
[0038] Figure 7 Schematically shows a structural block diagram of a traffic intrusion detection device according to an embodiment of the present disclosure; and
[0039] Figure 8 Schematically shows a block diagram of an electronic device suitable for implementing the traffic intrusion detection method according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, for the sake of explanation, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure. However, obviously, one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present disclosure.
[0041] The terms used herein are merely for describing specific embodiments and are not intended to limit the present disclosure. The terms "including", "comprising", etc. used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0042] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0043] In the case of using expressions such as "at least one of A, B, and C, etc.", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include, but is not limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).
[0044] In the technical solution of the present disclosure, the processing of collection, storage, use, processing, transmission, provision, disclosure, and application, etc. of the user's personal information involved all comply with the provisions of relevant laws and regulations, necessary confidentiality measures are taken, and it does not violate public order and good customs. In the technical solution of the present disclosure, before obtaining or collecting the user's personal information, the authorization or consent of the user is obtained.
[0045] Figure 1 The flowchart of the traffic intrusion detection method according to an embodiment of the present disclosure is schematically shown.
[0046] As Figure 1 shown, the traffic intrusion detection method of this embodiment includes operations S110 to S140.
[0047] In operation S110, a plurality of traffic data packets are obtained.
[0048] In the embodiments of the present disclosure, the traffic data packet is the object of traffic intrusion detection. The traffic data packet can be obtained from a network firewall or from an intrusion detection system deployed in an edge device.
[0049] In operation S120, multiple feature vectors are respectively extracted from multiple traffic data packets based on a neural network.
[0050] In the embodiments of the present disclosure, for each of the multiple traffic data packets, in a set of input channels of the neural network, in-packet modeling is performed on the traffic data packet to process the traffic data packet. The in-packet modeling is used to convert the original bytes included in the traffic data packet into representative features to achieve the extraction of the feature vector of the traffic data packet. Each set of output channels is not connected to form sparsification.
[0051] The neural network can be a convolutional neural network, adopting local connection. The distinguishable features in the bytes of the original traffic data packet are localized short fields. For example, "HTTP / 1.1", "id=xxx&Submit=Submit". The convolutional neural network is very good at perceiving local details. When processing localized attack features, the convolutional neural network can use a convolutional kernel with a fixed receptive field to perform feature extraction in a sliding manner in a local area, and can well capture local details. Compared with a dense fully connected neural network, the convolutional neural network adopting local connection can greatly reduce the number of parameters and the amount of calculation, and has higher operation efficiency.
[0052] In operation S130, by performing a Fourier transform on the data stream composed of multiple feature vectors, the frequency domain features of the data stream are extracted.
[0053] In the embodiments of the present disclosure, after obtaining the feature vector of a single traffic data packet, inter-packet modeling can be performed on the traffic data packet. The inter-packet modeling is used to capture the mutual relationship between traffic data packets. Multiple feature vectors are combined into a data stream, and discrete Fourier transform is used for inter-packet modeling to form a data packet sequence.
[0054] The data packet sequence is a time series. The inter-packet modeling analysis of traffic data packets based on the Fourier transform can fully extract the mutual relationship between traffic data packets and capture the timing features of the data packet sequence. Using the frequency domain characteristics to characterize the inter-packet relationship and sequence features, the data packet sequence is transformed from the time domain to the frequency domain to convert the data stream into a stream representation vector. The stream representation vector can include the frequency domain features of the data stream.
[0055] The patterns of normal traffic and attack traffic are significantly distinguishable in the frequency domain. For example, the frequencies of normal traffic tend to have a large and disorderly distribution range, while the frequency distributions of fixed types of attack traffic, such as botnet (Bot) and SQL injection attack traffic, tend to be relatively single and concentrated.
[0056] In the embodiments of the present disclosure, machine learning models such as recurrent neural networks or self-attention mechanisms can also be used to extract the frequency domain features of the data stream. Preferably, the Fourier transform may not require learning additional parameters and has a lower computational complexity.
[0057] In operation S140, the classifier detects the type of the data stream according to the frequency domain features.
[0058] In the embodiments of the present disclosure, the classifier can classify the flow representation vectors obtained through packet modeling and determine whether the traffic data packets corresponding to the flow representation vectors are normal traffic or different types of attack traffic.
[0059] According to the embodiments of the present disclosure, first, through in-packet modeling for each data packet, the feature vectors of each traffic data packet are extracted, and then the Fourier transform is used to extract the frequency domain features of the data stream composed of multiple feature vectors to capture the mutual relationship between multiple traffic data packets. By extracting features from data in two forms, namely data packets and data streams, the accuracy of traffic intrusion detection is improved.
[0060] Figure 2 A schematic diagram of a traffic intrusion detection method according to an embodiment of the present disclosure is schematically shown.
[0061] As Figure 2 shown, in-packet modeling analysis 220 is performed on multiple traffic data packets 210. The in-packet modeling analysis 220 is used to extract the feature vectors of each of the multiple traffic data packets 210 among the multiple traffic data packets 210 respectively. The in-packet modeling analysis 220 may include a feature extraction stage 221 and a feature compression stage 222. The dimension of the feature vectors obtained in the feature extraction stage 221 may be too large. To further reduce the amount of calculation and reduce the consumption of computing resources, the dimension of the feature vectors can be compressed through the feature compression stage 222 to obtain low-dimensional feature vectors 230.
[0062] Multiple low-dimensional feature vectors 230 form a data stream, and inter-packet modeling analysis 240 is performed. Through the Fourier transform, the frequency domain features of the low-dimensional feature vectors 230 are extracted. The classifier 250 detects the type of the traffic data packet according to the frequency domain features.
[0063] Through the embodiments of the present disclosure, the input channels of a convolutional neural network are grouped, and intra-packet modeling is performed within each group of input channels to process each traffic data packet separately, thereby reducing the connectivity of nodes in the neural network and saving computational effort. Dividing the input channels into different groups and processing the traffic data packets separately can be understood as decomposing the overall problem into the sum of multiple different sub-problems. Each group performs model operations within the group, establishing connectivity only within each group, thereby achieving sparsification and reducing computing resources. Using local connections within the group to replace dense fully connected connections can reduce the model complexity, reducing the complexity from O(n) to O(n / G), where G is the number of groups of input channels.
[0064] Figure 3 Schematically shows a flowchart of extracting feature vectors according to an embodiment of the present disclosure.
[0065] As Figure 3 shown, the steps of operation S120 for extracting multiple feature vectors from multiple traffic data packets based on a neural network include operations S310 to S340.
[0066] In operation S310, multiple initial feature vectors are extracted from multiple traffic data packets respectively through multiple groups of input channels of the convolutional layer of the neural network.
[0067] In operation S320, for the initial feature vectors of each traffic data packet among the multiple traffic data packets, the initial feature vectors are compressed through the fully connected layer of the neural network to obtain feature vectors, where the dimension of the feature vectors is smaller than that of the initial feature vectors.
[0068] In the embodiments of the present disclosure, the multiple traffic data packets may include M data packets, and the convolutional layer of the neural network may include M input channels and N output channels, where M and N are positive integers.
[0069] The steps of extracting multiple initial feature vectors from multiple traffic data packets respectively through multiple groups of input channels of the convolutional layer of the neural network in operation S310 include: dividing the M input channels into M groups of channels, each group of channels includes one input channel, and each input channel is provided with one convolutional kernel; performing depth convolution operations on the M data packets respectively through the M convolutional kernels of the M groups of channels to obtain M first feature vectors; performing pointwise convolution operations on the M first feature vectors, and passing through N output channels to obtain N initial feature vectors.
[0070] Exemplarily, the process of performing depth convolution operations on the M data packets respectively through the M convolutional kernels of the M groups of channels to obtain M first feature vectors can be calculated according to the following formula:
[0071]
[0072] Kd,m is the convolution kernel of the m-th input channel among M input channels, Q l+d-1,m is a one-dimensional convolutional feature map of length l in the m-th channel, and d is the pixel point in the one-dimensional convolutional feature map traversed by the convolution kernel. The one-dimensional convolutional feature map is used to represent the features of the traffic data packet.
[0073] A convolution kernel is applied to each input channel respectively for depth convolution operation. The number of output channels of the depth convolution can be the same as the number of input channels of the depth convolution. The depth convolution process does not expand the dimension of the feature map. Further, the feature information of different channels can be effectively fused through pointwise convolution to ensure the information flow between channels.
[0074] Exemplarily, the process of performing pointwise convolution operation on M first feature vectors through N output channels to obtain N initial feature vectors can be calculated according to the following formula:
[0075]
[0076] W d,m,n is the convolution kernel of the m-th input channel among M input channels and the n-th output channel among N output channels, Q l+d,m is a one-dimensional convolutional feature map of length l in the m-th input channel, and d is the pixel point of the one-dimensional convolutional feature map traversed by the convolution kernel.
[0077] The number of input channels of the first convolutional layer in the pointwise convolution is equal to the number of output channels of the convolutional layer in the depth convolution. The convolutional layer in the depth convolution includes M input channels and M output channels. The first convolutional layer in the pointwise convolution also includes M input channels. The convolutional layer of the pointwise convolution can include N output channels.
[0078] In the embodiments of the present disclosure, group convolution can also be performed on the pointwise convolution process to reduce node connectivity, reduce the amount of calculation, and improve the model efficiency. For example, the M channels input to the pointwise convolutional layer are divided into G groups, and pointwise convolution operations are performed within each group. To further increase the information exchange between channels, the input and output channels of a group convolution can be decomposed into two sub-convolutions. For example, it can be considered that the input and output channels of each group convolution form a group convolution channel matrix, and each group convolution channel matrix is split into two sub-convolution channel matrices. Channel shuffle is used between the two sub-convolutions to increase the information interaction between channels.
[0079] The steps of performing pointwise convolution operations on M first eigenvectors and obtaining N initial eigenvectors through N output channels include: dividing M input channels and N output channels into G groups to obtain G grouped convolution channel matrices, where G is a positive integer; for each grouped convolution channel matrix among the G grouped convolution channel matrices, splitting the grouped convolution channel matrix into a first sub-convolution channel matrix and a second sub-convolution channel matrix, the first sub-convolution channel matrix includes C output channels, and the second sub-convolution channel matrix includes C input channels, where C is a positive integer; respectively performing pointwise convolution on M first eigenvectors through G first sub-convolution channel matrices to obtain G×C second eigenvectors; performing channel shuffle on the G×C second eigenvectors to obtain G×C third eigenvectors; and respectively performing pointwise convolution on the G×C third eigenvectors through G second sub-convolution channel matrices to obtain N initial eigenvectors.
[0080] Exemplarily, the grouped convolution channel matrix can be split into a first sub-convolution channel matrix and a second sub-convolution channel matrix according to the following formula for splitting into two sub-convolutions:
[0081]
[0082] is the grouped convolution channel matrix of the L-th layer in the convolutional layer, having e input channels and f output channels, is the first sub-convolution channel matrix, is the second sub-convolution channel matrix, and the (L - 1)-th layer is an intermediate layer with C output channels, where M = G×e, N = G×f, and e and f are positive integers.
[0083] For example, the pointwise convolution includes 8 input channels and 8 output channels. The 8 input channels and 8 output channels are divided into 2 groups, and each group includes 4 input channels and 4 output channels. The 4 input channels and 4 output channels of each group can form a grouped convolution channel matrix Grouped convolution channel matrix can be split into a first sub-convolution channel matrix and a second sub-convolution channel matrix The first sub-convolution channel matrix includes 4 input channels and 2 output channels, and the second sub-convolution channel matrix includes 2 input channels and 4 output channels.
[0084] For each group of the first sub-convolution channel matrix and the second sub-convolution channel matrix Perform channel shuffle on the two sets of data output by the first sub-convolution channel matrix Randomly shuffle and reassign the channels of the two groups, thereby creating the flow of information between groups.
[0085] Pointwise convolution is performed on two split sub-convolution channel matrices. This can not only establish channel sparse connections through grouped convolution, significantly reducing the computational complexity of pointwise convolution, but also maintain connectivity between groups through inter-group channel shuffling, thereby further enhancing the representation ability of the output features.
[0086] Figure 4A Schematically shows a schematic diagram of extracting an initial feature vector according to an embodiment of the present disclosure.
[0087] As Figure 4A shown, during the depth convolution 410, 4 first feature vectors are output. During the pointwise convolution 420, the 4 input channels and 4 output channels corresponding to the 4 first feature vectors are split into two groups, obtaining two 2×2 grouped convolution channel matrices. Each 2×2 grouped convolution channel matrix includes 2 input channels and 2 output channels. Each 2×2 grouped convolution channel matrix is split into two 1×1 sub-convolution channel matrices. After passing through the first 1×1 sub-convolution channel matrix, the two groups of second feature vectors are shuffled in channels, obtaining two groups of third feature vectors. After passing through the second 1×1 sub-convolution channel matrix, the initial feature vector is obtained.
[0088] In the embodiment of the present disclosure, the step of extracting multiple initial feature vectors from multiple traffic data packets respectively through multiple groups of input channels of the convolutional layer of the neural network in operation S310 includes: for each traffic data packet of the multiple traffic data packets, obtaining the header data and payload data of each traffic data packet; constructing a first convolutional model for the header data and a second convolutional model for the payload data, where the number of input channels and output channels of the first convolutional model is less than that of the second convolutional model; and extracting an initial header feature vector and an initial payload feature vector from the header data and the payload data respectively through the first convolutional model and the second convolutional model.
[0089] Both the header data and the payload data of the traffic data packet can be used as detection input data. The header data and the payload data of the traffic data packet can be processed separately, applying different numbers of channels respectively. The semantic information carried by the header data is relatively less, while the payload data carries more semantic information and diverse features. The payload data may contain important information related to the attack content. Through traffic intrusion detection, application layer attacks located within the payload can be detected, such as SQL injection, cross-site scripting attacks, remote malicious code exploitation, and so on. A higher-dimensional feature space can be established for the payload data for representation.
[0090] For example, for the packet header data, only the network layer header data and the transport layer header data of the traffic data packet can be detected. The data link layer header does not contain valid information with attack characteristics. In addition, to mitigate overfitting in detection training, the source port, checksum, identification, sequence number, and acknowledgment number can be excluded from the detection input, and the IP address can be anonymized. For example, the length of the packet header data of a single traffic data packet is fixed at 30 bytes, while the length of the payload data can be 80 bytes.
[0091] In the first convolutional model for the packet header data and the second convolutional model for the payload data, the starting layers of the first and second convolutional models can both be standard convolution, which is used to initially extract shallow features from the original traffic data packet. The first and second convolutional models can then perform maxpooling operations and successively stack Depthwise convolutional units and Pointwise convolutional units to expand the receptive field. Each Pointwise convolutional unit has two sub-convolutions. The first and second convolutional models can also improve the training efficiency through Batch Normalization (BN) and linearly activate the output of each layer through the ReLU activation function.
[0092] The content of the first convolutional model for the packet header data and the second convolutional model for the payload data can be shown in the following table.
[0093]
[0094] In the embodiments of the present disclosure, in the case where there are N initial feature vectors among multiple initial feature vectors. The steps of compressing the initial feature vectors through the fully connected layer of the neural network in operation S320 to obtain feature vectors include: dividing the N initial feature vectors into multiple byte segments, each byte segment including N feature values; and respectively performing intra-group compression on the N feature values of each byte segment among the multiple byte segments through the fully connected layer to obtain multiple feature vectors.
[0095] After the convolution operation is completed, the dimension of the initial feature vector output by the convolutional layer may be too large, for example, greater than 1000. If the initial feature vector is directly used as the feature vector of the traffic data packet, a large amount of computational cost will be generated in the inter-packet modeling analysis stage. Adding a fully connected layer after the convolutional layer for feature compression can generate a concise packet feature vector. Feature compression can also implicitly filter out redundant features with low effectiveness through end-to-end supervised training.
[0096] During the feature compression process, the initial header feature vector and the initial payload feature vector are respectively divided into multiple short byte segments (seg1, seg2...), and operations are performed within the short byte segments. The fully connected layer is divided into g sparse connections, where g is the number of groups. Group feature compression can divide the traffic data packet into different byte segments according to the network layer and the transport layer, and evenly divide the initial header feature vector and the initial payload feature vector.
[0097] Since the compression methods for the initial header feature vector and the initial payload feature vector are similar. This disclosure only exemplarily shows the compression process of the payload feature vector.
[0098] Exemplarily, the initial payload feature vector output by the convolutional layer can be represented by the following formula:
[0099]
[0100] B is the initial payload feature vector output by the convolutional layer, and b i is the i-th byte of the initial payload feature vector. The initial payload feature vector has a total of k bytes, and the number of output channels of the convolutional layer is n. For each byte b i , the eigenvalue b i1 ...b in with n channels. Perform a splitting operation on the initial payload feature vector B, and split the initial payload feature vector B into multiple short byte segments seg i to achieve grouping.
[0101] The initial payload feature vector B can be divided according to the following formula:
[0102]
[0103] seg i is the i-th byte segment with n-channel features [b (i-1)R : b iR , R is the number of elements in each group, and g is the number of groups. After grouping, the feature seg i of each byte segment will be flattened into [b 1,(i-1) R,...b 1,iR ,...b n,(i-1) ,...b n,iR to perform intra-group feature compression and obtain the reduced feature of each group. The fusion layer will fuse the reduced features output by different groups and output a low-dimensional feature vector, such as a feature vector with a dimension less than 20.
[0104] Figure 4BSchematically shows a schematic diagram of compressing an initial feature vector according to an embodiment of the present disclosure. As shown in FIG. 4, feature extraction is performed on the traffic data packet 430 to obtain an initial packet header feature vector 440 and an initial payload feature vector 450. The initial packet header feature vector 440 and the initial payload feature vector 450 are divided. The initial packet header feature vector 440 can be divided into 3 short byte segments seg 1 , seg 2 and seg 3 , and the initial payload feature vector 450 can be divided into 5 short byte segments seg 1 , seg 2 , seg 3 , seg 4 and seg 5 . For example, if the length of the initial payload feature vector is 50 bytes, the initial payload feature vector is divided into 5 segments: [0:10], [10:20], [20:30], [30:40], [40:50]. Intra-group feature compression is performed on the features of each short byte segment to obtain the reduced features of each group. The reduced features output by different groups will be fused through the fusion layer 460, and a low-dimensional feature vector 470 will be output.
[0105] Figure 5 Schematically shows a flowchart of extracting frequency domain features according to an embodiment of the present disclosure.
[0106] As Figure 5 shown, in operation S130, the steps of performing a Fourier transform on the data stream to obtain frequency domain variables include operations S510 to S520.
[0107] In operation S510, a Fourier transform is performed on the data stream to obtain frequency domain variables.
[0108] In operation S520, the amplitude of the frequency domain variable is calculated to obtain frequency domain features.
[0109] In an embodiment of the present disclosure, let S be a two-way stream containing M feature vectors:
[0110] S = [p 1 , p 2 ,..., p m ,..., p M
[0111] p m is the feature vector of the mth data packet in the data stream.
[0112] The data stream S is converted into a data packet sequence, and the discrete Fourier transform is used to characterize the inter-packet relationship and sequence features of the data packets using frequency domain characteristics, so as to realize the conversion of the data stream S into a stream representation vector y m .
[0113] Exemplarily, a Discrete Fourier Transform (DFT) can decompose an input sequence into a combination of a set of harmonics in the frequency domain:
[0114]
[0115] x m is an input sequence of length m, where m ∈ [0, M - 1].
[0116] When applying the Discrete Fourier Transform in the step of performing a Fourier transform on the data stream in operation S510 to obtain the frequency domain variables, the two-dimensional Discrete Fourier Transform of the data packet sequence can be calculated according to the following formula:
[0117] y m = F seq (F h (p m ))), 0 ≤ m ≤ M - 1
[0118] Multiple y m are frequency domain variables, F h is an extraction function for extracting the mutual relationship between the internal features of multiple traffic data packets, F seq is an extraction function for extracting the sequential relationship of multiple traffic data packets in the data stream, p m is the feature vector of the m-th traffic data packet among the M traffic data packets in multiple traffic data packets.
[0119] The output of the two-dimensional Discrete Fourier Transform can be expressed as follows:
[0120] y m = a m + b m j
[0121]
[0122]
[0123] where U is the feature dimension of a single data packet, M is the data packet sequence dimension, and x UM is the data packet sequence.
[0124] Exemplarily, the amplitude of the frequency domain variables can be calculated according to the following formula:
[0125]
[0126] Due to the symmetry of the frequency domain, the first half of the variable Y m is used as the frequency domain feature. The final output of the Discrete Fourier Transform is a flow characterization vector Y m . The classifier will classify Y mClassify to determine whether the traffic data packet is normal traffic or different types of attack traffic.
[0127] The present disclosure also provides verification of the traffic intrusion detection method of the present disclosure to verify the detection performance. The verification process includes testing and evaluating both the detection effect and the detection efficiency. The experiment uses two public data sets for the experiment: CIC-IDS2017 and CIC-IDS2018.
[0128] Compare the traffic intrusion detection method of the present disclosure with multiple algorithms, such as AE+MLP, LCB+MLP, LCB+LSTM, 1D-CNN, and MobileNet. Among them, AE+MLP uses manually designed features for detection, uses an autoencoder to compress the input features in a semi-supervised manner, and uses a classifier to classify specific attack categories. LCB (Light-weighted Convolution Block) is a lightweight convolutional neural network module. LCB+MLP uses LCB for in-packet modeling of data packets and directly outputs the traffic data packet vector to the classifier without inter-packet modeling. LCB+LSTM uses LCB for in-packet modeling of data packets and uses Long Short Term Memory Neural Network (LSTM) for inter-packet modeling. 1D-CNN is a standard convolutional operation with two convolutional layers, 8 channels and 16 channels respectively. MobileNet is an efficient CNN architecture that uses depthwise separable convolutions to build a lightweight convolutional neural network model.
[0129] During the comparative experiment process, this section sets n p = 8 and n b = 80, where n p is the first N data packets used in a data stream, and n b is the number of bytes used in the payload. Name the traffic intrusion detection method of the present disclosure LiteIDS. The following table shows the multi-classification results of multiple detection methods.
[0130]
[0131]
[0132] Compared with AE+MLP, LiteIDS has higher precision, recall, and F1 value, which indicates that automatically extracting representative features from the original data packets is more effective than using manually designed features. LiteIDS has better detection effect compared with LCB+MLP, indicating the effectiveness of data packet inter-packet modeling through Fourier transform. LiteIDS transforms the original sequence of data packets from the time domain to the frequency domain, capturing richer frequency domain features while extracting the inter-packet relationship of data packets. Compared with the data packet inter-packet modeling method based on long short-term memory network (LCB+LSTM), LiteIDS can achieve similar detection effect while having less computational complexity and higher detection performance. Moreover, compared with the standard convolutional 1D-CNN and MobileNet without introducing simplification operations, the effectiveness of LiteIDS does not decrease significantly, proving the effectiveness and feasibility of the lightweight operation of this method.
[0133] In addition, compared with other deep learning-based anomaly detection methods, LiteIDS also has the characteristics of low computational cost and small model size. This disclosure evaluates the detection efficiency of the verification of the traffic intrusion detection method of this disclosure, and the evaluation metrics include MACC, the number of parameters, and the model size. MACC (Multiply-Accumulate Operation) is the sum of all multiplications and additions during the model operation process, directly reflecting the computational complexity in the model. The number of parameters and the model size reflect the number of variables that need to be trained in the model.
[0134] The model efficiency is evaluated by calculating the average value of the above evaluation metrics for a single traffic data packet. During the detection process of a single traffic data packet, since the load detection module of the traffic data packet does not work when the traffic data packet does not have a valid payload, the computational complexity of the load is multiplied by a probability factor λ, where λ is the probability that a single traffic data packet has a valid payload. For example, the proportion of data packets with a length not exceeding 64 bytes accounts for about 60% of the total number of traffic data packets, and the probability factor λ can be set to 0.4. During the process of data packet inter-packet modeling, the computational complexity of a data stream can be evenly distributed to each traffic data packet to calculate the average value of the computational efficiency of a single traffic data packet.
[0135] The following table shows the average MACC, the number of parameters, the model size, and the F1 value of a single traffic data packet for different detection algorithms. When evaluating the model complexity of Kitsune, the calculation is performed under the conditions of n = 198, k = 10, and k = 20, where n is the number of features of each traffic data packet and k is the number of autoencoders. The compression ratio of the hidden layer of the autoencoder is 3 / 4.
[0136] Algorithm MACC Number of parameters Average model size F1 score LiteIDS 4,263 1,107 4.3k 99.27% LCB+LSTM 6,192 3,260 13.1k 99.41% 1D-CNN 73,920 16,866 67.5k 99.69% MobileNet 3,805,000 268,321 1.03k 99.92% Kitsune(k = 20) 3800 3,800 15.2k Kitsune(k = 10) 6160 6,160 25.2k
[0137] The MACC value of LiteIDS is only 5.76% of that of 1D-CNN and 0.12% of that of MobileNet. It can be seen therefrom that LiteIDS can reduce the computational complexity of the standard one-dimensional convolution by up to 94.24% without significant decrease in detection accuracy. In addition, the MACC value of LiteIDS is only 68.84% of that of LCB+LSTM, indicating that using Fourier transform for inter-packet modeling and sequence feature extraction of data packets has lower computational complexity than LSTM. LiteIDS also has a MACC value equivalent to or lower than that of Kitsune, indicating that this method has equivalent or higher model efficiency than Kitsune. At the same time, LiteIDS also has the smallest number of parameters and average model size, which are 6.56% of 1D-CNN, 0.41% of MobileNet, and 33.95% of LCB+LSTM respectively, indicating that this method has the most concise parameters and smaller model occupancy during deployment.
[0138] In addition, the present disclosure evaluates the performance of running the traffic intrusion detection method of the present disclosure. The running environment is a desktop computer and a software router, and the software router also performs routing and forwarding tasks simultaneously. Figure 6 Schematically shows the data stream processing rate of different devices according to an embodiment of the present disclosure.
[0139] As Figure 6 shown, as the number of data packets n p in the data stream increases, the stream processing rate decreases continuously in both environments. When n p = 8, the stream processing rates of the system are 3010 flow / s and 2342 flow / s. According to the experimental results in the hyperparameter sensitivity section, even when n p = 4, the model still has a good F1 value. Therefore, the highest stream processing rates of the system in the two devices can reach 3293 flow / s and 2564 flow / s. The experimental results show that the traffic intrusion detection method of the present disclosure can run under a software router with the current model complexity.
[0140] Based on the above traffic intrusion detection method, the present disclosure also provides a traffic intrusion detection device. The following will be combined with Figure 7 to describe this device in detail.
[0141] Figure 7 Schematically shows the structural block diagram of the traffic intrusion detection device according to an embodiment of the present disclosure.
[0142] As Figure 7 shown, the traffic intrusion detection device 700 of this embodiment includes an acquisition module 710, a first extraction module 720, a second extraction module 730, and a detection module 740.
[0143] The obtaining module 710 is configured to obtain a plurality of traffic data packets. In one embodiment, the obtaining module 710 may be configured to perform the operation S110 described above, which will not be elaborated herein.
[0144] The first extraction module 720 is configured to respectively extract a plurality of feature vectors from the plurality of traffic data packets based on a neural network. In one embodiment, the first extraction module 720 may be configured to perform the operation S120 described above, which will not be elaborated herein.
[0145] The second extraction module 730 is configured to extract the frequency-domain features of the data stream by performing a Fourier transform on the data stream constituted by the plurality of feature vectors. In one embodiment, the second extraction module 730 may be configured to perform the operation S130 described above, which will not be elaborated herein.
[0146] The detection module 740 is configured to detect the type of the data stream according to the frequency-domain features through a classifier. In one embodiment, the detection module 740 may be configured to perform the operation S140 described above, which will not be elaborated herein.
[0147] According to an embodiment of the present disclosure, any plurality of the obtaining module 710, the first extraction module 720, the second extraction module 730, and the detection module 740 may be combined and implemented in one module, or any one of them may be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present disclosure, at least one of the obtaining module 710, the first extraction module 720, the second extraction module 730, and the detection module 740 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or any other reasonable way of integrating or packaging circuits, etc., implemented by hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Alternatively, at least one of the obtaining module 710, the first extraction module 720, the second extraction module 730, and the detection module 740 may be at least partially implemented as a computer program module, which can perform corresponding functions when the computer program module is run.
[0148] Figure 8 A block diagram of an electronic device suitable for implementing a particle Monte Carlo dose calculation method according to an embodiment of the present disclosure is schematically shown.
[0149] As Figure 8As shown, the electronic device 800 according to an embodiment of the present disclosure includes a processor 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage section 808 into a random access memory (RAM) 803. The processor 801 may include, for example, a general-purpose microprocessor (e.g., CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 801 may also include on-board memory for caching purposes. The processor 801 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.
[0150] In the RAM 803, various programs and data required for the operation of the electronic device 800 are stored. The processor 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. The processor 801 performs various operations of the method flow according to an embodiment of the present disclosure by executing the programs in the ROM 802 and / or the RAM 803. It should be noted that the program may also be stored in one or more memories other than the ROM 802 and the RAM 803. The processor 801 may also perform various operations of the method flow according to an embodiment of the present disclosure by executing the programs stored in one or more memories.
[0151] According to an embodiment of the present disclosure, the electronic device 800 may further include an input / output (I / O) interface 805, and the input / output (I / O) interface 805 is also connected to the bus 804. The electronic device 800 may further include one or more of the following components connected to the I / O interface 805: an input section 806 including a keyboard, a mouse, etc.; an output section 807 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a LAN card, a modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the I / O interface 805 as needed. A removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 810 as needed so that a computer program read from it can be installed into the storage section 808 as needed.
[0152] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist separately without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the method according to an embodiment of the present disclosure is implemented.
[0153] According to an embodiment of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present disclosure, the computer-readable storage medium may include one or more memories other than the ROM 802 and / or RAM 803 and / or ROM 802 and RAM 803 described above.
[0154] An embodiment of the present disclosure also includes a computer program product, which includes a computer program, and the computer program contains program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to enable the computer system to implement the traffic intrusion detection method provided by the embodiment of the present disclosure.
[0155] When the computer program is executed by the processor 801, it executes the above functions defined in the system / apparatus of the embodiment of the present disclosure. According to an embodiment of the present disclosure, the above-described systems, apparatuses, modules, units, etc. can be implemented by computer program modules.
[0156] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and be downloaded and installed through the communication part 809, and / or be installed from the removable medium 811. The program code contained in the computer program can be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0157] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 809, and / or be installed from the removable medium 811. When the computer program is executed by the processor 801, it executes the above functions defined in the system of the embodiment of the present disclosure. According to an embodiment of the present disclosure, the above-described systems, devices, apparatuses, modules, units, etc. can be implemented by computer program modules.
[0158] In accordance with embodiments of the present disclosure, program code for executing the computer programs provided by the embodiments of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. The programming languages include, but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (for example, by using an Internet service provider to connect through the Internet).
[0159] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0160] Those skilled in the art can understand that the features recited in the various embodiments and / or claims of the present disclosure can be combined or combined in various ways, even if such combinations or combinations are not explicitly recited in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features recited in the various embodiments and / or claims of the present disclosure can be combined and combined in various ways. All such combinations and / or combinations fall within the scope of the present disclosure.
[0161] The embodiments of the present disclosure have been described above. However, these embodiments are merely for illustrative purposes and not for limiting the scope of the present disclosure. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. The scope of the present disclosure is defined by the appended claims and their equivalents. Without departing from the scope of the present disclosure, those skilled in the art can make various substitutions and modifications, and these substitutions and modifications should all fall within the scope of the present disclosure.
Claims
1. A traffic intrusion detection method, comprising: Obtaining a plurality of traffic data packets; Extracting a plurality of feature vectors from the plurality of traffic data packets respectively based on a neural network; By performing a Fourier transform on the data stream composed of the plurality of feature vectors, extracting the frequency domain features of the data stream; and Using a classifier to detect the type of the data stream according to the frequency domain features; Wherein, for each traffic data packet of the plurality of traffic data packets, in a set of input channels of the neural network, in-packet modeling is performed on the traffic data packet, and the in-packet modeling is used to convert the original bytes included in the traffic data packet into a feature vector; Wherein, inter-packet modeling is performed on the traffic data packets, and the inter-packet modeling is used to capture the mutual relationship between the traffic data packets to form a data packet sequence, and the data packet sequence is a time sequence.
2. The traffic intrusion detection method according to claim 1, wherein, The extracting a plurality of feature vectors from the plurality of traffic data packets respectively based on a neural network includes: Extracting a plurality of initial feature vectors from the plurality of traffic data packets respectively through multiple groups of input channels of the convolutional layer of the neural network; and For the initial feature vector of each traffic data packet among the plurality of traffic data packets, through the fully connected layer of the neural network, compressing the initial feature vector to obtain the feature vector, wherein the dimension of the feature vector is smaller than that of the initial feature vector.
3. The traffic intrusion detection method according to claim 2, wherein, The plurality of traffic data packets include M data packets, and the convolutional layer of the neural network includes M input channels and N output channels, where M and N are positive integers; The extracting a plurality of initial feature vectors from the plurality of traffic data packets respectively through multiple groups of input channels of the convolutional layer of the neural network includes: Dividing the M input channels into M groups of channels, each group of channels includes one input channel, and each of the input channels is provided with a convolutional kernel; Performing depth convolution operations on the M data packets respectively through the M convolutional kernels of the M groups of channels to obtain M first feature vectors; Performing pointwise convolution operations on the M first feature vectors, and obtaining N of the initial feature vectors through the N output channels.
4. The traffic intrusion detection method according to claim 3, wherein, The performing depth convolution operations on the M data packets respectively through the M convolutional kernels of the M groups of channels to obtain M first feature vectors includes: According to Perform a deep convolution operation on M data packets, where is the convolution kernel of the m-th input channel among the M input channels, is the one-dimensional convolution feature map of length l in the m-th channel, and d is the pixel point in the one-dimensional convolution feature map traversed by the convolution kernel.
5. The traffic intrusion detection method according to claim 3, wherein, The performing pointwise convolution operations on the M first feature vectors, and obtaining N of the initial feature vectors through the N output channels includes: According to Perform a pointwise convolution operation on the M first eigenvectors, where is the convolution kernel for the m-th input channel among the M input channels and the n-th output channel among the N output channels, is the one-dimensional convolution feature map of length l in the m-th input channel, and d is the pixel point of the one-dimensional convolution feature map traversed by the convolution kernel.
6. The traffic intrusion detection method according to claim 4, wherein, The performing pointwise convolution operations on the M first feature vectors, and obtaining N of the initial feature vectors through the N output channels further includes: Dividing the M input channels and the N output channels into G groups to obtain G grouped convolution channel matrices, where G is a positive integer; For each of the G grouped convolutional channel matrices, split the grouped convolutional channel matrix into a first sub-convolutional channel matrix and a second sub-convolutional channel matrix, where the first sub-convolutional channel matrix includes C output channels and the second sub-convolutional channel matrix includes C input channels, and C is a positive integer; Perform pointwise convolution on the M first feature vectors through the G first sub-convolutional channel matrices respectively to obtain G×C second feature vectors; Perform channel shuffle on the G×C second feature vectors to obtain G×C third feature vectors; and Perform pointwise convolution on the G×C third feature vectors through the G second sub-convolutional channel matrices respectively to obtain N initial feature vectors.
7. The traffic intrusion detection method according to claim 6, wherein, The splitting of the grouped convolutional channel matrix into a first sub-convolutional channel matrix and a second sub-convolutional channel matrix includes: According to Split the grouped convolution channel matrix into a first sub-convolution channel matrix and a second sub-convolution channel matrix, which are split into two sub-convolutions, where is the grouped convolution channel matrix of the L-th layer in the convolutional layer, having e input channels and f output channels, is the first sub-convolution channel matrix, is the second sub-convolution channel matrix, and the (L - 1)-th layer is an intermediate layer with C output channels, where M = G×e, N = G×f, and e and f are positive integers.
8. The traffic intrusion detection method according to claim 2, wherein, The initial feature vectors include an initial packet header feature vector and an initial payload feature vector; the extracting of multiple initial feature vectors from multiple traffic packets through multiple input channels of the convolutional layer of the neural network includes: For each traffic packet of the multiple traffic packets, obtain the packet header data and payload data of each traffic packet; Construct a first convolutional model for the packet header data and a second convolutional model for the payload data, where the number of input channels and output channels of the first convolutional model is less than the number of input channels and output channels of the second convolutional model; and Extract the initial packet header feature vector and the initial payload feature vector from the packet header data and the payload data through the first convolutional model and the second convolutional model respectively.
9. The traffic intrusion detection method according to claim 2, wherein, The multiple initial feature vectors include N initial feature vectors; the compressing of the initial feature vectors through the fully connected layer of the neural network to obtain the feature vectors includes: Divide the N initial feature vectors into multiple byte segments, and each byte segment includes N feature values; and Perform intra-group compression on the N feature values of each byte segment in the multiple byte segments through the fully connected layer to obtain the multiple feature vectors.
10. The traffic intrusion detection method according to claim 2, wherein, The extracting of the frequency domain features of the data stream by performing Fourier transform on the data stream composed of the multiple feature vectors includes: Perform Fourier transform on the data stream to obtain frequency domain variables; and Calculate the amplitude of the frequency domain variables to obtain frequency domain features.
11. The traffic intrusion detection method according to claim 10, wherein, The performing of Fourier transform on the data stream to obtain frequency domain variables includes: According to Calculate the frequency domain variable y m , where F h is an extraction function for extracting the mutual relationship between the internal features of the multiple traffic data packets, and F seq is an extraction function for extracting the sequential relationship of the multiple traffic data packets in the data stream, and p m is the feature vector of the m-th traffic data packet among the multiple traffic data packets.
12. A traffic intrusion detection device, including: An acquisition module for acquiring multiple traffic packets; A first extraction module for extracting multiple feature vectors from multiple traffic packets respectively based on a neural network; And A second extraction module, configured to extract frequency-domain features of the data stream by performing a Fourier transform on the data stream composed of the multiple feature vectors; and a detection module, configured to detect the type of the data stream according to the frequency-domain features by using a classifier; wherein, the first extraction module is further configured to, for each of the multiple traffic data packets, perform in-packet modeling on the traffic data packet within a set of input channels of the neural network, and the in-packet modeling is used to convert the original bytes included in the traffic data packet into feature vectors; wherein, the second extraction module is further configured to perform inter-packet modeling on the traffic data packets, and the inter-packet modeling is used to capture the mutual relationship between the traffic data packets to form a data packet sequence, and the data packet sequence is a time sequence.
13. An electronic device comprising: one or more processors; a storage device for storing one or more programs, wherein, when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the method according to any one of claims 1 to 11.
14. A computer-readable storage medium, having stored thereon executable instructions, which when executed by a processor cause the processor to execute the method according to any one of claims 1 to 11.
15. A computer program product, comprising a computer program, which when executed by a processor implements the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Network intrusion detection method based on improved convolutional neural network
CN111275165A
Network traffic classification device and method based on short-time Fourier transform
CN113449768A