Lightweight malicious traffic classification method based on deep learning
The deep learning-based lightweight classification method for IoT devices addresses resource constraints and long-term dependency issues in malicious traffic detection, achieving efficient and accurate classification by reducing network layers and automating parameter adjustment.
Patent Information
- Application Number
- CN202510807000.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-07-15
AI Technical Summary
IoT devices are limited by computing and storage resources in malicious traffic detection. The existing detection methods ignore long-term dependencies, resulting in a degradation of detection performance. Model parameter adjustment depends on manual experience, and insufficient detection accuracy and generalization capabilities.
A lightweight malicious traffic classification method based on deep learning is adopted. By constructing a gated residual block network containing two parallel feature extraction paths, combining end-to-end PCA dimensional adaptive module and feature mapping, feature extraction and classification are realized, network layers are reduced, model parameters are automatically adjusted, detection accuracy and generalization capabilities are improved.
Implement efficient and accurate malicious traffic classification on IoT devices, reduce computing volume and storage needs, solve resource constraints, and improve detection performance.
Smart Images

Figure CN120321048A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of traffic detection, and particularly relates to a lightweight malicious traffic classification method based on deep learning. Background Art
[0002] With the rapid development of technologies such as wireless communication and edge computing, Internet of Things (IoT) devices have been widely used, enabling efficient communication and data exchange. IoT devices have revolutionized fields such as healthcare, transportation, and smart homes by providing unprecedented convenience and efficiency. However, the rapid growth of the IoT has also brought serious security problems.
[0003] In recent years, more and more IoT devices have started to be exposed to public networks and usually directly interact with the physical world to collect privacy data or control physical environment variables, making them one of the main targets of malicious attacks. IoT malicious traffic refers to data packets with malicious intent transmitted through IoT devices, which may contain viruses, Trojans, ransomware, etc., to steal information, damage systems, or launch cyberattacks. Malicious traffic attacks not only threaten privacy and data security but may also cause physical damage to critical infrastructure. Summary of the Invention
[0004] The purpose of this application is to provide a lightweight malicious traffic classification method based on deep learning to solve or alleviate the problems existing in the above-mentioned prior art.
[0005] To achieve the above purpose, this application provides the following technical solutions: This application provides a lightweight malicious traffic classification method based on deep learning, including: performing feature extraction on the traffic data to be detected through a feature extraction unit that includes two parallel feature extraction paths and multiple gated residual blocks connected in series in each feature extraction path, to obtain a feature sequence of the traffic data to be detected ; wherein, the gated residual block screens the traffic information in the traffic data to be detected through a set information selector; based on the feature sequence calculate the class probabilities of the traffic data to be detected for pre-divided traffic types.
[0006] Preferably, perform feature extraction on the traffic data to be detected through one or more feature extraction units connected in series to obtain a feature sequence of the traffic data to be detected 。
[0007] Preferably, in the feature extraction unit, the input feature sequence the feature sequence obtained by passing through multiple gated residual blocks of one feature extraction path , and the feature sequence the reverse feature sequence of The feature sequence obtained by multiple gated residual blocks passing through another parallel feature extraction path Perform feature concatenation to obtain the output feature of the feature extraction unit.
[0008] Preferably, the gated residual block is a single-layer network architecture including a cascaded causal convolutional layer and an information processing layer, and multiple information processors are included in the information processing layer; among them, the information processor at least includes a weight processor, an information selector, and a random dropper, and the weight processor is used to perform weight normalization on the traffic information in the traffic data to be detected, and the random dropper is used to randomly discard the traffic information in the traffic data to be detected.
[0009] Preferably, the information selector uses a gated linear function, and the random dropper uses a Droput function; in the th gated residual block corresponding to the feature extraction path, the feature sequence obtained by preprocessing After passing through padding operations, the feature sequence obtained by causal dilated convolution with a dilation rate of is further processed by weight normalization, a gated linear function, and a Droput function to obtain the feature sequence ; The feature sequence after convolution operation is added to the feature sequence to obtain the output feature of the th gated residual block; among them, , is a positive integer.
[0010] Preferably, in the information selector, the input feature sequence C undergoes feature screening operation through the Sigmoid function, and the output feature obtained by the feature screening operation is multiplied by the feature sequence C to obtain the output feature of the information selector.
[0011] Preferably, the feature sequence generated by dimensionality reduction operation on the original feature sequence of the traffic data to be detected obtained through the constructed end-to-end PCA dimensionality adaptive module is subjected to feature extraction through the feature extraction unit to obtain the feature sequence of the traffic data to be detected .
[0012] Preferably, in the end-to-end PCA dimensionality adaptive module, based on the principal component analysis method, the feature dimension of the original feature sequence is gradually decreased at a first fixed step size, and the first classification accuracy of the corresponding feature dimension is determined through the feature sequence obtained by each decreasing operation; Search for features around the corresponding feature sequence for one or more feature dimensions corresponding to the local maximum of the first classification accuracy at a second fixed step length, and calculate the second classification accuracy of the feature sequence corresponding to the feature dimensions obtained by the feature search based on the principal component analysis method; wherein, the second fixed step length is less than the first fixed step length. Extract features from the feature sequence corresponding to the smallest feature dimension among one or more feature dimensions corresponding to the local maximum of the second classification accuracy through a feature extraction unit to obtain the feature sequence of the traffic data to be detected. .
[0013] Preferably, map the feature length of the feature sequence to through a feature mapping operation to calculate the class probability that the traffic data to be detected is a pre-divided traffic type; wherein, is the total number of pre-divided traffic types.
[0014] Preferably, sequentially perform feature mapping on the feature sequence through global average pooling operation and pointwise convolution operation, and map the feature length to and input the feature sequence into a Softmax classifier to calculate the class probability that the traffic data to be detected is a pre-divided traffic type.
[0015] Beneficial effects: The lightweight malicious traffic classification method based on deep learning provided by the embodiments of the present application extracts features from the traffic data to be detected through a feature extraction unit including two parallel feature extraction paths, and each feature extraction path has multiple gated residual blocks connected in series to obtain the feature sequence of the traffic data to be detected. , and calculate the class probability that the traffic data to be detected is a pre-divided traffic type based on the feature sequence to realize the classification of the traffic data to be detected and determine whether the traffic data to be detected is malicious traffic. During the detection process, two parallel feature extraction paths are formed by gated residual blocks provided with information selectors to extract features from the traffic data to be detected, effectively reducing the problem of gradient disappearance when processing traffic data, and improving the ability to capture long-term dependence relationships in traffic data sequences, effectively ensuring the efficient and accurate classification of traffic data under the condition of using fewer network layers. Description of the Drawings
[0016] The specification drawings constituting a part of the present application are used to provide a further understanding of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. Among them: Figure 1Flow diagram of a lightweight malicious traffic classification method based on deep learning provided according to some embodiments of the present application; Figure 2 Logical diagram of a lightweight malicious traffic classification method based on deep learning provided according to some embodiments of the present application; Figure 3 Logical diagram of a gated residual block provided according to some embodiments of the present application; Figure 4 Logical diagram of a gated linear function provided according to some embodiments of the present application. Detailed implementation manners
[0017] The present application will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments. Each example is provided by way of explanation of the present application rather than limitation of the present application. In fact, those skilled in the art will clearly understand that modifications and variations can be made to the present application without departing from the scope or spirit of the present application. For example, features shown or described as part of one embodiment can be used in another embodiment to yield yet another embodiment. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the embodiments of the present invention shall fall within the scope of protection of the embodiments of the present invention.
[0018] Currently, in the detection of malicious traffic in the Internet of Things, limited by the limited computing and storage resources of Internet of Things devices, it is difficult to deploy a detection model in Internet of Things devices; moreover, existing malicious traffic detection methods focus on extracting time series features from network traffic and ignore dynamically selecting effective long-term dependencies, resulting in the detection model retaining information irrelevant to the classification task, thereby leading to a decline in the detection performance of the detection model.
[0019] In addition, the feature dimension reduction methods adopted by existing detection models often require manual experience to select appropriate parameters. The adjustment of model parameters is restricted by manual selection and cannot be automatically aligned with the ultimate optimization goal of the detection model during training, resulting in a reduction in the detection accuracy and generalization ability of the detection model.
[0020] Based on this, the embodiments of the present application provide a lightweight malicious traffic classification method based on deep learning, which classifies the grayscale image of the traffic data to be detected through a constructed malicious traffic classification model to obtain an output result of whether the traffic to be detected is benign traffic or malicious traffic. As Figures 1 to 4 shown, the method includes: Step S101, perform feature extraction on the traffic data to be detected through a feature extraction unit including two parallel feature extraction paths and a plurality of gated residual blocks connected in series in each feature extraction path to obtain a feature sequence of the traffic data to be detected .
[0021] Internet of Things (IoT) devices have been widely used, but they are vulnerable to malicious traffic attacks from malicious programs. It is very necessary to detect malicious traffic on the IoT. However, the storage space and computing power resources of IoT devices are limited, and the traffic scale in the IoT is large. Moreover, the model parameters and computational complexity of traditional neural networks are relatively high, making it impossible to be deployed on IoT devices. Therefore, by constructing a gated residual block with a single-layer network architecture containing an information selector, two parallel feature extraction paths are formed to extract features from the traffic data to be detected. During this process, the information selector is used to screen the traffic information in the traffic data to be detected. Under the premise of ensuring the efficient and accurate classification of traffic data, the number of network layers of the detection network, that is, the single-layer network architecture, is effectively reduced, realizing the lightweight of the detection model, enabling it to be deployed on resource-constrained IoT devices to detect malicious traffic.
[0022] When extracting features from the traffic data to be detected, it can be completed through one or more cascaded feature extraction units. The output of the previous feature extraction unit serves as the input of the next feature extraction unit, and the output of the last feature extraction unit serves as the feature sequence of the traffic data to be detected. The more cascaded feature extraction units there are, the more deep information of the traffic data to be detected can be extracted, and the higher the classification accuracy of the traffic data to be detected.
[0023] In each feature extraction path of the feature extraction unit, the number of cascaded gated residual blocks can be adjusted according to the different storage space and computing power resources of IoT devices. The fewer the number of cascaded gated residual blocks in the feature extraction path, the coarser the features of the traffic data to be detected are extracted, and the worse the ability to capture the long-term dependence relationship in the feature sequence of the traffic data to be detected.
[0024] In this application, the original feature sequence of the data to be detected obtained is dimension-reduced by the constructed end-to-end PCA dimension adaptive module to generate a feature sequence, which is then subjected to feature extraction by the feature extraction unit, and finally the feature sequence of the traffic data to be detected is obtained. Thereby, it effectively overcomes the problem that the existing detection model selects parameters through manual experience, avoids manual interference during the model training process, makes the adjustment of model parameters no longer restricted by humans, enables the model to automatically align with the final optimization goal (traffic classification) during training, and improves the detection accuracy and generalization ability of the model.
[0025] Among them, in the end-to-end PCA dimension adaptation module, first, based on the principal component analysis method, the feature dimension of the original feature sequence is gradually decreased according to the first fixed step size, and the first classification accuracy of the corresponding feature dimension is determined through the feature sequence obtained by each decreasing operation. That is to say, starting from the total feature dimension of the traffic data to be detected, the feature dimension is decreased according to the first fixed step size by the principal component analysis method, and after each decrease, the accuracy of the model on the dataset after dimensionality reduction is evaluated.
[0026] Then, one or more feature dimensions corresponding to the local maximum value of the first classification accuracy are searched for features around the corresponding feature sequence according to the second fixed step size smaller than the first fixed step size, and the second classification accuracy of the feature sequence corresponding to the feature dimension obtained by the feature search is calculated based on the principal component analysis method. That is to say, for one or more feature dimensions corresponding to the local maximum value of the first classification accuracy obtained by the first fixed step size, fine search is carried out around the corresponding feature dimensions according to the fine step size (the second fixed step size). For each dimension obtained by the fine search, the principal component analysis method is used to reduce the input dataset to the refined dimension, and then the accuracy (the second classification accuracy) of the model on the dataset of the refined dimension is calculated.
[0027] Finally, the feature sequence corresponding to the smallest feature dimension among one or more feature dimensions corresponding to the local maximum value of the second classification accuracy is subjected to feature extraction by the feature extraction unit to obtain the feature sequence of the traffic data to be detected . That is to say, the smallest refined dimension is found from one or more refined dimensions corresponding to the local maximum value of the second classification accuracy, and the corresponding feature sequence is input into the feature extraction unit for feature extraction, and the corresponding output result is used as the feature sequence of the traffic data to be detected . Thereby, in the end-to-end PCA dimension adaptation module, the accuracies of different feature dimensions on the dataset are automatically compared to determine the optimal feature dimension, avoiding the influence of human factors and effectively improving the detection efficiency of the model.
[0028] After the end-to-end PCA dimension adaptation module performs dimensionality reduction on the original feature sequence of the traffic data to be detected, the obtained feature sequence of the traffic data to be detected is input into the feature extraction unit for feature extraction. Specifically, the input feature sequence after passing through multiple cascaded gated residual blocks (forward gated residual blocks) of a feature extraction path, the feature sequence is obtained. The feature sequence input into the feature extraction unit in another feature extraction path, first, it is reversed in the time dimension to obtain the reversed feature sequence ; the reversed feature sequence After passing through multiple cascaded gated residual blocks (inverse gated residual blocks), a feature sequence is obtained. ; Finally, the feature sequence The feature sequence obtained by passing through multiple gated residual blocks is concatenated with the reversed feature sequence The feature sequence obtained by passing through multiple gated residual blocks to obtain the output feature of this feature extraction unit.
[0029] In this application, the gated residual block is constructed by a single-layer network structure including cascaded causal convolutional layers and information processing layers. The feature sequence obtained by preprocessing is sequentially processed by the causal convolutional layer and the information processing layer to process the feature information of the traffic data to be detected, and the corresponding output feature is obtained. Among them, multiple information processors in the information processing layer are used to process the traffic information in the traffic data to be detected. Specifically, the information processor includes at least a normalization processor, an information selector, and a random discard unit. The weight processor is used to perform weight normalization on the traffic information in the traffic data to be detected, the information selector is used to screen the traffic information in the traffic data to be detected, and the random discard unit is used to randomly discard the traffic information in the traffic data to be detected.
[0030] In this application, the information selector uses a gated linear function, and the random discard unit uses a Dropout function. In the th gated residual block in series in the feature extraction path, the feature sequence obtained by preprocessing first passes through ( , is a positive integer) padding operations, and then the feature sequence obtained by causal dilated convolution with a dilation rate of passes through weight normalization, a gated linear function, and a Dropout function to obtain the feature sequence ; Then, the feature sequence after convolution operation is added to the feature sequence to obtain the output feature of the th gated residual block.
[0031] It should be noted that there is no sequential relationship among weight normalization, the gated linear function, and the Dropout function. The feature sequence obtained by causal dilated convolution can be processed in the information processing layer in the order of the weight processor, the information selector, and the random discard unit, or in the order of the information selector, the weight processor, and the random discard unit, or in the order of the random discard unit, the information selector, and the weight processor, etc.
[0032] In a specific example, in the information selector, the input feature sequence C undergoes a feature screening operation through the Sigmoid function, and after multiplying the output feature obtained from the feature screening operation by the feature sequence C, the output feature of the information selector is obtained. Here, the input feature of the gated linear function is defined as , where is the input feature 's feature vector, are respectively the height, width, and number of channels of the input feature . The output feature obtained by performing a convolution operation on the input feature through the Sigmoid function is multiplied by the input feature to obtain the output feature . Specifically, according to the formula: the output feature of the gated linear function is determined; where the height, width, and number of channels of the output feature are respectively . Furthermore, the gradient of the output feature is: Thereby, during the process of the model gradient backpropagation, and is not directly attenuated by the activation function, effectively avoiding the problem of gradient disappearance that may occur during the backpropagation process as the number of network layers increases, enabling the model gradient to propagate smoothly forward and helping the gradient to flow better in the deep network.
[0033] In a specific example, after performing a dimensionality reduction operation on the original feature sequence of the obtained traffic data to be detected through the end-to-end PCA dimensionality adaptation module, the feature sequence of the traffic data to be detected obtained (here the feature sequence is the feature sequence obtained through preprocessing) first undergoes 2 padding operations in the first gated residual block of a feature extraction path, and then passes through a 1D causal dilated convolution with a dilation rate of 1 and 3 convolution kernels to obtain an output feature sequence with the same length as the feature sequence ; then, the obtained feature sequence with the same length of the output feature sequence undergoes weight normalization, a gated linear function, and a Droput function to obtain the feature sequence ; meanwhile, for the feature sequence (here the feature sequence is the feature sequence ) Perform a convolution operation and add the resulting feature sequence to the feature sequence to obtain the output feature of this gated residual block .
[0034] Next, use the output feature of the first gated residual block as the input of the second gated residual block (i.e., the feature sequence obtained from the preprocessing corresponding to the second gated residual block ), and perform feature extraction operations in the second gated residual block. Thereby, in the model, the flow features in the flow data to be detected are screened out through the gated linear function, and the model structure is redesigned using the gated residual block containing the gated linear function, effectively reducing the number of unnecessary residual blocks in the model, accelerating the inference rate of the flow while reducing the computational amount, and maintaining the classification performance of the model.
[0035] Correspondingly, in another feature extraction path, after performing a dimensionality reduction operation on the original feature sequence of the flow data to be detected obtained by the end-to-end PCA dimensionality adaptive module, the resulting feature sequence of the flow data to be detected is first reversed in the time dimension to obtain a reversed feature sequence , and the reversed feature sequence is the feature sequence obtained from the preprocessing corresponding to the first gated residual block in this feature extraction path . Here, the reversed feature sequence passes through multiple cascaded gated residual blocks containing the gated linear function in this feature extraction path, and the output feature of the last gated residual block is reversed in the time dimension to obtain the output feature sequence of this feature extraction path .
[0036] Here, it should be noted that the feature sequence of the traffic data to be detected is a set composed of the features of the traffic data to be detected. The features of the traffic data to be detected mainly include: flow duration, packet header length, protocol type, duration, TCP flags (number of FIN flags, number of SYN flags, number of PSH flags), etc. Among them, the flow duration (flow_duration) records the total duration from connection establishment to termination. Malicious botnets or DDoS attacks usually exhibit abnormally short or long flow durations; the packet header length (Header_Length) refers to the size of the header of each network packet. An abnormal header length may indicate a malformed packet attack or protocol abuse; the protocol type (Protocoltype) specifies the network protocol used by the traffic, such as TCP, UDP, or ICMP, etc. Malicious traffic tends to use specific protocols to bypass detection or execute attacks. For example, some malware may use UDP for C2 communication; the duration (Duration) is similar to the flow duration, but sometimes refers to a finer-grained packet or session duration, which helps to further analyze the activity and periodicity of the traffic; the number of FIN (fin_flag_number) flags counts the number of occurrences of the FIN (end) flag in the connection. A normal TCP connection is closed through the FIN flag, and an abnormal number of FINs means an incomplete connection closure or scanning behavior; the SYN (syn_flag_number) flag is usually used for the establishment of a TCP connection. The number of SYN flags records the number of occurrences of the STN (synchronization) flag. A large number of SYN requests without subsequent responses may point to a SYN Flood attack (a type of denial-of-service attack); the PSH (psh_flag_number) flag indicates that the receiving end should immediately push the data to the application instead of waiting for the buffer to fill. The number of PSH flags counts the number of occurrences of the PSH (push) flag. An abnormal number of PSHs may be related to the data transfer mode, such as rapid data sending or instant messaging of malware.
[0037] Step S102, based on the feature sequence Calculate the class probability that the traffic data to be detected belongs to the pre-divided traffic types.
[0038] Benign traffic refers to the data stream in a computer network or online platform that originates from legitimate users or legitimate programs; malicious traffic refers to the data stream in a computer network or online platform that originates from malicious activities, malicious programs, or malicious behaviors. The purpose of benign traffic is legal, harmless, compliant with norms, and does not have any malicious nature; the purpose of malicious traffic may include, but is not limited to, unauthorized access, data leakage, system damage, fraud, virus propagation, denial-of-service attacks (DDos), etc. The operation of benign traffic is consistent with the normal expected behavior of the system or service; malicious traffic usually violates the regulations of the system or service and may cause damage or improper benefits to it.
[0039] In this application, the data features of the traffic data to be detected are extracted by constructing two parallel feature extraction paths, and each feature extraction path serially connects multiple feature extraction units including gated residual blocks with gated linear functions, and then according to the obtained feature sequence of the traffic data to be detected , the class probability that the traffic data to be detected is a pre-divided traffic type is calculated.
[0040] Generally, the obtained feature sequence of the traffic data to be detected can be directly After passing through a fully connected operation and input into a Softmax classifier, the class probability that the traffic data to be detected is a pre-divided traffic type is calculated. However, directly inputting the feature sequence After passing through a fully connected operation and input into a Softmax classifier to classify the traffic data to be detected, the required computational amount is relatively high, and the traffic inference speed is very low, which is not conducive to deployment on resource-constrained Internet of Things devices.
[0041] Based on this, in this application, the feature length of the feature sequence is mapped to ( is the total number of pre-divided traffic types) to calculate the class probability that the traffic data to be detected is a pre-divided traffic type. Specifically, the feature sequence is successively subjected to global average pooling operation and pointwise convolution operation for feature mapping, and the feature sequence with the feature length mapped to is input into a Softmax classifier to calculate the class probability that the traffic data to be detected is a pre-divided traffic type.
[0042] Thereby, the data computational amount is greatly reduced, and the traffic inference speed is improved; effectively solves the problem of large traffic scale in the Internet of Things and difficulty in real-time traffic detection; through further lightweighting of the model structure of the traffic detection model, effectively solves the problems of limited storage space and limited computing power resources of Internet of Things devices, enabling it to be deployed on resource-constrained Internet of Things devices.
[0043] In the embodiment of the present application, the constructed lightweight malicious traffic classification model is connected in series with 1 end-to-end adaptive PCA dimension adjustment module, 2 feature extraction units, 1 global average pooling layer, 1 pointwise convolution layer, and 1 Softmax classifier. In the feature extraction unit, there are two parallel feature extraction paths, and each feature extraction path includes 2 serially connected gated residual blocks. Among them, the number of serially connected gated residual blocks in each feature extraction path can be adaptively adjusted according to the requirements of calculation accuracy and calculation speed.
[0044] The original feature sequence of the obtained traffic data to be detected first undergoes a dimensionality reduction operation through the end-to-end adaptive PCA dimension adjustment module, and then enters 2 serially connected feature extraction units for feature extraction, and outputs a feature sequence ; the feature sequence Then, it successively undergoes global average pooling operation and pointwise convolution operation for feature mapping, and maps its length to the number of pre-divided traffic types ; finally, the feature sequence with the feature length mapped to is input into the Softmax classifier to calculate the category probability that the traffic data to be detected is a pre-divided traffic type, so as to realize the classification of the traffic data to be detected, and obtain the result that it is benign traffic or malicious traffic.
[0045] In the global average pooling layer, an average operation is performed on each channel of the feature sequence to generate a one-dimensional vector with the same number as the number of feature channels, so as to effectively avoid parameter overfitting; then the one-dimensional vector with the same number as the number of feature channels is input into the pointwise convolution layer. There are multiple neurons in the pointwise convolution layer, and each neuron corresponds to a traffic type, so as to effectively combine the one-dimensional vector with the same number as the number of feature channels with the traffic category discrimination information (pre-divided traffic types), effectively avoid information loss, and improve the classification accuracy of the traffic data to be detected.
[0046] In this application, after the lightweight malicious traffic classification model is constructed, the input feature sequence dataset is divided into a training set and a test set; the lightweight malicious traffic classification model is trained using the training set; the trained lightweight malicious traffic classification model is tested using the test set.
[0047] In practical applications, it can be set in the form of software, such as designed as an independent APP or an embedded software that can be called at any time, and applied in computer terminals to achieve the deployment of a lightweight malicious traffic classification model. Among them, the computer terminal includes a memory, a processor, and a computer program stored on the memory and executable on the processor. For example, the computer terminal can be a smart phone, a tablet computer, a notebook computer, etc. that can execute programs. In some embodiments, the processor can be a Central Processing Unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor is generally used to control the overall operation of the computer device. In this embodiment, the processor is used to run the program code stored in the memory or process data. When the processor executes the program, it can implement the steps of the lightweight malicious traffic classification method of the present invention.
[0048] In addition, the lightweight malicious traffic classification method based on deep learning described above can also be designed into a readable storage medium, such as a USB key. The readable storage medium stores computer program instructions. When the computer program instructions are read and run by a processor, the steps of the lightweight malicious traffic classification method are executed. Made in the form of a USB key, it can be plugged into a computer, for example, by electronic plugging. The computer reads and executes the computer program instructions in the USB key, and can solve the technical problem that the current deep learning model has a high computational amount and parameters and cannot be deployed on resource-constrained Internet of Things devices.
[0049] Whether non-embedded or embedded, they can be classified into corresponding lightweight malicious traffic classification devices. The lightweight malicious traffic classification device uses a lightweight malicious traffic classification model to classify the original traffic sequence through the designed lightweight malicious traffic classification model, and obtains an output result of whether it is benign traffic or malicious traffic.
[0050] In the description of the present invention, the terms "first" and "second" are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present invention, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically defined.
[0051] The terms "one embodiment", "some embodiments", "an example", "a specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0052] The foregoing are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. A lightweight malicious traffic classification method based on deep learning, characterized in that Including: The traffic data to be detected is subjected to feature extraction by a feature extraction unit that includes two parallel feature extraction paths, and a plurality of gated residual blocks are connected in series in each feature extraction path, so as to obtain a feature sequence of the traffic data to be detected. Among them, the gated residual block screens the traffic information in the traffic data to be detected through a set information selector. Based on the feature sequence Calculate the category probability that the traffic data to be detected belongs to the pre-divided traffic types.
2. The method according to claim 1, wherein The traffic data to be detected is subjected to feature extraction through one or more serially connected feature extraction units to obtain a feature sequence of the traffic data to be detected .
3. The method according to claim 1, wherein In the feature extraction unit, the input feature sequence The feature sequence obtained through multiple gated residual blocks of a feature extraction path , and the feature sequence The reverse feature sequence of The feature sequence obtained through multiple gated residual blocks of another parallel feature extraction path Perform feature concatenation to obtain the output feature of the feature extraction unit.
4. The method according to claim 1, characterized in that The gated residual block is a single-layer network architecture including a cascaded causal convolutional layer and an information processing layer, and multiple information processors are included in the information processing layer; wherein, the information processor at least includes a weight processor, an information selector, and a random dropper, and the weight processor is used to perform weight normalization on the traffic information in the traffic data to be detected, and the random dropper is used to randomly discard the traffic information in the traffic data to be detected.
5. The method according to claim 4, wherein The information selector uses a gated linear function, and the random dropper uses a Droput function; In the th gated residual block in the concatenated feature extraction path, the feature sequence obtained by preprocessing after passing through padding operations, and after the feature sequence obtained by causal dilated convolution with a dilation rate of is processed by weight normalization, gated linear function, and Droput function, the feature sequence is obtained; Feature sequence The feature sequence after convolution operation and the feature sequence are added together to obtain the output feature of the th gated residual block; where , is a positive integer.
6. The method according to claim 5, characterized in that, In the information selector, the input feature sequence C undergoes a feature screening operation through a Sigmoid function, and after multiplying the output feature obtained from the feature screening operation by the feature sequence C, the output feature of the information selector is obtained.
7. The method according to claim 1, wherein The feature sequence generated by performing a dimensionality reduction operation on the original feature sequence of the data to be detected obtained through the constructed end-to-end PCA dimensionality adaptive module is subjected to feature extraction by a feature extraction unit to obtain the feature sequence of the traffic data to be detected 。 8. The method according to claim 7, wherein In the end-to-end PCA dimension adaptive module, based on the principal component analysis method, the feature dimension of the original feature sequence is gradually decreased at a first fixed step length, and the first classification accuracy of the corresponding feature dimension is determined through the feature sequence obtained by each decreasing operation; One or more feature dimensions corresponding to the local maximum value of the first classification accuracy are searched for features around the corresponding feature sequence at a second fixed step length, and the second classification accuracy of the feature sequence corresponding to the feature dimension obtained by the feature search is calculated based on the principal component analysis method; wherein, the second fixed step length is less than the first fixed step length; The feature sequence corresponding to the smallest feature dimension among one or more feature dimensions corresponding to the local maximum of the second classification accuracy is subjected to feature extraction by the feature extraction unit to obtain the feature sequence of the traffic data to be detected .
9. The method according to claim 1, wherein Map the feature length of the feature sequence through a feature mapping operation to to calculate the class probability that the traffic data to be detected is a pre-divided traffic type; where is the total number of pre-divided traffic types.
10. The method according to claim 9, wherein The feature sequence is successively subjected to global average pooling operation and pointwise convolution operation for feature mapping, and the feature sequence with the feature length mapped to is input into the Softmax classifier to calculate the class probability that the traffic data to be detected is the traffic type pre-divided.
Citation Information
Patent Citations
DGA domain name detection model and method based on gated convolution and LSTM
CN115242484A
Lightweight malicious traffic classification method based on deep learning
CN117336057A
Ultrahigh-order QAM signal nonlinear compensation method and system based on time characteristic memory neural network
CN119807630A
Video super-resolution reconstruction method and system
CN120013766A
Conditional Computation For Continual Learning
US20210150345A1