Malicious traffic detection method based on selective state space model

Through selective state space model and deep learning technology, the characteristics of malicious traffic data are extracted and processed, and the existing detection methods are solved in the face of low detection accuracy when they have not seen malicious traffic, and the detection accuracy of high accuracy is achieved.

CN119995920AActive Publication Date: 2025-05-13ZHEJIANG UNIV

Patent Information

Application Number
CN202411862094.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-17
Publication Date
2025-05-13
Estimated Expiration
2044-12-17

AI Technical Summary

Technical Problem

When facing malicious traffic that has never been seen, existing malicious traffic detection methods are prone to misjudgment, and it is difficult to have strong local feature processing and global information processing capabilities at the same time, resulting in low detection accuracy.

Method used

A malicious traffic detection method based on a selective state space model is adopted to convert the original traffic data into grayscale map data, potential features are extracted, and detection models are constructed through the Ghost module, S2M module and Mamba module, pre-training and comparative learning are carried out to enhance the model's ability to identify malicious traffic data.

Benefits of technology

It realizes high-accuracy zero-sample malicious traffic detection, enhances the detection ability of new malicious traffic, and improves the generalization performance of the model and the usability and robustness in actual scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119995920A_ABST
    Figure CN119995920A_ABST
Patent Text Reader

Abstract

The invention discloses a malicious traffic detection method based on a selective state space model. An existing method is high in misjudgment rate and poor in detection effect. According to the method, zero-sample malicious traffic detection is realized through conversion of original traffic data into a grey-scale map, potential feature extraction, detection model training, known malicious traffic data pre-training and the like. The method comprises the following steps: firstly, preparing known malicious flow data and carrying out data enhancement on the known malicious flow data, then using the known malicious flow data and synthesized malicious flow data after data enhancement for pre-training of a zero-sample malicious flow detection model together, and finally identifying the type of malicious flow data never seen through the pre-trained model. Malicious traffic data which are never seen and obtained from the network are preprocessed and then input into the pre-trained zero sample malicious traffic detection model, and whether the current traffic is malicious traffic is judged. According to the method, effective semantic information is reserved, redundant semantic information is removed, the effectiveness of feature extraction is improved, and then the accuracy of the model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of network security and network defense, in particular to the field of malicious traffic detection, and specifically relates to a malicious traffic detection method based on a selective state space model. Background Art

[0002] As network attack methods continue to evolve, general threat methods not only include traditional malware and viruses, but also expand to more covert advanced persistent threats and zero-day attacks. These attacks often use complex encryption technology, obfuscation technology, and social engineering methods, making it difficult for traditional detection methods to identify and respond to these increasingly complex network threats. Therefore, the research and development of advanced malicious traffic detection technology has become particularly important. In the field of malicious traffic detection, zero-sample learning has attracted attention for its potential in dealing with unknown threats. In the context of malicious traffic detection, this means that the detection system needs to be able to identify and respond to new malicious traffic that has not appeared before, which may use advanced encryption technology to evade traditional detection methods.

[0003] In order to cope with the challenges facing network security, research on malicious traffic detection technology has begun to develop in the direction of intelligence. The current detection methods are mainly divided into three categories: behavior analysis-based methods, machine learning-based methods, and deep learning methods. Among them, the behavior analysis-based method identifies potential malicious activities by monitoring abnormal behavior patterns in network traffic. This method does not rely on known malicious features, but analyzes the statistical characteristics and behavior patterns of network traffic to detect unknown or variant malicious traffic. However, as the research deepens, this method may produce false positives due to the complexity and variability of the network environment.

[0004] Therefore, with the development of network technology, detection methods based on machine learning have begun to be widely studied. Such methods train models to identify the characteristics of malicious traffic to achieve automated threat identification and response. However, these methods are not dynamic and adaptable enough. Network attackers often update their attack strategies and techniques, which makes it possible that models trained based on historical data may not be able to effectively identify emerging attack patterns.

[0005] As attackers advance in technology, the encryption and obfuscation methods of malicious traffic are also constantly upgrading, making it increasingly difficult for traditional feature-based detection methods to adapt to new threat situations. Therefore, researchers have begun to explore more advanced deep learning detection methods to more effectively identify and defend against unknown malicious traffic. These methods improve the accuracy and efficiency of detection by building complex models to learn the complex patterns of network traffic. However, traditional deep learning methods also have limitations, especially when facing unknown malicious traffic. These methods usually rely on a large amount of labeled data, which is often impractical to obtain in practical applications. Therefore, it is particularly important to develop a method that can detect unknown malicious traffic without direct samples. Zero-shot learning is a special machine learning method whose core goal is to correctly identify new categories without having seen any labeled samples of new categories during training. At present, the research on zero-shot malicious traffic detection is still in its infancy. Some existing methods try to improve the model's ability to identify unknown malicious traffic through techniques such as transfer learning and meta-learning. For example, some research works build cross-domain feature representations so that the model can learn features useful for unknown categories from known categories.

[0006] When faced with malicious traffic that has never been seen before, malicious traffic often uses various techniques to evade detection, and traditional malicious traffic detection methods mainly rely on known features for detection, which usually performs poorly in this case. At the same time, unknown malicious traffic on the network is emerging in an endless stream, and previous malicious traffic detection methods are unable to cope with these unknown malicious traffic, often resulting in serious misjudgments.

[0007] The Chinese invention patent application with application number CN202411018887.7 discloses a method for detecting encrypted malicious traffic based on traffic interaction behavior and attention mechanism. The method realizes the detection capability of unknown encrypted malicious traffic by constructing a traffic interaction graph and using a multi-head self-attention model to extract traffic features. Its limitation is that it relies on the constructed traffic interaction graph and the features extracted by the self-attention mechanism. When facing new unknown malicious traffic, it may not be able to effectively capture all key features, resulting in a decrease in detection accuracy. In addition, the method requires a large amount of computing resources to process and compare a large amount of network traffic data, which may encounter cost and efficiency problems in actual deployment. The Chinese invention patent application with application number CN202410254307.8 discloses a method for identifying malicious encrypted traffic based on multi-granularity feature extraction. The method constructs multiple classifiers to analyze the abnormal traffic of encrypted traffic and obtains abnormal encrypted traffic of unknown malicious or attack. Its limitation is that the multi-granularity feature extraction involved in the method is very complex, which increases the difficulty of implementation, and it relies on a large amount of labeled data to train the model, and the generalization ability on unknown or zero-sample malicious traffic may be limited. The Chinese invention patent application with application number CN202311794220.1 discloses a malicious encrypted traffic classification method and terminal in a network information system scenario. This method designs a more efficient lightweight network model, uses a classification model to classify and identify unknown encrypted traffic, and analyzes it in combination with performance evaluation indicators. Its limitation is that the ability to extract potential features from traffic data is insufficient, which may affect the quality and depth of feature extraction.

[0008] There are many shortcomings in the existing technology. First, the existing methods are usually effective in detecting known malicious traffic, but they often make misjudgments when facing malicious traffic that has never been seen before; second, the existing methods are usually unable to have strong local feature processing and global information processing capabilities at the same time, and it is difficult to achieve high-accuracy detection; third, most of the existing methods use neural network models that have good effects in other fields, and do not design special models for the characteristics of malicious traffic data with less semantic information. Summary of the invention

[0009] The purpose of the present invention is to provide a malicious traffic detection method based on a selective state space model in a new malicious traffic detection scenario, aiming at the situation where malicious traffic that has never been seen before is easy to escape detection, and the existing detection methods have low accuracy and low practical usability. Through the steps of converting raw traffic data into grayscale images, extracting potential features of traffic data, training detection models, and pre-training a large amount of known malicious traffic data, high-accuracy zero-sample malicious traffic detection is achieved. The present invention aims to propose a zero-sample malicious traffic detection method for malicious traffic types that have never been seen before, further improve the detection accuracy of malicious traffic data that has never been seen before from the aspects of data preprocessing, potential feature extraction, detection model design, and model pre-training, and enhance the usability and robustness in actual scenarios.

[0010] The method of the present invention specifically comprises: Step (1) Prepare training data for the malicious traffic classification model: Considering that the data volume and model performance present a power-law distribution, and that large data volume plays an important role in the emergent effect on model performance, download known malicious traffic data from the Internet as raw traffic data; convert the raw traffic data into grayscale data that is more convenient for classification through slicing, truncation, zero padding, and pixel value mapping operations. The specific data conversion includes:

[0011] (1-1) Data slicing: The original PCAP file is split into data streams. The data packets in the PCAP file are grouped according to the five-tuple information and arranged in chronological order; the five-tuple information is the source IP address, destination IP address, source port, destination port, and protocol. This step ensures that each data stream is a set of data packets with the same five-tuple information, thereby maintaining the relevance and order between data packets.

[0012] (1-2) Data uniform length processing: the first part of each file is bytes are retained and discarded from the file Bytes and all subsequent information. If the file is less than Bytes, 0x00 bytes are added to the end of the file. The purpose of this is to standardize the length of all data stream files to facilitate subsequent image conversion and model processing.

[0013] (1-3) Pixel value mapping: Map each byte of data value to a grayscale pixel value, where 0x00 corresponds to black and 0xff corresponds to white. This step converts the data stream file from binary form to image form, so that it can be processed by deep learning models such as convolutional neural networks.

[0014] Step (2) Perform data enhancement on malicious traffic data: Malicious traffic samples generated using the variational autoencoder model. The variational autoencoder model consists of an encoder and a decoder. The encoder is used to map the malicious traffic data to the probability distribution of the latent space and sample from the latent space through the reparameterization technique; the decoder is used to reconstruct the malicious traffic data. The reconstructed malicious traffic data is the synthesized malicious traffic sample.

[0015] Step (3) Design a malicious traffic detection model: The malicious traffic detection model is constructed based on a Ghost module, an S2M module, a Mamba module, a global pooling layer and a linear layer, wherein each convolutional layer is connected to a batch normalization layer and a nonlinear activation layer.

[0016] The malicious traffic detection model consists of a convolution kernel size of The Ghost module consists of a Ghost module, multiple cross-designed S2M modules and S2MGhostMamba modules, a global pooling layer and a linear layer. The Ghost module first uses The convolution layer compresses the number of channels of the input image, and then obtains more feature maps through the depthwise separable convolution layer, splicing different feature maps together to form a new output;

[0017] The S2M module separates the token mixer and channel mixer based on the MobileNetV2 module, and simplifies the module structure through reparameterization techniques. The final S2M module is obtained. The specific steps are to use The depthwise separable convolutional layer integrates spatial information along the channel direction, and then uses a residual connection and two 1×1 convolutional layers to learn the relationship between different channel features; The S2MGhostMamba module consists of a convolution kernel of size A Ghost module with a convolution kernel size of 1×1, Mamba modules, a Ghost module with a convolution kernel size of 1×1, and a convolution kernel size of The Mamba module consists of a Ghost module. The Mamba module includes neural network layers such as RMS normalization layer, linear layer and selective state space model.

[0018] Step (4) Pre-train malicious traffic detection model: (4-1) Parameter initialization: Randomly initialize the weight parameters of the learning network and bias parameters , initialize the iteration round , set the initial learning rate , training sample batch size , maximum number of iterations ; (4-2) Data batching: according to the set sample batch size The dataset Divide evenly batches, and the malicious traffic data subset of each batch is represented as , and its corresponding label set is ; (4-3) Data input: Randomly select a subset of data from a batch The data is sent to the malicious traffic detection model, and the feature representation of the data is extracted through the Ghost module, S2M module, and Mamba module. The global pooling layer and linear layer are input to obtain the predicted label set of the batch data. ; (4-4) Contrastive learning feature enhancement: Calculate the contrast loss function value based on the feature representation extracted by the Ghost module, S2M module and Mamba module , in the feature space, the distance between samples of different categories is maximized and the distance between samples of the same category is minimized, a more discriminative feature representation is learned, and the enhanced feature representation is output.

[0019] (4-5) Parameter update: According to the true label of the batch data and the predicted label set Calculate the loss function value , and according to the loss function value Update model parameters; (4-6) Single round training: When the Round If all batches of data are input into the detection model, it means that the round of training is completed and enter step (4-7), otherwise return to step (4-3); (4-7) Training end judgment: When the loss function In continuous The reduction in wheel is less than ,in The minimum number of convergence rounds to determine whether the detection model has converged, For judgment If the threshold value is no longer reduced, it indicates that the detection model has converged, and step (4-9) is executed; otherwise, step (4-8) is executed; (4-8) If ,but , continue iterating and return to step (4-2); if , indicating that the detector training is completed and proceeding to step (4-9); (4-9) Model saving: Save the optimal weight parameters of the detector model and the optimal bias parameters .

[0020] Step (5) Detecting malicious traffic data that has never been seen before: (5-1) Preprocessing the malicious traffic data that has never been seen before, using the same method as step (1); (5-2) Input a convolution kernel size of The Ghost module performs preliminary feature extraction; the Ghost module first uses the ordinary The convolution layer compresses the number of channels of the input image, and then performs a depth-separable convolution layer to obtain more feature maps, and then splices different feature maps together to form a new output; (5-3) Use multiple S2M modules for deep local feature extraction; (5-4) Use S2MGhostMamba module interspersed between multiple S2M modules for further global feature extraction: The S2MGhostMamba module passes the feature map through a convolution kernel size of The Ghost module is used to perform local feature modeling, and then the number of channels is adjusted through a Ghost module with a convolution kernel size of 1×1 to obtain the output , represents the field of real numbers, and Represent the height and width of the effective receptive field respectively, represents the dimension; then Expand to non-overlapping one-dimensional vectors Among them, the effective receptive field patch is based on a fixed area Divide into small area, the number of one-dimensional vectors obtained by expansion is , the height of each patch and width That is, the height and width of each one-dimensional vector, , , is the size of the convolution kernel. Then use the Mamba module to perform convolution on each pixel in each patch. Encode and get A one-dimensional vector with global feature information .Will A one-dimensional vector with global feature information Collapse back to a high-dimensional vector .

[0021] The number of channels is adjusted back to the original size through a Ghost module with a convolution kernel size of 1×1 and mapped back to the low-dimensional feature space. Then, it is spliced ​​with the original input feature map along the channel direction through a residual connection, and finally, it is concatenated with the original input feature map through a convolution kernel size of The Ghost module performs feature fusion to obtain output features;

[0022] (5-5) Use the detector to output the final detection results: use a 1×1 Ghost module for feature compression, use a global pooling layer to adjust the size of the feature map to a fixed size, flatten the feature map into a one-dimensional vector, add a Dropout layer to prevent overfitting, and finally use a linear layer to output the detection results.

[0023] Compared with the prior art, the present invention has the following beneficial effects: 1. Based on the principle that model performance and data volume present a power-law distribution, the present invention uses a large amount of known malicious traffic data to pre-train the model, and uses a variational autoencoder to enhance the malicious traffic data, thereby achieving an emergence effect and showing good results on malicious traffic data that has never been seen before.

[0024] 2. The present invention learns highly discriminative feature representations in feature space through contrastive learning, and enhances the model's ability to identify malicious traffic data by maximizing the feature distance between samples of different categories and minimizing the feature distance between samples of the same category. Specifically, the contrast loss function is used to drive the model to learn more discriminative feature embedding, so that malicious traffic samples from different categories are easier to distinguish in the feature space, thereby improving the generalization performance of the model and the detection accuracy of new malicious traffic.

[0025] 3. The present invention converts the traffic data into grayscale data that is more convenient for classification and detection through operations such as flow data slicing, truncation, zero padding, and pixel value mapping, so that the potential features of the flow data can be better analyzed and extracted.

[0026] 4. The present invention takes into account the characteristics of malicious traffic data with less effective semantic information and more redundant information, uses the Mamba module with a selective state space model, and utilizes the mechanism of meta-learning to learn how to retain as much effective semantic information as possible during the feature extraction process and filter out useless information, thereby improving the effectiveness of the extracted features and further improving the accuracy of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 It is a schematic diagram of the overall framework of malicious traffic detection of the present invention; Figure 2 It is a schematic diagram of the data preprocessing step of the present invention; Figure 3is a schematic diagram of a malicious traffic detection model in the method of the present invention; Figure 4 is a schematic diagram of the Ghost module in the method of the present invention; Figure 5 is a schematic diagram of the improved MobileNetV2 module (S2M module) in the method of the present invention; Figure 6 It is a schematic diagram of the S2MGhostMamba module in the method of the present invention; Figure 7 It is a schematic diagram of the Mamba module in the method of the present invention. DETAILED DESCRIPTION

[0028] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments.

[0029] In this embodiment, data from 20 malicious traffic data sets and normal traffic data sets are randomly selected from the CIC-AndMal2017 and USTC-TFC2016 data sets as pre-training sets, and four types of malicious traffic data that have never been seen, namely plankton, selfmite, shuanet and zeus, are selected as test sets. The zero-sample malicious traffic detection method based on convolutional neural network and selective state space model proposed in the present invention is described. The overall framework is as follows: Figure 1 shown.

[0030] Step (1) Obtain zero-sample malicious traffic types (including 20 malicious traffic types, including beanbot, biige, dowgin, ewind, fakeinst, fakemart, fakenotify, feiwo, gooligan, jifake, kemoge, koodous, mazarbot, miuref, mobidash, nandrobox, smssniffer, tinba, youmi and zsone) from the CIC-AndMal2017 and USTC-TFC2016 datasets and traffic data of normal traffic types for model pre-training, and convert the original traffic data into grayscale data that is more convenient for classification through slicing, truncation, zero padding and pixel value mapping operations. The specific data conversion is as follows: Figure 2 As shown:

[0031] 1. Data slicing: Split the original PCAP file into data streams. Group the data packets in the PCAP file according to the five-tuple information (source IP address, destination IP address, source port, destination port, and protocol) and arrange them in chronological order.

[0032] 2. Data uniform length processing: The first part of each data stream file in the CIC-AndMal2017 and USTC-TFC2016 datasets is (In this example ) bytes are retained and discarded Bytes and beyond. For insufficient length Bytes file, followed by 0x00 bytes to achieve a uniform length, which is convenient for subsequent image conversion and model processing.

[0033] 3. Pixel value mapping: Map each byte data value to a grayscale pixel value, where 0x00 corresponds to black (0) and 0xff corresponds to white (255) (in this example, the final output image data format is PNG).

[0034] Step (2) Perform data enhancement on malicious traffic data: A large number of malicious traffic samples are generated using a variational autoencoder model. The variational autoencoder model consists of an encoder and a decoder. The encoder is used to map the malicious traffic data to the probability distribution of the latent space (in this example, the probability distribution is a Gaussian distribution) and sample from the latent space through the reparameterization technique; the decoder is used to reconstruct the malicious traffic data. The reconstructed malicious traffic data is the synthesized malicious traffic sample.

[0035] Step (3) Design a malicious traffic detection model: like Figure 3 As shown in Figure 1, the model is built based on Ghost module, S2M module, Mamba module, global pooling layer and linear layer, where the convolution layer of each Ghost module is connected to batch normalization layer and nonlinear activation layer. S2M module and Mamba module are used to extract local features and global features respectively.

[0036] Specifically, the model consists of a convolution kernel size of (In this example ) Ghost module, multiple cross-designed S2M modules and S2MGhostMamba modules, a global pooling layer and a linear layer. Among them, the Ghost module consists of a (In this example ) convolution layer and a point convolution layer, the final Ghost module is as follows Figure 4 As shown in Figure 2; the S2M module separates the token mixer and channel mixer based on the MobileNetV2 module, and simplifies the structure of the module through the reparameterization technique. The final S2M module is as follows: Figure 5 As shown, the specific steps include using (In this example ) integrates spatial information along the channel direction, and then uses a residual connection and two 1×1 convolutional layers to learn the relationship between different channel features. Figure 6 As shown, a convolution kernel size is The Ghost module (in this example ), a Ghost module with a convolution kernel size of 1×1, Mamba modules (in this example ), a Ghost module with a convolution kernel size of 1×1 and a convolution kernel size of Ghost module (in this example ). Among them, Mamba modules are as follows Figure 7 As shown, it specifically includes neural network layers such as root mean square normalization layer, linear layer and selective state space model.

[0037] Step (4) pre-trains a malicious traffic detection model, using the malicious traffic data obtained in step (1) to train the malicious traffic detection model. The details are as follows:

[0038] (4-1) Parameter initialization: Randomly initialize the weight parameters and bias parameters in the model. Set the initial learning rate to (In this example ), the training sample batch size is (In this example ), the maximum number of iterations is (In this example ), that is, in each iteration, the model will process 64 samples until 200 iterations are completed.

[0039] (4-2) Data batching: divide the preprocessed data set into batches Each batch contains 64 malicious traffic data subsets and their corresponding labels. In this example, 1000 samples are selected. (1000 / 64), each batch contains 64 samples, and the last batch may contain fewer samples.

[0040] (4-3) Data input: In each round of training, a batch of data subsets is randomly selected In this example, the 64 samples are fed into the model one by one. The model extracts the feature representation of each sample through the Ghost module, S2M module and Mamba module, and inputs the global pooling layer and linear layer to obtain the predicted label set of the batch data. .

[0041] (4-4) Contrastive learning feature enhancement: Calculate the contrast loss function value based on the feature representation extracted by the Ghost module, S2M module and Mamba module , in the feature space, the distance between samples of different categories is maximized and the distance between samples of the same category is minimized, a more discriminative feature representation is learned, and the enhanced feature representation is output.

[0042] (4-5) Parameter update: According to the true label of the batch data (In this embodiment ) and the label set predicted by the model (In this embodiment ) Calculate the loss function value , using a gradient descent-based optimizer (in this example, an AdamW optimizer with a decoupled weight decay regularization method is used) to update the weight parameters and bias parameters .

[0043] (4-6) Single round training: When the (In this example ) If all batches of data are input into the model, it means that the round of training is completed and enters step ((4-7), otherwise it returns to step (4-3).

[0044] (4-7) Training end judgment: When the loss function In continuous The reduction in wheel is less than ,in (In this embodiment ) is the minimum number of convergence rounds to determine whether the detection model is converged. (In this embodiment ) for judgment A threshold value that basically no longer decreases indicates that the detection model has converged, and step (4-9) is executed; otherwise, step (4-8) is executed.

[0045] (4-8) If ,but , continue iterating and return to step (3-2); if , the detector training is completed and enters step (3-8).

[0046] (4-9) Model saving: Save the optimal weight parameters of the detector model and the optimal bias parameters .

[0047] Step (5) Identify the types of malicious traffic data: The malicious traffic data is preprocessed in the same way as step (1), and then input into the model for feature extraction and detection. The details are as follows:

[0048] (5-1) Preprocess the malicious traffic data through slicing, truncation, zero padding, and pixel value mapping operations to obtain grayscale image data. The specific method is the same as step (1).

[0049] (5-2) Input the grayscale image data into a (In this example )’s Ghost module (in this example, the Ghost module consists of a convolutional layer with a convolution kernel size of 3×3 and a 3×3 depth-separable convolutional layer) for preliminary feature extraction.

[0050] (5-3) Use multiple S2M modules for deep local feature extraction.

[0051] (5-4) Multiple S2MGhostMamba modules are interspersed between multiple S2M modules (in this example, the number and types of modules are 5 S2M modules, 2 S2MGhostMamba modules, 1 S2M module, 4 S2MGhostMamba modules, 1 S2M module, and 3 S2MGhostMamba modules, respectively) to perform further global feature extraction: The specific steps of the S2MGhostMamba module include passing the feature map through a convolution kernel size of Ghost module (in this example ) to perform local feature modeling, and then adjust the number of channels through a Ghost module with a convolution kernel size of 1×1 to obtain the output Then Expand to non-overlapping one-dimensional vectors .in, (In this example )and is the number of one-dimensional vectors, (In this example , )and (In this example , ) are the width and height of each one-dimensional vector. Then use Mamba modules (in this example ) for each Encode and get Then, a Ghost module with a convolution kernel size of 1×1 is used to adjust the number of channels back to the original size and map it back to the low-dimensional feature space. Then, a residual connection is used to splice the original input feature map along the channel direction, and finally, a convolution kernel size of Ghost module (in this example ) to perform feature fusion to obtain output features.

[0052] (5-5) Use the detector to output the final detection results: The specific steps of the detector include using a 1×1 Ghost module to compress features, a global pooling layer to adjust the size of the feature map to a fixed size, flattening the feature map into a one-dimensional vector, adding a Dropout layer to prevent overfitting, and finally outputting the detection results through a linear layer.

[0053] The contents described in the above examples are merely an enumeration of implementation forms of the present invention. The protection scope of the present invention should not be limited to the specific forms described in the embodiments. The protection scope of the present invention should also include similar inventive methods conceived on the basis of the present invention.

Claims

1. A malicious traffic detection method based on a selective state space model, characterized by: Step (1) preparing training data for the malicious traffic classification model: downloading known malicious traffic data from the Internet as raw traffic data; converting the raw traffic data into grayscale image data after processing for pre-training of the malicious traffic detection model; Step (2) Perform data enhancement on malicious traffic data: Malicious traffic samples generated using a variational autoencoder model; The variational autoencoder model consists of an encoder and a decoder; the encoder is used to map malicious traffic data into the probability distribution of the latent space and sample from the latent space through the reparameterization technique; The decoder is used to reconstruct malicious traffic data; the reconstructed malicious traffic data is the synthesized malicious traffic sample; Step (3) Design a malicious traffic detection model: The malicious traffic detection model is constructed based on the Ghost module, the S2M module, the Mamba module, the global pooling layer and the linear layer, wherein each convolutional layer is connected to a batch normalization layer and a nonlinear activation layer; Step (4) Pre-train malicious traffic detection model: (4-1) Parameter initialization: Randomly initialize the weight parameters of the learning network and bias parameters , initialize the iteration round , set the initial learning rate , training sample batch size , maximum number of iterations ; (4-2) Data batching: according to the set sample batch size The dataset Divide evenly batches, and the malicious traffic data subset of each batch is represented as , and its corresponding label set is ; (4-3) Data input: Randomly select a subset of data from a batch The data is sent to the malicious traffic detection model, and the feature representation of the data is extracted through the Ghost module, S2M module, and Mamba module. The global pooling layer and linear layer are input to obtain the predicted label set of the batch data. ; (4-4) Contrastive learning feature enhancement: Calculate the contrast loss function value based on the feature representation extracted by the Ghost module, S2M module and Mamba module , maximize the distance between samples of different categories and minimize the distance between samples of the same category in the feature space, learn more discriminative feature representations, and output enhanced feature representations; (4-5) Parameter update: According to the true label of the batch data and the predicted label set Calculate the loss function value , and according to the loss function value Update model parameters; (4-6) Single round training: When the Round If all batches of data are input into the detection model, it means that the round of training is completed and enter step (4-7), otherwise return to step (4-3); (4-7) Training end judgment: When the loss function In continuous The reduction in wheel is less than ,in The minimum number of convergence rounds to determine whether the detection model has converged, For judgment If the threshold value is no longer reduced, it indicates that the detection model has converged, and step (4-9) is executed; otherwise, step (4-8) is executed; (4-8) If ,but , continue iterating and return to step (4-2); if , indicating that the detector training is completed and proceeding to step (4-9); (4-9) Model saving: Save the optimal weight parameters of the detector model and the optimal bias parameters ; Step (5) Detecting malicious traffic data that has never been seen before: (5-1) Preprocessing malicious traffic data that has never been seen before; (5-2) Input a convolution kernel size of The Ghost module performs preliminary feature extraction; the Ghost module first uses the ordinary The convolution layer compresses the number of channels of the input image, and then performs a depth-separable convolution layer to obtain more feature maps, and then splices different feature maps together to form a new output; (5-3) Use multiple S2M modules for deep local feature extraction; (5-4) Use S2MGhostMamba module interspersed between multiple S2M modules for further global feature extraction: (5-5) Use the detector to output the final detection results: use a 1×1 Ghost module for feature compression, use a global pooling layer to adjust the size of the feature map to a fixed size, flatten the feature map into a one-dimensional vector, add a Dropout layer to prevent overfitting, and finally use a linear layer to output the detection results.

2. The method for detecting malicious traffic based on a selective state space model as claimed in claim 1, characterized in that: The same method is used in step (1) of processing the original traffic data and step (5-1) of preprocessing the malicious traffic data that has never been seen, including: (1-1) Data slicing: The original PCAP file is split into data streams. The data packets in the PCAP file are grouped according to five-tuple information and arranged in chronological order. The five-tuple information is source IP address, destination IP address, source port, destination port and protocol. (1-2) Data uniform length processing: the first part of each file is bytes are retained and discarded from the file Bytes and all subsequent information; if the file is less than Bytes, add 0x00 bytes to the end of the file; (1-3) Pixel value mapping: Map the data value of each byte to a grayscale pixel value, where 0x00 corresponds to black and 0xff corresponds to white.

3. The method for detecting malicious traffic based on a selective state space model as claimed in claim 1, characterized in that: The malicious traffic detection model described above consists of a convolution kernel size of The network consists of a Ghost module, multiple cross-designed S2M modules and S2MGhostMamba modules, a global pooling layer and a linear layer; among them: The Ghost module consists of a The Ghost module first uses The convolution layer compresses the number of channels of the input image, and then obtains more feature maps through the depthwise separable convolution layer, splicing different feature maps together to form a new output; The S2M module described above separates the token mixer and channel mixer based on the MobileNetV2 module, and simplifies the module structure through reparameterization techniques to obtain the final S2M module. The specific steps are to use The depthwise separable convolutional layer integrates spatial information along the channel direction, and then uses a residual connection and two 1×1 convolutional layers to learn the relationship between different channel features; The S2MGhostMamba module consists of a convolution kernel size of A Ghost module with a convolution kernel size of 1×1, Mamba modules, a Ghost module with a convolution kernel size of 1×1, and a convolution kernel size of The Ghost module consists of the Mamba module, which includes neural network layers such as RMS normalization layer, linear layer and selective state space model.

4. The method for detecting malicious traffic based on a selective state space model as claimed in claim 1, characterized in that: Step (5-3) is as follows: The S2MGhostMamba module passes the feature map through a convolution kernel size of The Ghost module is used to perform local feature modeling, and then the number of channels is adjusted through a Ghost module with a convolution kernel size of 1×1 to obtain the output , represents the field of real numbers, and Represent the height and width of the effective receptive field respectively, represents the dimension; then Expand to non-overlapping one-dimensional vectors ; Among them, the effective receptive field patch is based on a fixed area Divide into small area, the number of one-dimensional vectors obtained by expansion is , the height of each patch and width That is, the height and width of each one-dimensional vector, , , is the size of the convolution kernel; then the Mamba module is used for each pixel in each patch Encode and get A one-dimensional vector with global feature information ;Will A one-dimensional vector with global feature information Collapse back to a high-dimensional vector ; The number of channels is adjusted back to the original size through a Ghost module with a convolution kernel size of 1×1 and mapped back to the low-dimensional feature space; then it is spliced ​​with the original input feature map along the channel direction through a residual connection, and finally a convolution kernel size of The Ghost module is used to perform feature fusion to obtain output features.

Citation Information

Patent Citations

  • Malicious encrypted traffic identification method based on multi-granularity feature extraction

    CN117978530A

  • Malicious encrypted traffic classification method and terminal in network information system scene

    CN118590252A

  • Encrypted malicious traffic detection method based on traffic interaction behavior and attention mechanism

    CN118827211A

  • Smearing type image segmentation method and device, electronic equipment and medium

    CN116071370A

  • Malicious traffic identification method and system based on data enhancement and feature fusion

    CN116318928A

Cited By

  • Malicious encrypted traffic detection method based on diffusion model

    CN121098548A

  • A malicious encrypted traffic detection method based on diffusion model

    CN121098548B

  • Malicious traffic category detection method and related equipment

    CN122247711A

  • An encrypted traffic analysis method based on protocol-aware state space and end-cloud collaboration

    CN122660920A