Network traffic data classification method and device, computer equipment and medium
By adopting the Mamba block-based autoencoder and multimodal learning method in network traffic classification, the problems of low accuracy and low computing efficiency in the prior art are solved, and high accuracy and high efficiency network traffic classification are achieved.
Patent Information
- Application Number
- CN202510073593.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art has problems of low accuracy and low computing efficiency in network traffic classification, especially when facing encrypted traffic and dynamically changing networks, it is difficult to adapt to modern network environments.
The autoencoder based on Mamba block and a bidirectional recurrent neural network are used, combined with the multimodal learning method, and multimodal representation of network traffic data is extracted, and the training is carried out through the final full connection layer to generate the classification results of network traffic data.
It improves the accuracy and computing efficiency of network traffic classification, and can effectively detect new attack traffic in complex network environments, with a classification accuracy of 97.24%.
Smart Images

Figure CN120017596A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer network technology, and in particular to a method, device, computer equipment and medium for classifying network flow data. Background Art
[0002] Network traffic classification is widely used in areas such as network management and threat detection.
[0003] With the increasing complexity of network traffic types, traditional traffic classification methods based on rules and manual feature extraction have become difficult to adapt to modern network environments, especially when facing encrypted traffic and dynamically changing networks, the accuracy and robustness have dropped significantly. Although deep learning has been widely used in traffic classification, existing methods, such as Transformer, have high computational complexity and low efficiency when processing large-scale data. In addition, most of the existing methods are single-modal models, which do not make full use of traffic features. As a pre-trained model based on the SSM structure, the Mamba model can improve classification efficiency while maintaining high accuracy through low computational complexity. Combining the Mamba model with multimodal learning not only effectively solves the shortcomings of existing methods in terms of accuracy and computational efficiency, but also improves classification performance, especially in complex network environments and the detection of new attack traffic.
[0004] Currently, Mamba block-based models and multimodal learning models are rarely used in traffic classification, and there are few methods that combine the two. Summary of the invention
[0005] In view of this, an embodiment of the present invention provides a method for classifying network traffic data to solve the technical problems of low accuracy and low computational efficiency in the classification processing of network traffic in the prior art. The method includes:
[0006] Split network traffic data into bidirectional streams, extract original bytes from the bidirectional streams, and extract packet length sequences from the bidirectional streams;
[0007] Construct an autoencoder based on Mamba blocks, use the autoencoder to perform unsupervised pre-training on the original bytes, generate a pre-trained autoencoder, use the pre-trained autoencoder to encode the original bytes, generate a first representation corresponding to the original bytes, use a bidirectional recurrent neural network to learn the packet length sequence, and generate a second representation corresponding to the packet length sequence;
[0008] After cascading the first representation and the second representation, a multimodal representation is generated, and the multimodal representation is input into the final fully connected layer for training to generate the classification result of the network traffic data.
[0009] The embodiment of the present invention also provides a network traffic data classification device to solve the technical problems of low accuracy and low computational efficiency in the classification processing of network traffic in the prior art. The device includes:
[0010] A data conversion module is used to segment network traffic data into bidirectional streams, extract original bytes in the bidirectional streams, and extract packet length sequences in the bidirectional streams;
[0011] A representation generation module is used to construct an autoencoder based on the Mamba block, use the autoencoder to perform unsupervised pre-training on the original bytes, generate a pre-trained autoencoder, use the pre-trained autoencoder to encode the original bytes, generate a first representation corresponding to the original bytes, use a bidirectional recurrent neural network to learn the packet length sequence, and generate a second representation corresponding to the packet length sequence;
[0012] The data classification module is used to generate a multimodal representation by cascading the first representation and the second representation, input the multimodal representation into the final fully connected layer for training, and generate a classification result of the network traffic data.
[0013] An embodiment of the present invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements any of the above-mentioned methods for classifying network traffic data when executing the computer program, so as to solve the technical problems of low accuracy and low computational efficiency in the classification processing of network traffic in the prior art.
[0014] An embodiment of the present invention also provides a computer-readable storage medium, which stores a computer program for executing any of the above-mentioned network traffic data classification methods to solve the technical problems of low accuracy and low computational efficiency in the classification processing of network traffic in the prior art.
[0015] Compared with the prior art, the at least one technical solution adopted in the embodiments of this specification can achieve the following beneficial effects:
[0016] To solve the problems of high computational complexity and insufficient feature utilization of deep learning methods in the field of traffic classification, the Mamba model and multimodal methods are combined to improve computational efficiency while fully utilizing features to improve classification accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0018] Figure 1 is a flow chart of a method for classifying network traffic data provided by an embodiment of the present invention;
[0019] Figure 2 It is an overall framework diagram of a method for implementing the above-mentioned network traffic data classification provided by an embodiment of the present invention;
[0020] Figure 3 is a structural diagram of a Mamba block provided by an embodiment of the present invention;
[0021] Figure 4 is a structural block diagram of a computer device provided by an embodiment of the present invention;
[0022] Figure 5 It is a structural block diagram of a network traffic data classification device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0023] The embodiments of the present application are described in detail below with reference to the accompanying drawings.
[0024] The following describes the implementation methods of the present application through specific specific examples, and those skilled in the art can easily understand other advantages and effects of the present application from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The present application can also be implemented or applied through other different specific implementation methods, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that, in the absence of conflict, the following embodiments and the features in the embodiments can be combined with each other. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without making creative work belong to the scope of protection of the present application.
[0025] In an embodiment of the present invention, a method for classifying network traffic data is provided, such as Figure 1 and Figure 2 As shown, the method includes:
[0026] Step S101: dividing the network traffic data into bidirectional streams, extracting the original bytes in the bidirectional streams, and extracting the packet length sequence in the bidirectional streams;
[0027] Step S102: constructing an autoencoder based on the Mamba block, performing unsupervised pre-training on the original bytes using the autoencoder to generate a pre-trained autoencoder, encoding the original bytes using the pre-trained autoencoder to generate a first representation corresponding to the original bytes, and using a bidirectional recurrent neural network to learn the packet length sequence to generate a second representation corresponding to the packet length sequence;
[0028] Step S103: After cascading the first representation and the second representation, a multimodal representation is generated, and the multimodal representation is input into the final fully connected layer for training to generate a classification result of the network traffic data.
[0029] In the specific implementation, in order to remove the background traffic and convert it into a bidirectional flow for classification, the following steps are performed to segment the network traffic data into a bidirectional flow:
[0030] After removing the background traffic from the network traffic data, the IP address and port corresponding to the source and the IP address and port corresponding to the destination are separated from the network traffic data; the IP address, port, protocol of the source, IP address and port of the destination are combined to generate a bidirectional flow, wherein the bidirectional flow is data that can flow simultaneously or alternately in two directions.
[0031] Specifically, the network traffic data is segmented into bidirectional flows using the IP address, protocol, and port triples, while the background traffic is eliminated.
[0032] In the specific implementation, in order to obtain a mode in the bidirectional stream, the following steps are performed to extract the original bytes in the bidirectional stream:
[0033] For the first K packets of each bidirectional flow, perform the following operations until all K packets are processed: extract M bytes from the packet header, and pad the end of the packet header with zeros if the length of the packet header is less than M bytes; extract N bytes from the packet payload, and pad the end of the payload with zeros if the length of the payload is less than N bytes; arrange the K×(M+N) bytes to generate an n×n matrix as the original bytes of the bidirectional flow, where The values of K, M, and N are positively correlated with the accuracy.
[0034] In specific implementation, in order to obtain another mode in the bidirectional stream, the following steps are performed to extract the packet length sequence in the bidirectional stream:
[0035] Extract the packet length sequence of each packet from the first m packets of each bidirectional stream.
[0036] Specifically, extract the original bytes and packet length sequence from the bidirectional stream obtained in the previous step. Specifically, for the first K packets of the bidirectional stream, extract M bytes from the packet header and N bytes from the payload. Packets with insufficient length are padded with zeros. Then, arrange these K(M+N) bytes into an n×n matrix, where In addition, it is necessary to extract the packet length sequence from the first m packets of each bidirectional flow. The values of K, M, and N depend on the requirements. The larger the value, the more data there is. Improving the accuracy requires better hardware processing level.
[0037] In specific implementation, the following steps are used to build an autoencoder based on the Mamba block:
[0038] The Mamba block is used as an encoder to learn the representation of the original bytes; the Mamba block is used as a decoder to convert the feature vector output by the encoder into a target sequence; the encoder and the decoder are connected to generate an autoencoder based on the Mamba block.
[0039] In specific implementation, the following steps are used to perform unsupervised pre-training of the original bytes using the autoencoder to generate a pre-trained autoencoder:
[0040] The token is used as a mask to randomly replace some bytes in the original bytes, and the bytes generated after the replacement are used as the training set of the autoencoder; the training set is used to perform unsupervised training on the autoencoder to generate a pre-trained autoencoder.
[0041] Specifically, the autoencoder composed of Mamba blocks is used to perform unsupervised pre-training on the raw bytes. Figure 3 As shown in the figure, the Mamba blocks are used as encoders and decoders to form an autoencoder, and the unsupervised pre-training method is used for training. During the training process, a certain proportion of the original bytes are randomly masked. In this process, the "mask" token is used to replace some of the original bytes, and the model tries to restore them using adjacent bytes. This method allows the use of a large amount of unlabeled data for pre-training, ensuring the encoding ability of the model while reducing the need for labeled data. The encoder composed of Mamba blocks will be used for subsequent byte representation learning.
[0042] In specific implementation, the following steps are used to calculate the overall loss value of the model, and the minimum loss value is used as the optimization goal of the model:
[0043] After generating the first representation corresponding to the original byte, the first representation is input into the first fully connected layer for training and the loss value during training is obtained through the first Softmax loss function as the first loss value Loss1; after generating the second representation corresponding to the packet length sequence, the second representation is input into the second fully connected layer for training and the loss value during training is obtained through the second Softmax loss function as the second loss value Loss2; when the multimodal representation is input into the final fully connected layer for training, the loss value during training is obtained through the third Softmax loss function as the third loss value Loss3; the total loss value Loss is calculated. all =Loss1+Loss2+Loss3, the total loss value is used as the overall loss value of the classification model of the network traffic data; the overall loss value is minimized by adjusting the parameters of the classification model of the network traffic data.
[0044] Specifically, after pre-training, we can get ByteMamba (pre-trained autoencoder), which is a Mamba layer that can encode network stream bytes. It will be used for the representation learning of the original bytes in the network stream in the classification task and as the first modality of multimodal learning. Specifically, the original bytes are encoded using the pre-trained encoder to obtain the corresponding representation m1 (first representation). Then, m1 is input into a fully connected layer and trained using a supervised method with a cross entropy loss function (first Softmax loss function), and the loss of this part is recorded as Loss1.
[0045] The bidirectional LSTM network is used to learn the packet length sequence of the bidirectional flow and obtain its representation m2 (the second representation). After that, m2 is also input into a fully connected layer and trained using a supervised method with a cross entropy loss function (the second Softmax loss function). The loss of this part is recorded as Loss2
[0046] The representations m1 (first representation) and m2 (second representation) are cascaded to obtain a multimodal representation m about the bidirectional flow, which is input into a fully connected layer for training. The loss function is the cross entropy loss. This part of the loss is recorded as Loss3, and then input into the Softmax layer to output the classification result.
[0047] The overall loss function of the model is Loss all =Loss1+Loss2+Loss3, the optimization goal of model training is to minimize the overall loss function.
[0048] This method can be used to combine the original bytes of network traffic and the packet length sequence features to complement each other to improve the characterization ability of network traffic and thus improve the classification accuracy. On a data set containing 23 types of traffic, the classification accuracy can reach 97.24%.
[0049] In this embodiment, a computer device is provided, such as Figure 4 As shown, it includes a memory 401, a processor 402, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, any of the above-mentioned network traffic data classification methods is implemented.
[0050] Specifically, the computer device may be a computer terminal, a server or a similar computing device.
[0051] In this embodiment, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program for executing any of the above-mentioned network traffic data classification methods.
[0052] Specifically, computer-readable storage media include permanent and non-permanent, removable and non-removable media, and information storage can be achieved by any method or technology. Information can be computer-readable instructions, data structures, modules of programs or other data. Examples of computer-readable storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable storage media does not include temporary computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0053] Based on the same inventive concept, an embodiment of the present invention also provides a classification device for network traffic data, as described in the following embodiments. Since the principle of solving the problem by the classification device for network traffic data is similar to that of the classification method for network traffic data, the implementation of the classification device for network traffic data can refer to the implementation of the classification method for network traffic data, and the repeated parts will not be repeated. As used below, the term "unit" or "module" can be a combination of software and / or hardware that implements predetermined functions. Although the device described in the following embodiments is preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.
[0054] Figure 5 is a structural block diagram of a network traffic data classification device according to an embodiment of the present invention. Figure 5 As shown, it includes: a data conversion module 501, a representation generation module 502 and a data classification module 503, and the structure is described below.
[0055] The data conversion module 501 is used to divide the network traffic data into bidirectional streams, extract the original bytes in the bidirectional streams, and extract the packet length sequence in the bidirectional streams;
[0056] A representation generation module 502 is used to construct an autoencoder based on a Mamba block, perform unsupervised pre-training on the original bytes using the autoencoder to generate a pre-trained autoencoder, encode the original bytes using the pre-trained autoencoder to generate a first representation corresponding to the original bytes, and use a bidirectional recurrent neural network to learn the packet length sequence to generate a second representation corresponding to the packet length sequence;
[0057] The data classification module 503 is used to generate a multimodal representation by cascading the first representation and the second representation, and input the multimodal representation to the final fully connected layer for training to generate a classification result of the network traffic data.
[0058] In one embodiment, the data conversion module includes:
[0059] A data extraction unit is used to remove background traffic from the network traffic data, and then segment the network traffic data to obtain the IP address and port corresponding to the source end and the IP address and port corresponding to the destination end;
[0060] The bidirectional flow generation unit is used to combine the source IP address, the source port, the protocol, the destination IP address and the destination port to generate a bidirectional flow, wherein the bidirectional flow is data that can flow simultaneously or alternately in two directions.
[0061] In one embodiment, the data conversion module further includes:
[0062] The loop unit is used to perform the following operations on the first K packets of each bidirectional flow until all K packets are processed:
[0063] A packet header extraction unit, used for extracting M bytes from the packet header of the packet, and padding the end of the packet header with zeros if the length of the packet header is less than M bytes;
[0064] A payload extraction unit, for extracting N bytes from a payload of the packet, and padding the payload with zeros if the length of the payload is less than N bytes;
[0065] The original byte extraction unit is used to arrange the K×(M+N) bytes to generate an n×n matrix as the original bytes of the bidirectional flow, where: The values of K, M, and N are positively correlated with the accuracy.
[0066] In one embodiment, the data conversion module further includes:
[0067] The packet length sequence extraction unit is used to extract the packet length sequence of each packet from the first m packets of each bidirectional flow.
[0068] In one embodiment, the representation generation module includes:
[0069] The encoder generation unit is used to use the Mamba block as an encoder, and the encoder is used to learn the representation of the raw bytes;
[0070] A decoder generation unit, used to use the Mamba block as a decoder, and the decoder is used to convert the feature vector output by the encoder into a target sequence;
[0071] The autoencoder generation unit is used to connect the encoder and the decoder to generate an autoencoder based on the Mamba block.
[0072] In one embodiment, the representation generation module further includes:
[0073] A training set generation unit is used to randomly replace some bytes in the original bytes by using the token as a mask, and use the bytes generated after the replacement as the training set of the autoencoder;
[0074] The encoder training unit is used to perform unsupervised training on the autoencoder using the training set to generate a pre-trained autoencoder.
[0075] In one embodiment, the above device also includes a loss value calculation module.
[0076] In one embodiment, the loss value calculation module includes:
[0077] A first loss value calculation unit, configured to generate a first representation corresponding to the original byte, input the first representation into a first fully connected layer for training, and obtain a loss value during training through a first Softmax loss function as a first loss value Loss1;
[0078] A second loss value calculation unit is used to generate a second representation corresponding to the packet length sequence, input the second representation into the second fully connected layer training and obtain the loss value during training through a second Softmax loss function as a second loss value Loss2;
[0079] A third loss value calculation unit, used for obtaining a loss value during training by using a third Softmax loss function when inputting the multimodal representation into the final fully connected layer for training as a third loss value Loss3;
[0080] Total loss value calculation unit, used to calculate the total loss value Loss all =Loss1+Loss2+Loss3, the total loss value is used as the overall loss value of the classification model of network traffic data;
[0081] The optimization target determination unit is used to minimize the overall loss value by adjusting the parameters of the classification model of the network traffic data.
[0082] The embodiments of the present invention achieve the following technical effects:
[0083] The multimodal network traffic classification method based on Mamba pre-training in the embodiment of the present invention is used for the classification task of network traffic, and can be used for intrusion detection and defense, network traffic management and optimization, etc. It mainly uses the autoencoder and bidirectional LSTM network composed of Mamba blocks as sub-models, respectively processes the original bytes and packet length sequences of the traffic, and completes the traffic classification task in a multimodal learning manner; solves the problems of high computational complexity and insufficient feature utilization of deep learning methods in the field of traffic classification, combines the Mamba model with the multimodal method, improves computational efficiency, and fully utilizes features to improve classification accuracy.
[0084] Obviously, those skilled in the art should understand that the modules or steps of the above-mentioned embodiments of the present invention can be implemented by a general computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices, and optionally, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, and in some cases, the steps shown or described can be executed in a different order from that here, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. In this way, the embodiments of the present invention are not limited to any specific combination of hardware and software.
[0085] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the embodiments of the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for classifying network traffic data, characterized in that: include: Segmenting network traffic data to generate bidirectional streams, extracting original bytes from the bidirectional streams, and extracting packet length sequences from the bidirectional streams; Constructing an autoencoder based on Mamba blocks, performing unsupervised pre-training on the original bytes using the autoencoder to generate a pre-trained autoencoder, encoding the original bytes using the pre-trained autoencoder to generate a first representation corresponding to the original bytes, and using a bidirectional recurrent neural network to learn the packet length sequence to generate a second representation corresponding to the packet length sequence; After cascading the first representation and the second representation, a multimodal representation is generated, and the multimodal representation is input into a final fully connected layer for training to generate a classification result of the network traffic data.
2. The method for classifying network traffic data according to claim 1, characterized in that: Split network traffic data into bidirectional streams, including: After removing the background traffic in the network traffic data, segmenting the network traffic data to obtain the IP address and port corresponding to the source end and the IP address and port corresponding to the destination end; The IP address of the source, the port of the source, the protocol, the IP address of the destination, and the port of the destination are combined to generate a bidirectional flow, wherein the bidirectional flow is data that can flow simultaneously or alternately in two directions.
3. The method for classifying network traffic data according to claim 1, characterized in that: Extracting the raw bytes in the bidirectional stream includes: For the first K packets of each bidirectional flow, perform the following operations until all K packets are processed: Extracting M bytes from the header of the packet, and padding the end of the header with zeros if the length of the header is less than M bytes; Extracting N bytes from the payload of the packet, and padding the end of the payload with zeros if the length of the payload is less than N bytes; The n×n matrix generated by arranging K×(M+N) bytes is used as the original bytes of the bidirectional stream, where: The values of K, M, and N are positively correlated with the accuracy.
4. The method for classifying network traffic data according to claim 1, characterized in that: Extracting a packet length sequence in the bidirectional stream, comprising: A packet length sequence of each packet is extracted from first m packets of each of the bidirectional streams.
5. The method for classifying network traffic data according to claim 1, characterized in that: Build a Mamba block-based autoencoder, including: Using a Mamba block as an encoder, the encoder is used to learn the representation of the raw bytes; Using the Mamba block as a decoder, the decoder is used to convert the feature vector output by the encoder into a target sequence; The encoder is connected to the decoder to generate a Mamba block based autoencoder.
6. The method for classifying network traffic data according to claim 1, characterized in that: Using the autoencoder to perform unsupervised pre-training on the original bytes to generate a pre-trained autoencoder, comprising: Using token as a mask to randomly replace some bytes in the original bytes, and using the bytes generated after the replacement as a training set of the autoencoder; The autoencoder is unsupervisedly trained using the training set to generate a pre-trained autoencoder.
7. The method for classifying network traffic data according to any one of claims 1 to 6, characterized in that: Also includes: After generating the first representation corresponding to the original byte, input the first representation into the first fully connected layer training and obtain the loss value during training through the first Softmax loss function as the first loss value Loss1; After generating a second representation corresponding to the packet length sequence, inputting the second representation into a second fully connected layer for training and obtaining a loss value during training through a second Softmax loss function as a second loss value Loss2; When the multimodal representation is input into the final fully connected layer for training, a loss value during training is obtained through a third Softmax loss function as a third loss value Loss3; Calculate the total loss value Loss all =Loss1+Loss2+Loss3, the total loss value is used as the overall loss value of the classification model of network traffic data; The overall loss value is minimized by adjusting the parameters of the classification model of the network traffic data.
8. A network traffic data classification device, characterized in that: include: A data conversion module, used for dividing the network traffic data into bidirectional streams, extracting the original bytes in the bidirectional streams, and extracting the packet length sequence in the bidirectional streams; A representation generation module, used to construct an autoencoder based on a Mamba block, use the autoencoder to perform unsupervised pre-training on the original bytes to generate a pre-trained autoencoder, use the pre-trained autoencoder to encode the original bytes to generate a first representation corresponding to the original bytes, use a bidirectional recurrent neural network to learn the packet length sequence, and generate a second representation corresponding to the packet length sequence; A data classification module is used to generate a multimodal representation by cascading the first representation and the second representation, input the multimodal representation into a final fully connected layer for training, and generate a classification result of the network traffic data.
9. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method for classifying network traffic data according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program for executing the network traffic data classification method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Encrypted traffic identification method and system, terminal and storage medium
CN114448905A
Personnel intrusion identification method and system based on Mamba model and rule engine
CN118644816A
Cited By
Network traffic classification method, system and device, and storage medium
CN120915687A
Network traffic classification method and related device
CN121486340A