Confusion traffic classification method and system based on denoising representation learning

Through a method based on denoising representation learning, header and payload embedding are generated and noise is removed. Combined with the cross-gate fusion mechanism, the problem of the accuracy of obfuscating traffic classification under the new obfuscating architecture is solved, and more efficient traffic classification is achieved.

CN120238360APending Publication Date: 2025-07-01INST OF COMPUTING TECH CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510473805.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The existing obfuscated traffic classification method has decreased accuracy under the new traffic obfuscated architecture, making it difficult to effectively distinguish the traffic patterns of different network applications and services, and the noise impact is severe, resulting in low classification performance.

Method used

Using a method based on denoising representation learning, headers and payload embeddings are generated through the pre-trained packet embedding transformation model, and noise is removed using feature extraction and deconfusing attention mechanisms, and embeddings are merged through a cross-gate fusion mechanism to generate the final combined embeddings for classification.

Benefits of technology

It improves the accuracy of confusing traffic classification, reduces the impact of noise, enhances the accuracy of the semantic representation of traffic, and improves the reliability and scalability of the classification model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120238360A_ABST
    Figure CN120238360A_ABST
Patent Text Reader

Abstract

The invention discloses an obfuscated traffic classification method and system based on denoising representation learning, and the method comprises the steps: inputting to-be-classified obfuscated traffic into a pre-trained data packet embedding conversion model, and generating header embedding and effective load embedding in a header and an effective load of each data packet of the traffic; aiming at header embedding and effective load embedding, removing confusion noise in an effective load through denoising representation learning; aiming at the effective load embedding of header embedding and confusion removal, a cross-gate fusion mechanism is adopted, header embedding and load embedding of data packets are merged, the filtered header embedding and the filtered load embedding are spliced, final combined embedding is obtained, and confusion traffic classification is achieved. The method and the system provided by the invention have excellent transferability, expandability and robustness, are high in accuracy on a confused traffic classification task, and can identify network behaviors in confused traffic transmission and classify network service traffic based on data packet characterization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to network security technology, and particularly to a method for classifying obfuscated traffic based on denoising representation learning. Background Art

[0002] Network traffic classification has received extensive attention from the academic and industrial communities due to its importance in multiple network domain applications such as traffic optimization, host behavior analysis, and network security monitoring. To enhance traffic data privacy protection and improve network security, various traffic obfuscation technologies have been applied to actual network applications and services. Traffic after being processed by different obfuscation technologies is called obfuscated traffic, that is, the inherent characteristics of the traffic can be masked through various obfuscation operations (such as encryption, camouflage, and randomization), which poses a huge challenge to traffic identification and classification. Due to this difficulty, traffic obfuscation technologies are widely used in anonymous networks such as Tor (Onion Router) to enable obfuscated traffic to evade network monitoring, behavior analysis, and censorship. In addition, these technologies may also be used to conceal the illegal activities of cybercriminals. Therefore, it is particularly important to classify obfuscated traffic.

[0003] So far, the research on classifying obfuscated traffic can be systematically summarized into three categories: deep packet inspection (DPI), feature engineering-based methods, and deep learning-based methods. Deep packet inspection (DPI) relies on the content and header information of network packets and uses feature matching or sequence matching to identify and classify network traffic. Although such methods have achieved certain results in identifying obfuscated traffic, they require a large amount of manual analysis and feature extraction. In addition, feature engineering-based traffic classification methods, such as support vector machines (SVM), decision trees, and convolutional neural networks (CNN), also rely on manually extracted features during the classification process. To address this limitation, researchers have developed a variety of deep learning-based classification methods, including sequence-specified, byte-specified, and graph-specified methods. These deep learning models can automatically extract features from the input network packets or flows, thus effectively reducing the errors that may be caused by manual feature extraction.

[0004] Traditional traffic obfuscation technologies allow each user or application to select an obfuscation method according to requirements. However, it will also introduce new features in the obfuscated traffic, and the differences in these features depend on the specific obfuscation mechanism, that is, whether the obfuscation methods are consistent. Since existing methods utilize these newly generated features to perform traffic classification tasks, a relatively high classification accuracy can be achieved.

[0005] To increase the strength of privacy protection, some existing new traffic obfuscation architectures have been proposed and widely adopted. This architecture is equipped with classic encoder and decoder components. The encoder aims to generate similar features for different traffic using a unified obfuscation strategy, while the decoder aims to restore the original form of the obfuscated traffic to ensure normal network functions. As Figure 1 shown, the combination of the encoder-decoder components not only preserves the regular traffic functions but also disables the identification of obfuscated traffic. For the obfuscated traffic generated by this new architecture, the classification accuracy of the latest methods has decreased by approximately 20%-60%. Currently, the new architecture brings three changes to obfuscated traffic classification: i) Feature obfuscation. The obfuscation strategy generates a unified and identical traffic pattern in different applications and services. This obfuscation of different features makes it difficult to distinguish obfuscated traffic; ii) Classification complexity. The goal of traffic classification becomes more complex, rather than the previous binary classification between simple obfuscated traffic and non-obfuscated traffic. Today's classification task is to identify specific services or applications with an obfuscated state; iii) Adding noise. The new obfuscation strategy requires introducing byte sequences, that is, adding noise to the traffic, which will modify the semantic content of the original traffic and affect the classification method; iv) Ignoring time features. Adding interference packets will change the time features of traffic data, and the current methods are not designed for this, resulting in general performance in classifying such data packets.

[0006] In the current existing research on obfuscated traffic classification, the classification accuracy of the existing obfuscated traffic classification technologies for the obfuscated traffic generated by the new architecture has decreased by approximately 20%-60%, and the classification effect is poor. The reason for this situation is that the obfuscation strategy homogenizes the traffic patterns of different network applications and services, eliminating the unique features inherent in each type of traffic; the byte sequences introduced during the traffic obfuscation process will modify the original semantics, thus bringing a large amount of noise to the classification task. The above two reasons lead to the poor effect of the current classification methods.

[0007] In summary, the above new traffic obfuscation architecture not only leads to the problems of decreased classification accuracy of obfuscated traffic and low performance in classifying data packets. It is urgent to introduce an attention mechanism and propose an efficient obfuscated traffic classification model that can effectively reduce the noise impact introduced by the byte sequences added during the traffic obfuscation process. Summary of the Invention

[0008] To address the traffic classification challenges brought by the new obfuscation framework in the above-mentioned existing technologies, an obfuscated traffic classification method based on denoising representation learning is proposed.

[0009] In a first aspect, an embodiment of the present application provides an obfuscated traffic classification method based on denoising representation learning. The method includes:

[0010] Packet embedding conversion step: Input the obfuscated traffic to be classified into a pre-trained packet embedding conversion model, and generate header embeddings and payload embeddings in the headers and payloads of each packet of the traffic.

[0011] Obfuscation removal step: Based on the header embeddings and payload embeddings, remove the obfuscation noise in the payload through denoising representation learning, where the obfuscation noise includes: noise introduced by padding and noise introduced by interference packets.

[0012] Packet feature fusion step: For the header embeddings and the obfuscation-removed payload embeddings, adopt a cross-gate fusion mechanism to merge the headers and payload embeddings of the packets, splice the header embeddings and the filtered payload embeddings to obtain the final combined embeddings, and realize the classification of the obfuscated traffic.

[0013] In the embodiment of this application, the above packet embedding conversion step includes:

[0014] Use a pre-trained model based on BERT to perform embedding conversion on the headers and payloads of each packet respectively, and generate corresponding header and payload embedding representations.

[0015] In the embodiment of this application, the above obfuscation removal step includes:

[0016] Feature extraction attention step: Through the feature extraction attention mechanism, extract the information in the payload that is helpful for classification.

[0017] Obfuscation removal attention step: Through the obfuscation removal attention mechanism, eliminate the obfuscation noise in the payload.

[0018] In the embodiment of this application, the above packet feature fusion step includes:

[0019] Perform a linear transformation operation on the header and payload embeddings of each packet, adopt cross-gate feature fusion to merge the headers and payload embeddings of the packets, and generate a fused denoised representation.

[0020] In the embodiment of this application, the above feature extraction attention step includes:

[0021] Initialize the feature preference matrix, use the learnable preference matrix and the payload embeddings as inputs, calculate the attention scores, and perform attention score normalization to calculate and obtain the token deviation matrix.

[0022] Multiply the token deviation matrix by the payload embeddings to calculate and obtain the feature deviation embedding matrix.

[0023] In the embodiment of this application, the above obfuscation removal attention step includes:

[0024] Introduce an additional multi-head attention mechanism, and use a new matrix that summarizes the relationships between all categories and the embeddings to represent the payload embeddings;

[0025] Replace the padding category bias vector in the new feature preference matrix with a zero vector;

[0026] Use the modified attention mechanism to calculate the de-confused embeddings, where the confused embeddings contain the category features after removing the padding interference.

[0027] In the embodiments of the present application, the above-mentioned data packet feature fusion steps include:

[0028] After applying linear transformation operations to the header embeddings and the de-confused payload embeddings respectively, perform non-linear transformation through the PReLU activation function;

[0029] Generate a gating vector through the Sigmoid layer, and the gating vector is used to weight the original embedding vectors;

[0030] Filter the payload embeddings using the header gating vector, and filter the header embeddings using the payload gating vector.

[0031] In a second aspect, the embodiments of the present application provide a confused traffic classification system based on denoising representation learning, which adopts the confused traffic classification method based on denoising representation learning as described above. The system includes:

[0032] Data packet embedding conversion module: Input the confused traffic to be classified, and based on the pre-trained data packet embedding conversion model, generate header embeddings and payload embeddings in the headers and payloads of each data packet in the traffic;

[0033] Confusion removal module: Based on the header embeddings and the payload embeddings, remove the confusion noise in the payload through denoising representation learning, where the confusion noise includes: the noise introduced by padding and the noise introduced by interference packets;

[0034] Data packet feature fusion module: For the header embeddings and the payload embeddings after confusion removal, adopt a cross-gate fusion mechanism to merge the headers and payload embeddings of the data packets, splice the filtered headers and payload embeddings, and obtain the final combined embeddings to achieve the classification of confused traffic.

[0035] In a third aspect, the embodiments of the present application provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the confused traffic classification method based on denoising representation learning are implemented.

[0036] Fourthly, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the above-mentioned traffic classification method based on denoising representation learning are implemented.

[0037] Compared with the related prior art, it has the following outstanding beneficial effects:

[0038] 1) The method of the present invention aims at the embedding transformation of network packet traffic; uses a pre-trained model based on BERT to independently process the headers and payloads of each packet, and generates a header embedding EHeader and a payload embedding EPayload accordingly. These embeddings are then used in the process of deconfusion and feature fusion, and finally serve as the basis for traffic classification of obfuscated traffic;

[0039] 2) The method of the present invention aims at the removal of obfuscation of network packet traffic embedding; the deconfusion module learns potential different features by defining a bias matrix, and considers two types of noise introduced by interfering packets and obfuscation padding respectively. A deconfusion attention component is designed to enhance the traffic semantic representation in a more accurate manner for the noise problems caused by interfering packets and obfuscation padding in EPayload;

[0040] 3) The method of the present invention aims at the cross-gate feature fusion of the data headers and payload embeddings of network packets, considers the correlation between EPayload and EHeader, and uses a cross-gate feature fusion module to accurately locate this relationship. After cross-gate feature fusion, the payload embedding EPayload and the header embedding EHeader are concatenated into a single representation to fully promote the downstream classification task. Description of the Drawings

[0041] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation of the present application. In the drawings:

[0042] Figure 1 It is a schematic diagram of a novel traffic obfuscation framework in the prior art;

[0043] Figure 2 It is a schematic diagram of the overall method for classifying obfuscated traffic in an embodiment of the present invention;

[0044] Figure 3 It is a schematic diagram of the method for classifying obfuscated traffic in an embodiment of the present invention;

[0045] Figure 4 It is a flow chart for identifying and classifying obfuscated traffic in an embodiment of the present invention;

[0046] Figure 5Schematic diagram of the traffic classification system for obfuscation according to an embodiment of the present invention;

[0047] Figure 6 Schematic diagram of the computer hardware according to the present invention. Detailed implementation manners

[0048] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following items (pieces)" or a similar expression thereof refers to any combination of these items, including any combination of a single item (piece) or plural items (pieces). For example, at least one of a, b, or c may represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c may be single or multiple.

[0049] It should also be understood that the term "and / or" in this article is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood by referring to the context before and after.

[0050] It should also be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above - mentioned processes do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0051] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.

[0052] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0053] In addition, in each embodiment of the present invention, each functional unit can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0054] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that makes a contribution to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0055] To make the above features and effects of the present invention more clearly and understandably described, specific embodiments are given below and are described in detail in conjunction with the accompanying drawings of the specification. This specification discloses one or more embodiments containing the features of the present invention. The disclosed embodiments are only for illustrative purposes. The protection scope of the present invention is not limited to the disclosed embodiments, and the present invention is defined by the appended claims.

[0056] The following is a system embodiment corresponding to the above method embodiment, and this embodiment can be implemented in cooperation with the above embodiment. The relevant technical details mentioned in the above embodiment are still valid in this embodiment. To avoid repetition, they will not be elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied to the above embodiment.

[0057] The method of the present invention aims to propose an ONTformer obfuscated traffic denoising model, mainly for the classification problem of obfuscated traffic. Through a dual attention mechanism and a feature fusion method, byte sequences are extracted from network traffic data packets and representations for classification are generated. In actual use, the ONTformer model can be deployed on nodes in the network to identify and classify obfuscated network traffic data packets. The dual attention mechanism calculates the traffic representation of class preference by evaluating the correlation between traffic load and relevant labels, thereby extracting effective traffic features. The denoising attention mechanism learns potential features through a bias matrix to reduce the impact of noise introduced during the obfuscation process on the classification task. Finally, the present invention integrates the packet header and payload information through a cross-gating feature fusion module to generate a more robust obfuscated traffic classification model.

[0058] To address the traffic classification challenges posed by the new obfuscation framework, the method of the present invention introduces an attention mechanism and proposes an efficient obfuscated traffic classification model. The classification model of the present invention can effectively reduce the noise impact introduced by the byte sequences added during the traffic obfuscation process. Specifically, the method of the present invention first evaluates the correlation between the traffic load and the relevant tags to calculate the traffic representation prioritizing categories, so as to appropriately extract effective traffic features. In addition, the method of the present invention also designs a denoising model to purify the traffic representation and adopts the cross-gated feature fusion technology to integrate the data packet header and payload information, and finally generates a powerful obfuscated traffic classification model.

[0059] As Figure 2 shown, the method steps mainly include:

[0060] Train the data packet embedding converter: Use pre-trained models based on BERT (Header-BERT and Payload-BERT) to process the header and payload of each data packet respectively to generate corresponding embedding representations;

[0061] Learn the relationship between data packet embeddings and labels: Define a category bias matrix to learn the different potential features of traffic data of different categories.

[0062] Remove data packet obfuscation noise: Consider two types of noise introduced by interfering data packets and obfuscation padding respectively. Enhance the traffic semantic representation in a more accurate way through the de-obfuscation attention component.

[0063] Fuse data packet features: Apply linear transformation operations to both EPayload and EHeader, and then apply the PReLU function. By applying different scaling factors to each channel, it plays a role similar to the attention mechanism.

[0064] Concatenate data packet features: The payload embedding EPayload and the header embedding EHeader are concatenated into a single representation to fully facilitate the downstream classification task.

[0065] The method of the embodiments of the present application will be described in detail below in conjunction with specific embodiments:

[0066] Embodiment 1

[0067] As Figure 3 shown, the embodiments of the present application provide an obfuscated traffic classification method based on denoising representation learning. The steps of the present invention include:

[0068] Data packet embedding conversion step 101: Input the obfuscated traffic to be classified into the pre-trained data packet embedding conversion model to generate a header embedding and a payload embedding in the header and payload of each data packet of the traffic;

[0069] Confusion removal step 102: Based on header embedding and payload embedding, through denoising representation learning, remove the confusion noise in the payload, where the confusion noise includes: the noise introduced by padding and the noise introduced by interfering packets;

[0070] Data packet feature fusion step 103: For the header embedding and the payload embedding after confusion removal, adopt a cross-gate fusion mechanism to merge the header and payload embeddings of the data packet, splice the header embedding and the filtered payload embedding to obtain the final combined embedding, and realize the classification of the confused traffic.

[0071] In the embodiment of the present application, the above data packet embedding conversion step 101 includes:

[0072] Use the pre-trained model based on BERT to perform embedding conversion on the header and payload of each data packet respectively, and generate the corresponding header and payload embedding representations.

[0073] Specifically, as Figure 3 shown, use the pre-trained model based on BERT to independently process the header and payload of each data packet, and correspondingly generate the header embedding EHeader and the payload embedding EPayload. These embeddings are then used in the process of deconfusion and feature fusion, and finally serve as the basis for the classification of the confused traffic.

[0074] In the embodiment of the present application, the above confusion removal step 102 includes:

[0075] Feature extraction attention step: Through the feature extraction attention mechanism, extract the information in the payload that is helpful for classification;

[0076] Deconfusion attention step: Through the deconfusion attention mechanism, eliminate the confusion noise in the payload.

[0077] Specifically, as Figure 3 shown, the deconfusion step learns the potentially different features by defining a bias matrix, and respectively considers the two types of noise introduced by interfering data packets and confusion padding. For the noise problem caused by the confusion padding in BPayload, a deconfusion attention component is designed to enhance the traffic semantic representation in a more accurate way.

[0078] In the embodiment of the present application, the above feature extraction attention step includes:

[0079] Initialize the feature preference matrix, use the learnable preference matrix and the payload embedding as inputs, calculate the attention scores, and perform attention score normalization to calculate and obtain the token bias matrix;

[0080] Multiply the token bias matrix by the payload embedding to calculate and obtain the feature bias embedding matrix.

[0081] More specifically, the core objective of the de - obfuscation module in the present invention is to eliminate the inaccurate representations introduced by the obfuscation padding and interference data packets in the data embedding module. This module achieves this goal through two components: feature extraction attention and de - obfuscation attention.

[0082] The specific steps of feature extraction attention are as follows:

[0083] 1. Initialize the category preference matrix:

[0084] First, initialize a category preference matrix \(A_0 = [A_{10}, A_{20}, \ldots, A_{C0}, A_{P0}\)

[0085] , \(A_{Chaff}]\), where \(A_{10}\) to \(A_{C0}\) represent the deviations of \(C\) different categories, \(A_{Chaff}\) represents the deviation of the interference data packet, and \(A_{P0}\) represents the deviation of the obfuscation padding.

[0086] 2. Calculate the attention scores

[0087] Use the learnable matrix \(A\) as the query (Query), the payload embedding \(E_{Payload}\) as the key (Key), and calculate the elements \(A_i\cdot E_{jy}\) in the result matrix \(A\times E_{Payload}\) to measure the similarity between the deviation vector \(A_i\) of the \(i\) - th category and the \(j\) - th payload feature \(E_{jy}\).

[0088] 3. Normalize the attention scores:

[0089] Normalize the scores through the softmax operation to obtain the attention distribution probability distribution of each category for different payload features, that is, the token deviation matrix.

[0090] 4. Calculate the feature deviation matrix

[0091] Multiply the token deviation matrix by the payload embedding \(B_{Payload}\) to obtain the deviation embedding \(A'\) of the specific category feature.

[0092] In the embodiments of the present application, the above - mentioned de - obfuscation attention steps include:

[0093] Introduce an additional multi - head attention mechanism, and use a new matrix that summarizes the relationships between all categories and the embeddings to represent the payload embedding;

[0094] Replace the padding category deviation vector in the new feature preference matrix with a zero vector;

[0095] Use the modified attention mechanism to calculate the de - obfuscation embedding, and the obfuscation embedding contains the category features after removing the padding interference.

[0096] Specifically, the specific steps of de - obfuscation attention are as follows:

[0097] 1. Multi-Head Attention Mechanism:

[0098] An additional multi-head attention mechanism is introduced, and A' (summarizing the relationships between all categories and embeddings) is used to refine the payload embedding BPayload to eliminate the abnormal influence of noise.

[0099] 2. Replacement Strategy:

[0100] To avoid the attention mechanism from wrongly focusing on padding noise, the padding category bias vector AP0 in A' is replaced with a zero vector, thereby neutralizing the negative impact of the data packets mainly related to padding.

[0101] 3. Calculate the Deobfuscated Embedding:

[0102] The modified attention mechanism is used to calculate the deobfuscated embedding Adeobf', which contains the category features after removing padding interference.

[0103] In the embodiment of this application, the above data packet feature fusion step 103 includes:

[0104] Linear transformation operations are performed on the embeddings of the header and payload of each data packet, and cross-gated feature fusion is adopted to merge the header and payload embeddings of the data packet to generate a fused denoised representation.

[0105] Specifically, as Figure 4 shown, this method considers the correlation between BPayload and EHeader, and adopts a cross-gate feature fusion module to accurately locate this relationship. After cross-gate feature fusion, the payload embedding EPayload and the header embedding EHeader are concatenated into a single representation to fully facilitate the downstream classification task.

[0106] In the embodiment of this application, the above data packet feature fusion step includes:

[0107] After applying linear transformation operations to the header embedding and the deobfuscated payload embedding respectively, a non-linear transformation is performed through the PReLU activation function;

[0108] A gating vector is generated through the Sigmoid layer, and the gating vector is used to weight the original embedding vector;

[0109] The payload embedding is filtered using the header gating vector, and the header embedding is filtered using the payload gating vector.

[0110] Specifically, the present invention also uses a cross-gated feature fusion module to merge the header and payload embeddings of the data packet to capture the internal correlation between the two and provide a more effective representation for the downstream classification task. The specific process is as follows:

[0111] 1. Linear Transformation and PReLU Activation:

[0112] Apply linear transformation operations to the header embedding EHeader and the deobfuscated payload embedding Adeobf′ respectively, and then perform non-linear transformation through the PReLU activation function. The PReLU function not only serves as a non-linear transformation but also acts like an attention mechanism, adjusting the negative elements of each channel with different scaling factors.

[0113] 2. Generate gating vectors:

[0114] Generate gating vectors through the Sigmoid layer, which are used to weight the original embedding vectors (element-wise multiplication). The output values of the Sigmoid function are between [0,1] and can be regarded as "attention weights", indicating the importance of each element in the final representation.

[0115] 3. Cross filtering:

[0116] Use the header gating vector g1 to filter the payload embedding Adeobf′, and use the payload gating vector g2 to filter the header embedding EHeader. This dual processing not only serves as a non-linear transformation but also introduces an attention mechanism through different scaling factors of the negative elements in different channels, enabling the model to filter out unimportant information and retain important information.

[0117] 4. Final output:

[0118] Concatenate the filtered header and payload embeddings to obtain the final combined embedding E for downstream classification tasks.

[0119] Through these two modules, ONTformer can effectively extract pure and distinguishable feature representations from the obfuscated network traffic, thereby improving the performance of classification tasks.

[0120] As described above, the method of the present invention can be preferably implemented.

[0121] In summary, compared with the prior art, the method of the present invention proposes a dual attention mechanism to learn the implicit correlation between obfuscated traffic and corresponding labels. It can effectively capture the relationship between obfuscated traffic and corresponding labels. At the same time, the present invention can effectively reduce the semantic noise brought by padding and interfering packets during the traffic obfuscation process, thereby inferring pure and distinguishable packet headers and payload embeddings for further classification tasks. The invention has excellent transferability, scalability and robustness, has a high accuracy in the obfuscated traffic classification task, and can identify network behaviors in obfuscated traffic transmission and classify network service traffic based on packet representations.

[0122] Embodiment 2

[0123] As Figure 5As shown in the figure, an embodiment of the present application provides a traffic classification system for obfuscated traffic based on denoising representation learning, which adopts the above-mentioned traffic classification method for obfuscated traffic based on denoising representation learning. The system includes:

[0124] Packet embedding conversion module 201: Input the obfuscated traffic to be classified, and based on the pre-trained packet embedding conversion model, generate header embeddings and payload embeddings in the headers and payloads of each packet of the traffic.

[0125] Obfuscation removal module 202: Based on the header embeddings and payload embeddings, remove the obfuscation noise in the payload through denoising representation learning, where the obfuscation noise includes: noise introduced by padding and noise introduced by interfering packets.

[0126] Packet feature fusion module 203: For the header embeddings and the obfuscation-removed payload embeddings, adopt a cross-gate fusion mechanism to merge the headers and payload embeddings of the packets, splice the filtered header and payload embeddings, and obtain the final combined embeddings to achieve the classification of obfuscated traffic.

[0127] Embodiment III

[0128] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the above-mentioned traffic classification method for obfuscated traffic based on denoising representation learning are implemented.

[0129] Embodiment IV

[0130] An embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the above-mentioned traffic classification method for obfuscated traffic based on denoising representation learning are implemented.

[0131] In addition, the traffic classification method for obfuscated traffic based on denoising representation learning described in Figure 1 the embodiments of the present application can be implemented by an electronic device, such as a computer device. Figure 6 It is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present application.

[0132] In some of these embodiments, the computer device may further include a communication interface 83 and a bus 80. Among them, as Figure 6 shown, the processor 81, the memory 82, and the communication interface 83 are connected through the bus 80 and complete communication with each other.

[0133] Specifically, the above-mentioned processor 81 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured as one or more integrated circuits for implementing the embodiments of the present application.

[0134] The memory 82 can be used to store or cache various data files required for processing and / or communication, as well as possible computer program instructions executed by the processor 81.

[0135] The processor 81 reads and executes the computer program instructions stored in the memory 82 to implement any one of the above-mentioned denoising representation learning-based obfuscated traffic classification generation methods.

[0136] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0137] The above embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

Claims

1. A confusion traffic classification method based on denoising representation learning, characterized in that: The method comprises: Packet embedding conversion step: inputting the obfuscated traffic to be classified into the pre-trained packet embedding conversion model, generating header embedding and payload embedding in the header and payload of each packet of the traffic; The obfuscation removal step includes: removing the obfuscation noise in the payload through denoising representation learning based on the header embedding and the payload embedding, wherein the obfuscation noise includes: noise introduced by padding and noise introduced by interference packets; Data packet feature fusion step: for the header embedding and the effective load embedding with obfuscation removed, a cross-gate fusion mechanism is adopted to merge the header and load embedding of the data packet, and the header embedding and the filtered load embedding are spliced ​​to obtain the final combined embedding to achieve classification of obfuscated traffic.

2. The method for classifying obfuscated traffic based on denoising representation learning according to claim 1 is characterized in that: The data packet embedding conversion step comprises: Use the BERT-based pre-trained model to perform embedding conversion on the header and payload of each data packet respectively, and generate corresponding header and payload embedding representations.

3. The method for classifying obfuscated traffic based on denoising representation learning according to claim 1 is characterized in that: The obfuscation removal step comprises: Feature extraction attention step: Through the feature extraction attention mechanism, information in the payload that helps classification is extracted; De-obfuscating attention step: The obfuscating noise within the payload is removed through the de-obfuscating attention mechanism.

4. The method for classifying obfuscated traffic based on denoising representation learning according to claim 1 is characterized in that: The data packet feature fusion step comprises: A linear transformation operation is performed on the embeddings of the header and payload of each packet, and cross-gated feature fusion is used to merge the header and payload embeddings of the packet to generate a fused denoised representation.

5. The method for classifying obfuscated traffic based on denoising representation learning according to claim 3 is characterized in that: The feature extraction attention step includes: Initialize a feature preference matrix, use the learnable preference matrix and the payload embedding as input, calculate an attention score, normalize the attention score, and calculate a token bias matrix; The token bias matrix is ​​multiplied by the payload embedding to calculate a feature preference embedding matrix.

6. According to claim 3, the confusion traffic classification method based on denoising representation learning is characterized in that: The de-obfuscated attention step includes: Introducing an additional multi-head attention mechanism, using a new matrix summarizing the relationship between all categories and embeddings to represent the load embedding; replacing the filled category bias vector in the new feature preference matrix with a zero vector; The modified attention mechanism is used to compute the de-obfuscated embedding, which contains the category features after removing the padding interference.

7. The method for classifying obfuscated traffic based on denoising representation learning according to claim 4 is characterized in that: The data packet feature fusion step comprises: After applying linear transformation operations to the head embedding and the deobfuscated payload embedding, nonlinear transformation is performed through the PReLU activation function; Generate a gating vector through a Sigmoid layer, wherein the gating vector is used to weight the original embedding vector; Filter the payload embeddings using the header gating vector, and filter the header embeddings using the payload gating vector.

8. A confusion traffic classification system based on denoising representation learning, using the confusion traffic classification method based on denoising representation learning as claimed in any one of claims 1 to 7, characterized in that: The system comprises: Packet embedding conversion module: inputs the obfuscated traffic to be classified, and generates header embedding and payload embedding in the header and payload of each packet of the traffic based on the pre-trained packet embedding conversion model; An obfuscation removal module: based on the header embedding and the payload embedding, removing the obfuscation noise in the payload through denoising representation learning, wherein the obfuscation noise includes: noise introduced by padding and noise introduced by interference packets; Data packet feature fusion module: for the header embedding and the effective load embedding with obfuscation removal, a cross-gate fusion mechanism is adopted to merge the header and load embedding of the data packet, and the filtered header and load embedding are spliced ​​to obtain the final combined embedding to achieve the classification of obfuscated traffic.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the confusion traffic classification method based on denoising representation learning described in any one of claims 1 to 7 are implemented.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the confusion traffic classification method based on denoising representation learning as described in any one of claims 1 to 7 are implemented.