Encrypted malicious Trojan flow detection method based on mask auto-encoder and multistage flow modeling
By building a multi-level traffic modeling matrix and combining a time-sequence convolution network and a multi-level attention mechanism, using a self-supervised pre-training strategy, the problem of traditional methods identifying malicious Trojan traffic in an encrypted traffic environment is solved, and efficient and low-cost encrypted malicious Trojan traffic detection is achieved.
Patent Information
- Application Number
- CN202510427070.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-04
AI Technical Summary
The existing traditional detection methods are difficult to effectively identify encrypted malicious Trojan traffic, especially in network environments where traffic encryption technology is widely used. In addition, deep learning methods require a large amount of tag data for supervision and training, resulting in high deployment costs and insufficient model robustness.
Using a method based on masked autoencoder and multi-level traffic modeling, a multi-level traffic modeling matrix is constructed, combined with a time-sequence convolution network and a multi-level attention mechanism, and a self-supervised pre-training strategy is used to achieve accurate identification of encrypted malicious Trojan traffic.
It realizes efficient identification of encrypted malicious Trojan traffic without label data, reduces dependence on label data, and improves the classification recognition ability and robustness of the model.
Smart Images

Figure CN120263482A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of cyberspace security, and relates to an encrypted malicious trojan traffic detection method based on a masked autoencoder and multi-level traffic modeling. Background Art
[0002] As one of the most common types of malware, trojans are highly harmful. Currently, the mainstream malicious trojans mainly include Remote Access Trojans (RATs), backdoor trojans, ransomware, and distributed denial-of-service (DDoS) trojans. Among them, trojan software with remote control capabilities is particularly harmful. Taking remote access trojans as an example, attackers first achieve initial penetration through methods such as spear-phishing, implant remote control malicious trojans on the victim's host, and then establish a command and control (C&C) channel. The controlled host periodically polls the C&C server through DNS tunneling or the HTTP protocol to obtain malicious instructions, such as destructive operations like keylogging, screenshotting, and file transfer, resulting in a serious risk of information leakage.
[0003] In addition, traffic encryption has become an inevitable trend. According to Google's Transparency Report, by 2024, the proportion of encrypted traffic in its services has exceeded 95%. However, while traffic encryption technology protects privacy and data security, it also provides opportunities for malicious traffic. Malicious traffic hides its malicious behavior through encryption technology, thus evading security detection, rendering traditional detection methods ineffective, and posing a huge challenge to China's network supervision. Given the high risk and concealment of encrypted malicious trojans, detecting their traffic has important practical value.
[0004] Traditional network traffic analysis methods mainly identify network services based on basic traffic fingerprints, such as shallow features like communication protocols and port numbers. However, due to the complexity and diversity of the current network environment, these methods are no longer applicable. To solve this problem, many studies have adopted machine learning algorithms based on statistical features to achieve encrypted traffic analysis. However, these methods rely on expert prior knowledge to select specific features, and statistical feature drift can also significantly affect the robustness of the model. In recent years, end-to-end traffic analysis methods based on deep learning can directly extract features from the raw packet bytes to achieve efficient traffic analysis. However, deep learning frameworks also face challenges. On the one hand, in the scenario of long packets, the byte sequence may mask the key discriminant information in other packets, resulting in certain performance fluctuations in traditional deep learning methods. On the other hand, deep learning-based methods usually require a large amount of labeled data for supervised training, and data annotation is a time-consuming and laborious task, leading to high costs for model deployment.
[0005] In recent years, pre-training methods based on large amounts of unlabeled data have demonstrated excellent performance in the fields of Computer Vision (CV) and Natural Language Processing (NLP). Compared with traditional supervised learning methods, pre-training methods under self-supervised learning have two main advantages: (1) Pre-training methods can obtain more effective feature representations, thereby reducing the dependence of downstream tasks on labeled data. (2) Through pre-training, better model initialization parameters can be obtained, thus accelerating model convergence and improving the performance of classification tasks. Frontier pre-training models from the CV field, such as Masked AutoEncoders (MAE), provide a new technical path for malicious traffic identification and classification. MAE learns the deep latent representations of unlabeled images through self-supervised learning by randomly masking part of the input image and reconstructing the missing pixels.
[0006] Therefore, the present invention introduces an encrypted malicious trojan traffic detection method based on MAE with self-supervised learning ability. First, a Multi-level Flow Modeling (MFM) matrix is created to simulate the original traffic distribution. This matrix is composed of the bytes of original trojan traffic data packets and contains traffic information at different granularities in a structured manner. On this basis, an encrypted malicious trojan traffic classification model based on the MFM matrix is designed. The Temporal Convolutional Network (TCN) is used to extract temporal behavior features, and the packet-level and flow-level attention mechanisms are fused to capture the dependencies within and across data packets. The present invention is based on the MAE pre-training paradigm, including two-stage processes of pre-training and fine-tuning. In the pre-training stage, a part of the data packet bytes in the MFM matrix are randomly masked as the model input, and then the asymmetric encoder-decoder structure is used to reconstruct the masked byte content. In the fine-tuning stage, the encoder parameters obtained after the pre-training stage are loaded into the model, and a small amount of encrypted malicious trojan traffic label data is used to fine-tune the model for trojan traffic classification. Summary of the Invention
[0007] In order to strengthen the supervision of cyberspace security and achieve the precise identification of encrypted malicious Trojan traffic in a dynamic network traffic environment, the present invention proposes an encrypted malicious Trojan traffic detection method based on a masked autoencoder and multi-level traffic modeling. For encrypted malicious Trojan attack traffic, the original traffic bytes are converted into a multi-level traffic modeling matrix MFM, and then the matrix is sliced into non-overlapping small blocks. Each small block is mapped into a vector and position encoding information is added as the input to the encoder. The packet-level attention mechanism and the flow-level attention mechanism are used to capture the dependencies within and between data packets respectively, and the temporal convolutional network is fused to explicitly model the traffic behavior patterns contained in the data packet time interval and the data packet length sequence. A self-supervised pre-training strategy based on MAE is adopted to train the encoder using a large amount of available unlabeled data, and fine-tuning is performed using a small amount of Trojan traffic data in the downstream task to achieve efficient encrypted malicious Trojan traffic detection.
[0008] To achieve the above object, the present invention provides the following technical solutions:
[0009] An encrypted malicious Trojan traffic detection method based on a masked autoencoder and multi-level traffic modeling, comprising the following steps:
[0010] (1) Convert the original bytes of encrypted malicious Trojan traffic into a multi-level traffic modeling matrix, which directly reflects the byte distribution of the traffic, and extract the arrival time interval of data packets and the data packet length sequence features for explicitly modeling the temporal dependencies between data packets;
[0011] (2) Slice the matrix constructed in step (1) into non-overlapping small blocks, map each small block into a vector, and add position encoding information for use as the input to the encoder;
[0012] (3) Adopt a hierarchical attention mechanism, and use the packet-level attention mechanism and the flow-level attention mechanism to capture the dependencies within and between data packets respectively;
[0013] (4) Fuse the temporal convolutional network to explicitly model the traffic behavior patterns contained in the data packet time interval and the cumulative data packet length, and capture the temporal dependencies between data packets;
[0014] (5) Adopt a self-supervised pre-training strategy based on a masked autoencoder, train the encoder using a large amount of unlabeled data, and fine-tune using a small amount of Trojan traffic label data in the downstream task to achieve efficient encrypted malicious Trojan traffic detection.
[0015] Further, the step (1) specifically includes the following sub-steps:
[0016] (1.1) First, the original traffic is divided into flows based on the five-tuple information of the source IP address, destination IP address, source port, destination port, and protocol type. At the same time, the Ethernet header is removed to strip the link layer features, the port numbers are set to zero to avoid traffic fingerprint interference, and a randomized IP address replacement strategy is adopted, while preserving the traffic directionality as the identifier of the uplink and downlink traffic;
[0017] (1.2) Extract the arrival time interval sequence {IPT1, IPT2, …, IPT M} and the packet length sequence {±pkt_length1, ±pkt_length2, …, ±pkt_length M} of the first M packets, where IPT M represents the relative arrival time interval between the Mth packet and the previous packet, and ±pkt_length M represents the length of the Mth packet, and the directionality is preserved, where the packets sent from the victim to the attacker are recorded as positive, and the packets sent from the attacker to the victim are recorded as negative;
[0018] (1.3) Extract the first M consecutive packets, and through three-level feature abstraction at the byte level, packet level, and flow level, construct a two-dimensional matrix MFM with a size of H×W.
[0019] Further, perform logarithmic normalization on the arrival time interval sequence in step (1.2) to prevent excessive time span and improve numerical stability, and perform standardization on the packet length sequence to control the numerical range.
[0020] Further, the specific operation of step (1.3) is as follows:
[0021] Each packet is divided into two parts: a header and a payload. Each row in the matrix corresponds to a type of byte, divided into a header row or a payload row, and the original byte values are used as the initial features to prevent semantic loss. The header row retains the network layer, transport layer, and extensible header information, and the payload row stores the original byte stream of the application layer. When the length exceeds 240 bytes, tail truncation is performed. A single packet is constructed into a packet-level matrix with a size of H / M * W, and M packet-level matrices are stacked along the second dimension, i.e., the column direction, to form the complete flow representation MFM. When the actual number of bytes of the packet is insufficient, a zero-padding strategy is adopted to ensure dimension consistency.
[0022] Further, in step (2), the matrix constructed in step (1) is cut into non-overlapping small blocks, each small block is mapped into a vector, and position encoding information is added to be used as the input of the encoder.
[0023] Further, step (2) specifically includes the following sub-steps:
[0024] (2.1) Divide the MFM matrix into non - overlapping 2D patches of size P×P, obtaining N = HW / P 2 patches, denoted as and map these patches to D - dimensional embedding vectors through a linear layer;
[0025] (2.2) To preserve the position information, add the position encoding information to the patch embeddings as the input to the encoder. The position encoding uses common sine and cosine functions, and its calculation method is as follows:
[0026] PE(pos, 2i) = sin(pos / 10000 2i / d )
[0027] PE(pos, 2i + 1) = cos(pos / 10000 2i / d )
[0028]
[0029] where PE(pos, 2i) and PE(pos, 2i + 1) represent the position encoding vectors at the pos - th position of the i - th feature dimension in the current dimension respectively. 2i and 2i + 1 correspond to the even - numbered dimension and the odd - numbered dimension respectively, represents the i - th patch, represents a learnable linear transformation matrix used to map the patch to a D - dimensional embedding vector, E pos ∈ R N×D represents the position encoding matrix, x e ∈ R N×D represents the final input sequence with position encoding, which is used as the input to the subsequent encoder.
[0030] Further, the step (3) specifically includes the following sub - steps:
[0031] (3.1) The packet - level attention mechanism part uses a flow encoder based on Vision Transformer to achieve intra - packet feature interaction. To prioritize the learning of the internal dependencies between the packet header region and the payload region, restrict the self - attention calculation to be only between patches within the same packet, rather than performing global interaction on the entire MFM matrix, so as to generate a representation vector with packet - level semantic integrity. Through the multi - head attention mechanism, each patch within the packet can achieve dynamic interaction based on the correlation weights of different attention heads. Its calculation process can be expressed as Concat(head1, head2, …, head n ,), where each attention head is calculated through the attention function:
[0032] Q = x l W Q, K = x l W K , V = x l W V
[0033]
[0034] where x l ∈ R N×D represents the input of the l-th layer, and Q, K, V represent the Query, Key, and Value matrices in the attention mechanism respectively. is a learnable parameter matrix used to project the input into different spaces, and D k is the dimension of the attention head, and QK T represents the dot product of Query and Key, measuring the attention correlation.
[0035] (3.2) The flow-level attention mechanism adaptively transforms the feature granularity by performing row pooling on the features output by the packet-level attention mechanism x r = RowPooling(x' p ), and generates row-level blocks representing the headers and payloads. These row-level blocks are fed into the encoder as new inputs, and the multi-head attention is used to extract the relationships between packets. Then, all row-level features are aggregated through column pooling to obtain the final representation x MFM ∈ R D .
[0036] Furthermore, the temporal convolutional module in step (4) consists of multiple layers of one-dimensional dilated convolutions. Each layer includes LayerNorm normalization, causal convolution, and the non-linear activation function ReLU. Finally, the residual connection is used to stabilize the training of the deep network. After this process, the output of the module is averaged along the time dimension to obtain the global temporal features, which are projected through a fully connected layer to align with the output of the encoder.
[0037] Furthermore, step (5) includes the following sub-steps:
[0038] (5.1) In the pre-training stage, a masked autoencoder is used to construct an asymmetric encoder-decoder architecture to reconstruct the original byte data of the MFM matrix. Specifically, a certain proportion of random masking operations are performed on the MFM matrix, and only a small number of unmasked blocks are retained as the input of the encoder. The encoder extracts features from these visible blocks and outputs encoded tokens. Then, a lightweight decoder is used to combine the encoded tokens and masked tokens to reconstruct the covered regions of the MFM matrix. The optimization is performed by calculating the mean square error between the true value and the reconstructed value, and the calculation formula is as follows:
[0039] L rec= MSE(y real , y rec )
[0040] where L rec represents the reconstruction loss, MSE represents the mean squared error function, y real represents the true block content masked in the original MFM matrix, and y rec represents the output of the decoder, i.e., the reconstructed block content.
[0041] (5.2) In the fine-tuning stage, the encoder parameters obtained in the pre-training stage are loaded into the packet-level attention module and the flow-level attention module, which are used to extract local features within the data packet and global features of the traffic session, respectively. To accelerate convergence and compress the model size simultaneously, a parameter sharing strategy is implemented between the packet-level encoder and the flow-level encoder;
[0042] (5.3) In the classification stage, each block in the MFM matrix undergoes two-stage pooling operations of row pooling and column pooling for feature fusion to generate a global classification feature vector. The output vector of the temporal convolutional module is concatenated to obtain the final classification vector. This vector is flattened and then input into the MLP, and finally, the predicted distribution of the encrypted malicious trojan traffic categories is output C is the total number of categories. The classification loss is obtained by calculating the cross-entropy between the predicted distribution and the true label, and the calculation formula is as follows:
[0043]
[0044] where represents the predicted output of the model, indicating the predicted probability that the sample belongs to each category, and y is the true label of the sample.
[0045] An electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the encrypted malicious trojan traffic detection method based on the masked autoencoder and multi-level traffic modeling is implemented.
[0046] A computer-readable storage medium stores computer instructions, and when the computer instructions are executed by the processor, the encrypted malicious trojan traffic detection method based on the masked autoencoder and multi-level traffic modeling is implemented.
[0047] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0048] (1) The present invention constructs a matrix based on the original traffic bytes, includes traffic information of different granularities in a structured manner, does not require relying on expert knowledge to design statistical features, and realizes fine-grained modeling of traffic semantics.
[0049] (2) The present invention does not rely on a large amount of encrypted malicious trojan traffic label data, and can learn the potential characteristics of traffic from a large amount of unlabeled traffic data, so as to utilize a small number of trojan traffic label sample data in downstream tasks to achieve the classification of encrypted malicious trojan traffic.
[0050] (3) The present invention combines a temporal convolutional network and a multi-level attention mechanism, effectively fusing the structural semantic features and traffic behavior features of encrypted malicious trojan traffic, and improving the classification and recognition ability of encrypted malicious trojan traffic. The temporal convolutional network can effectively model the temporal relationship between data packets and capture the periodic and mutative behavior features in trojan traffic, while the multi-level attention mechanism focuses more on the global context information of traffic. The temporal convolutional network can be used as a supplement to the attention mechanism, and through multi-modal fusion, effectively enhance the overall temporal modeling ability to achieve higher-precision classification. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments recorded in the embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings.
[0052] Figure 1 is a flowchart of the encrypted malicious trojan traffic detection method based on a masked autoencoder and multi-level traffic modeling provided by the present invention,
[0053] Figure 2 is a schematic diagram of the multi-level traffic modeling matrix in the embodiment of the present invention,
[0054] Figure 3 is a schematic diagram of the pre-training - fine-tuning architecture in the embodiment of the present invention,
[0055] Figure 4 is the change of the F1-score value detected by the method of the present invention under different mask ratios.
[0056] Figure 5 is the result comparison of using the global attention mechanism and removing the TCN module. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0057] Combined with the drawings and embodiments, the present invention will be further described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and not to limit the scope of the present invention.
[0058] Embodiment: The present invention proposes a method for early and accurate identification of fine-grained behaviors of malicious trojans for continuous connections, and its method flowchart is as Figure 1As shown in the figure, it includes three parts. The first part is the pre-training stage. Using a large amount of unlabeled malicious traffic data on the network, an MFM matrix is constructed from the original traffic bytes, and a random masking operation is performed. Then, it is input into the encoder to learn the latent feature representation. The decoder is used to reconstruct the masked matrix bytes, and the reconstruction loss is measured by the mean square error to continuously optimize the learning ability of the encoder. The second part loads the encoder parameters obtained in the pre-training stage into the packet-level attention module and the flow-level attention module, and fine-tunes the model with a small amount of encrypted Trojan traffic label data. The third part uses a temporal convolutional network to learn the Trojan traffic behavior characteristics contained in the packet arrival time interval and the packet length sequence, explicitly models the temporal dependence between packets, and finally splices with the output of the fine-tuning stage for Trojan traffic detection.
[0059] Specifically, the method of the present invention has the following steps:
[0060] (1) Convert the original bytes of encrypted malicious Trojan traffic into a multi-level traffic modeling matrix, which directly reflects the byte distribution of the traffic, and extract the arrival time interval of the packets and the packet length sequence features for explicitly modeling the temporal dependence between packets;
[0061] The specific process of this step is as follows:
[0062] (1.1) First, divide the original traffic according to the five-tuple information of the source IP address, destination IP address, source port, destination port, and protocol type. At the same time, remove the Ethernet header to strip the link layer features, set the port number to zero to avoid traffic fingerprint interference, and adopt a randomized IP address replacement strategy to retain the traffic directionality as the identifier of the uplink and downlink traffic.
[0063] (1.2) Extract the arrival time interval sequence {IPT1, ITP2, …, IPT M} and the packet length sequence {±pkt_length1, ±pkt_length2, …, ±pkt_length M} of the first M packets, where IPT M represents the relative arrival time interval between the Mth packet and the previous packet, and ±pkt_length M represents the length of the Mth packet, and the directionality is retained, where the traffic sent from the victim to the attacker is recorded as positive, and the traffic sent from the attacker to the victim is recorded as negative. Perform logarithmic normalization on the arrival time interval sequence to prevent too large time span and improve numerical stability, and perform standardization on the packet length sequence to control the numerical range.
[0064] (1.3) Extract the first M consecutive packets and construct a two-dimensional matrix MFM with a size of H×W through byte-level, packet-level, and flow-level three-level feature abstraction. The specific operations are as follows:
[0065] Each data packet is divided into two parts: a header and a payload. Each row in the matrix corresponds to a type of byte, divided into a header row or a payload row, and the original byte values are used as initial features to prevent semantic loss. The header row retains network layer, transport layer, and extensible header information, and the payload row stores the original byte stream of the application layer. When the length exceeds 240 bytes, tail truncation is performed. A single data packet is constructed into a packet-level matrix of size H / M * W, and M packet-level matrices are stacked along the second dimension, i.e., the column direction, to form a complete flow representation MFM. When the actual number of bytes in the data packet is insufficient, a zero-padding strategy is adopted to ensure dimensional consistency.
[0066] In the present invention, the MFM matrix is set to contain 15 data packets, and finally a fixed-size MFM matrix of 120 * 40 bytes is constructed, which contains 15 packet-level units. Each packet-level unit includes 2 header rows, which are compatible with IP, TCP, UDP standard headers and optional fields, and 6 payload rows, retaining the core features of the application layer.
[0067] (2) Cut the matrix constructed in step (1) into non-overlapping small blocks, each small block is mapped into a vector, and position encoding information is added, which is used as the input of the encoder;
[0068] The specific process of this step is as follows:
[0069] (2.1) Divide the MFM matrix into non-overlapping 2D blocks (patches) of size P×P, obtaining N = HW / P 2 blocks, denoted as and map these blocks into D-dimensional embedding vectors through a linear layer;
[0070] (2.2) To retain the position information, add the position encoding information to the block embedding as the input of the encoder. The position encoding uses common sine and cosine functions, and its calculation method is as follows:
[0071] PE(pos, 2i) = sin(pos / 10000 2i / d )
[0072] PE(pos, 2i + 1) = cos(pos / 10000 2i / d )
[0073]
[0074] where PE(pos, 2i) and PE(pos, 2i + 1) respectively represent the position encoding vectors at the pos-th position in the i-th feature dimension in the current dimension. 2i and 2i + 1 respectively correspond to the even and odd dimensions, represents the i-th block, represents a learnable linear transformation matrix used to transform the block is mapped to a D-dimensional embedding vector, E pos ∈R N×D represents the position encoding matrix, x e ∈R N×D represents the final input sequence with position encoding, which is used as the input for the subsequent encoder.
[0075] In the present invention, D = 192, P = 2, N = 60 * 20 = 1200 are set. The block design ensures that each 2 * 2 block corresponds to the same type of features in the original byte stream. Each packet matrix contains one row of header blocks and three rows of payload blocks after performing the embedding operation, forming a structured feature representation.
[0076] (3) Adopt a hierarchical attention mechanism to capture the dependencies within and between data packets by using the packet-level attention mechanism and the flow-level attention mechanism respectively;
[0077] The specific process of this step is as follows:
[0078] (3.1) The packet-level attention mechanism part uses a flow encoder based on Vision Transformer to achieve in-packet feature interaction. In order to prioritize the learning of the internal dependencies in the header region and payload region of the data packet, the self-attention calculation is restricted to be only between blocks within the same data packet, rather than performing global interaction on the entire MFM matrix, so as to generate a representation vector with packet-level semantic integrity. Through the multi-head attention mechanism, each block within the data packet can achieve dynamic interaction based on the correlation weights of different attention heads. Its calculation process can be expressed as Concat(head1,head2,…,head n ,), where each attention head is calculated through the attention function:
[0079] Q = x l W Q , K = x l W K , V = x l W V
[0080]
[0081] where x l ∈R N×D represents the input of the l-th layer, and Q, K, V represent the Query, Key, and Value matrices in the attention mechanism respectively, is a learnable parameter matrix used to project the input into different spaces, D k is the dimension of the attention head, and QK T represents the dot product of Query and Key, measuring the attention correlation.
[0082] (3.2) The flow-level attention mechanism performs row pooling on the features output by the packet-level attention mechanism x r = RowPooling(x' p ), to achieve an adaptive transformation of the feature granularity. By performing mean pooling row by row to generate the row pooling output features, row-level blocks representing the headers and payloads are generated. These row-level blocks are fed into the encoder as new inputs, and multi-head attention is used to extract the relationships between packets. Subsequently, all row-level features are aggregated through column pooling to obtain the final representation x MFM ∈R D . This step achieves a time complexity of Q(N), effectively modeling flow-level features such as traffic behavior patterns and interaction time series while keeping the model lightweight.
[0083] In this paper, n = 16 and L = 4 are set. Diverse feature correlations are learned through 16 parallel attention heads, and a 4-layer alternating multi-head attention and feed-forward network are adopted to gradually enhance the intra-packet feature expression ability. Such a design conforms to the strong information correlation characteristics within traffic data packets, ensuring the model representation ability while reducing the computational complexity, and significantly improving the processing efficiency of traffic analysis.
[0084] (4) Incorporate a temporal convolutional network to explicitly model the traffic behavior patterns contained in the packet time intervals and cumulative packet lengths, and capture the temporal dependencies between packets. The temporal convolutional module consists of multiple layers of one-dimensional dilated convolutions. Each layer includes LayerNorm normalization, causal convolution, and the non-linear activation function ReLU. Finally, residual connections are used to stabilize the training of the deep network. After this processing, average pooling is performed on the module output along the time dimension to obtain global temporal features, which are projected through a fully connected layer to align with the encoder output.
[0085] In the present invention, a 3*3 convolutional kernel is used, and the dilation coefficients are 1, 2, and 4 in sequence. The output channels of each layer are 64. After the TCN module is processed, average pooling is performed on the module output along the time dimension to obtain global temporal features, which are projected through a fully connected layer to 192 dimensions and aligned with the encoder output.
[0086] (5) Adopt a self-supervised pre-training strategy based on a masked autoencoder to train the encoder using a large amount of unlabeled data, and fine-tune it using a small amount of trojan traffic label data in downstream tasks to achieve efficient encrypted malicious trojan traffic detection.
[0087] The specific process of this step is as follows:
[0088] (5.1) In the pre-training stage, a masked autoencoder is used to construct an asymmetric encoder-decoder architecture to reconstruct the original byte data of the MFM matrix. Specifically, a certain proportion of random masking operations are performed on the MFM matrix, and only a small number of unmasked blocks are retained as the input of the encoder. The encoder extracts features from these visible blocks and outputs encoded tokens. Then, a lightweight decoder is used to combine the encoded tokens and masked tokens to reconstruct the covered area of the MFM matrix. Optimization is carried out by calculating the mean square error between the true value and the reconstructed value, and the calculation formula is as follows:
[0089] L rec = MSE(y real , y rec )
[0090] Among them, L rec represents the reconstruction loss, MSE represents the mean square error function, y real represents the true block content masked in the original MFM matrix, and y rec represents the output of the decoder, that is, the reconstructed block content.
[0091] (5.2) In the fine-tuning stage, the encoder parameters obtained in the pre-training stage are loaded into the packet-level attention module and the flow-level attention module, which are used to extract local features within the data packet and global features of the traffic session respectively. In order to accelerate convergence and compress the model size at the same time, a parameter sharing strategy is implemented between the packet-level encoder and the flow-level encoder;
[0092] (5.3) In the classification stage, each block in the MFM matrix will undergo two-stage pooling operations of row pooling and column pooling for feature fusion to generate a global classification feature vector, and the output vector of the temporal convolutional module is concatenated to obtain the final classification vector. This vector is input into the MLP after being flattened, and finally the predicted distribution of the encrypted malicious trojan traffic category is output C is the total number of categories. The classification loss is obtained by calculating the cross-entropy between the predicted distribution and the true label, and the calculation formula is as follows:
[0093]
[0094] Among them, represents the predicted output of the model, indicating the predicted probability that the sample belongs to each category, and y is the true label of the sample.
[0095] In order to study the effectiveness of different masking ratios for encrypted malicious trojan traffic detection, the F1-score of the model for identifying encrypted malicious trojan traffic under different masking ratios is considered respectively as Figure 4 shown.
[0096] Overall, appropriately increasing the masking ratio helps improve the model performance. However, an excessively high masking ratio makes the decoder reconstruction task too difficult, leading to a rapid decline in the F1 value, indicating that the masking strategy needs to balance representation learning and task feasibility. Experiments show that when the masking ratio is 75%, the model performance can reach the optimal. A high masking ratio indicates that there is a large amount of information redundancy in encrypted traffic bytes, and the encrypted traffic classification task itself does not rely on a deep understanding of the data content. In the encrypted scenario, semantic-level parsing of the encrypted payload is not possible. The multi-level MFM matrix based on the image structure proposed in this paper has a more reasonable data representation method and can more effectively capture the characteristics of encrypted traffic data, further demonstrating the effectiveness of the design of the method in this paper.
[0097] To verify the effectiveness of the present invention in detecting encrypted malicious trojan traffic, the method of the present invention is compared with the following encrypted traffic classification models: (1) FlowPrint, a semi-supervised learning-based mobile application fingerprint generation method; (2) FS-Net, an end-to-end encrypted traffic classification method based on multi-layer bidirectional gated recurrent units; (3) ET-BERT, an encrypted traffic classification method based on a pre-trained framework; (4) PERT, an encrypted traffic classification method based on word embedding. The results are shown in Table 1 below.
[0098] Table 1 Comparative experiment results
[0099] Method Accuracy Precision Recall <![CDATA[F1]]> FlowPrint 86.80% 68.15% 50.02% 54.00% FS-Net 72.05% 75.02% 72.38% 71.31% ET-BERT 91.89% 89.23% 91.76% 90.28% PERT 90.37% 89.49% 88.92% 88.26% ETC-MAE 99.78% 98.93% 99.51% 99.26%
[0100] The experimental results show that ETC-MAE can effectively identify different encrypted malicious trojans, obtain excellent performance with fewer labeled data samples, and exhibit strong generalization ability.
[0101] Compared with the global attention mechanism, the packet-level and flow-level hierarchical attention architecture adopted in the present invention not only reduces the computational complexity but also achieves better results in classification performance, verifying the effectiveness of the packet-level and flow-level attention mechanisms. In addition, the present invention integrates multi-modal features, uses TCN to explicitly model the temporal relationship between packets, captures the traffic temporal behavior characteristics, and significantly improves the model performance. The experimental results are as Figure 5 shown. It can be seen that when the TCN module is removed, the classification performance of the model significantly decreases.
[0102] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: the specific implementation manners of the present invention can still be modified or equivalently replaced, and any modification or equivalent replacement without departing from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.
Claims
1. A method for detecting encrypted malicious Trojan traffic based on a masked autoencoder and multi-level traffic modeling, characterized in that It includes the following steps: (1) Convert the encrypted malicious trojan traffic raw bytes into a multi-level traffic modeling matrix, which directly reflects the byte distribution of the traffic, and extract the arrival time interval of the data packets and the sequence characteristics of the data packet lengths to explicitly model the temporal dependencies between the data packets; (2) Split the matrix constructed in step (1) into non-overlapping small blocks, each small block is mapped into a vector, and position encoding information is added to be used as the input of the encoder; (3) Adopt a hierarchical attention mechanism to capture the dependencies within the data packets and between the data packets by using the packet-level attention mechanism and the flow-level attention mechanism respectively; (4) Integrate the temporal convolutional network to explicitly model the traffic behavior patterns contained in the data packet time interval and the cumulative data packet length, and capture the temporal dependencies between the data packets; (5) Adopt a self-supervised pre-training strategy based on the masked autoencoder, use a large amount of unlabeled data to train the encoder, and fine-tune with a small amount of trojan traffic label data in the downstream task to achieve efficient encrypted malicious trojan traffic detection.
2. The encrypted malicious trojan traffic detection method based on a masked autoencoder and multi-level traffic modeling according to claim 1, wherein The specific steps of step (1) include the following sub-steps: (1.1) First, perform flow partitioning on the original traffic according to the five-tuple information of the source IP address, destination IP address, source port, destination port, and protocol type. At the same time, remove the Ethernet header to strip the link layer features, set the port number to zero to avoid traffic fingerprint interference, and adopt a randomized IP address replacement strategy to retain the traffic directionality as the identifier of the uplink and downlink traffic; (1.2) Extract the arrival time interval sequence {IPT1, IPT2, …, IPT M} and the packet length sequence {±pkt_length1, ±pkt_length2, …, ±pkt_length M}, where IPT M represents the relative arrival time interval between the Mth packet and the previous packet, and ±pkt_length M represents the length of the Mth packet, and the directionality is retained, where the packets sent from the victim to the attacker are recorded as positive, and the packets sent from the attacker to the victim are recorded as negative; (1.3) Extract the first M consecutive data packets, and construct a two-dimensional matrix MFM with a size of H×W through three-level feature abstractions at the byte level, packet level, and flow level.
3. The encrypted malicious trojan traffic detection method based on masked autoencoder and multi-level traffic modeling according to claim 2, wherein Perform logarithmic normalization on the arrival time interval sequence in step (1.2), and perform standardization on the data packet length sequence to control the numerical range.
4. The encrypted malicious Trojan traffic detection method based on masked autoencoder and multi-level traffic modeling according to claim 2, wherein, The specific operation of step (1.3) is as follows: Each data packet is divided into two parts: a header and a payload. Each row in the matrix corresponds to a type of byte, which is divided into a header row or a payload row, and the original byte value is used as the initial feature to prevent semantic loss. The header row retains the network layer, transport layer, and extensible header information, and the payload row stores the original byte stream of the application layer. When the length exceeds 240 bytes, tail truncation is performed. A single data packet is constructed into a packet-level matrix with a size of H / M*W, and M packet-level matrices are stacked along the second dimension, i.e., the column direction, to form a complete flow representation matrix MFM. When the actual byte number of the data packet is insufficient, a zero-padding strategy is adopted to ensure dimension consistency.
5. The encrypted malicious Trojan traffic detection method based on masked autoencoders and multi-level traffic modeling according to claim 1, characterized in that, In step (2), splitting the matrix constructed in step (1) into non-overlapping small blocks, each small block is mapped into a vector, and position encoding information is added to be used as the input of the encoder, specifically including the following sub-steps: (2.1) Divide the MFM matrix into non-overlapping 2D patches of size P×P, obtaining N = HW / P 2 patches, denoted as and map these patches into D-dimensional embedding vectors through a linear layer; (2.2) To retain the position information, add the position encoding information to the block embedding as the input of the encoder. The position encoding uses common sine and cosine functions, and its calculation method is as follows: PE(pos, 2i) = sin(pos / 10000 2i / d ) PE(pos, 2i + 1) = cos(pos / 10000 2i / d ) Among them, PE(pos, 2i) and PE(pos, 2i + 1) respectively represent the position encoding vectors at the pos-th position of the i-th feature dimension in the current dimension. 2i and 2i + 1 correspond to the even and odd dimensions respectively. represents the i-th block. represents a learnable linear transformation matrix used to map the block to a D-dimensional embedding vector, E pos ∈ R N×D represents the position encoding matrix, x e ∈ R N×D represents the final input sequence with position encoding, which is used as the input for the subsequent encoder.
6. The encrypted malicious Trojan traffic detection method based on masked autoencoder and multi-level traffic modeling according to claim 1, characterized in that The specific steps of step (3) include the following sub-steps: (3.1) The packet-level attention mechanism part uses a traffic encoder based on Vision Transformer to achieve in-packet feature interaction. Through the multi-head attention mechanism, each block in the data packet can achieve dynamic interaction based on the correlation weights of different attention heads. Its calculation process is expressed as Concat(head1, head2, …, headn,), where each attention head is calculated through an attention function: Q = x l W Q , K = x l W K , V = x l W V where x l ∈R N×D represents the input of the l-th layer, and Q, K, and V represent the Query, Key, and Value matrices in the attention mechanism respectively. W Q , W K , are learnable parameter matrices used to project the input into different spaces. D k is the dimension of the attention head. QK T represents the dot product of Query and Key, measuring the attention correlation. (3.2) The flow-level attention mechanism realizes the adaptive transformation of feature granularity by performing row pooling on the features output by the packet-level attention mechanism x r = RowPooling(x′ p ), generates row-pooled output features by performing mean pooling row by row, generates row-level blocks representing headers and payloads, and these row-level blocks are fed into the encoder as new inputs. Then, multi-head attention is used to extract the relationships between packets. After that, all row-level features are aggregated through column pooling to obtain the final representation x MFM ∈R D .
7. The encrypted malicious Trojan traffic detection method based on masked autoencoder and multi-level traffic modeling according to claim 1, wherein The temporal convolutional module in step (4) consists of multiple layers of one-dimensional dilated convolutions. Each layer contains LayerNorm normalization, causal convolution, and the non-linear activation function ReLU. Finally, residual connections are used to stabilize the training of the deep network. After this process, average pooling is performed on the module output along the time dimension to obtain global temporal features, which are projected through a fully connected layer to align with the encoder output.
8. The encrypted malicious trojan traffic detection method based on masked autoencoders and multi-level traffic modeling according to claim 1, characterized in that, Step (5) includes the following sub-steps: (5.1) In the pre-training stage, a masked autoencoder is used to construct an asymmetric encoder-decoder architecture to reconstruct the original byte data of the MFM matrix. Specifically, a random masking operation is performed on the MFM matrix, and a small number of unmasked blocks are retained as the input to the encoder. The encoder extracts features from these visible blocks and outputs encoded tokens. Then, a lightweight decoder is used to combine the encoded tokens and masked tokens to reconstruct the covered area of the MFM matrix. Optimization is performed by calculating the mean squared error between the true value and the reconstructed value. The calculation formula is as follows: L rec = MSE(y real , y rec ) Among them, L rec represents the reconstruction loss, MSE represents the mean square error function, and y real represents the true block content masked in the original MFM matrix, and y rec represents the output of the decoder, that is, the reconstructed block content. (5.2) In the fine-tuning stage, the encoder parameters obtained in the pre-training stage are loaded into the packet-level attention module and the flow-level attention module, which are used to extract local features within the data packet and global features of the traffic session respectively. To accelerate convergence and compress the model size simultaneously, a parameter sharing strategy is implemented between the packet-level encoder and the flow-level encoder; (5.3) Classification stage: Each block in the MFM matrix undergoes two-stage pooling operations, namely row pooling and column pooling, for feature fusion to generate a global classification feature vector. The output vector of the temporal convolutional module is concatenated to obtain the final classification vector, which is flattened and then input into the MLP, and finally the predicted distribution of encrypted malicious trojan traffic categories is output. C is the total number of categories. The classification loss is obtained by calculating the cross-entropy between the predicted distribution and the true label, and the calculation formula is as follows: Among them, represents the predicted output of the model, indicating the predicted probability that the sample belongs to each category, and y is the true label of the sample.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the encrypted malicious trojan traffic detection method based on masked autoencoder and multi-level traffic modeling as described in any one of claims 1 to 8 above.
10. A computer-readable storage medium having computer instructions stored thereon, characterized in that, When the computer instruction is executed by the processor, it implements the encrypted malicious trojan traffic detection method based on masked autoencoder and multi-level traffic modeling as described in any one of claims 1 - 8.
Citation Information
Cited By
Traffic generation method and device
CN120935037A
Network traffic classification method and system based on dual position coding and mixed mask mechanism
CN121547281A
A network traffic classification method and system based on double position encoding and hybrid mask mechanism
CN121547281B
Encrypted traffic classification method and system based on cross-modal comparative learning, and medium
CN121727866A
DoH tunnel detection method based on feature fusion and large language model
CN121864426A