Encryption Traffic Threat Detection Method and Device Based on Multi-Stream Enhanced Single-Stream Representation
By using multi-stream enhanced single-stream characterization method in encrypted traffic detection, combined with pre-training strategies for comparative learning and stream consistency discrimination tasks, the problem of dependence on labeled data in the prior art is solved, and the generalization ability and detection performance of the model are improved.
Patent Information
- Application Number
- CN202510296252.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-03-13
AI Technical Summary
The existing encrypted traffic detection model mainly relies on supervised learning, requires a large amount of labeled data, and the labeling process is time-consuming and labor-intensive, making it difficult to effectively utilize unlabeled data, resulting in insufficient generalization capabilities in new tasks, new fields, and new scenarios.
The encrypted traffic threat detection method based on multi-stream enhanced single-stream characterization is adopted. By setting up proxy tasks in the labelless data set for pre-training, including comparing learning tasks and stream consistency discrimination tasks, using similar flows in multiple streams to enhance single-stream characterization, and fine-tuning through a small amount of labeled data to optimize the classification performance of the model.
Effectively use unlabeled data to learn traffic characterization, improve the generalization ability of the model, reduce dependence on labeled data, and improve detection performance in new tasks, new fields, and new scenarios.
Smart Images

Figure CN119788444B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of malicious traffic detection, and in particular to an encrypted traffic threat detection method and device based on multi-flow enhanced single-flow characterization. Background Art
[0002] As a key supporting force for the booming development of the digital economy, the healthy and stable operation of the information network is premised on the accurate classification and proper management of network traffic. In today's era, with the continuous popularization of the network and the increasingly rich and diverse application scenarios, the complexity of network traffic has increased significantly. In the field of malicious traffic detection, the current malicious traffic shows a sharp growth trend, which has already posed a serious threat to network security and stability, and also brought extremely severe challenges to traditional traffic detection and classification methods.
[0003] Currently, most models based on supervised deep learning highly rely on a large amount of labeled data. However, the labeling process of network traffic usually requires the expertise of security experts, which is time-consuming and labor-intensive, and the labeling cost is high. For example, in order to collect malicious traffic data, researchers usually run malicious software in a sandbox environment to generate corresponding malicious traffic samples. This method requires building and maintaining a highly simulated virtual running environment and faces complex traffic capture and cleaning tasks. At the same time, the characteristics of malicious traffic may change dynamically over time, making it difficult for existing labeled data to maintain effectiveness in the long term.
[0004] In this context, self-supervised learning technology has gradually become a promising alternative. Self-supervised learning refers to learning by using the data's own information, and solving the problem of insufficient labeled data in supervised learning by learning feature representations from unlabeled data. Its core idea is to learn shallow representations from unlabeled data by setting proxy tasks, and finally fine-tune with labeled data to make full use of the data. In the field of computer vision, self-supervised learning methods construct positive samples by performing data enhancement (such as random cropping and color jitter) on samples. Positive samples are different observation angles of the original image samples, and their semantics are similar, while other samples are negative samples. The shallow representation of the image is learned by pulling the distance of the positive samples closer and the distance of the negative samples farther away. In the field of natural language processing, taking BERT as an example, it learns more in-depth contextual semantic information through MLM (random mask, predicting masked words through context information) and NSP (neighboring sentence prediction, by judging whether a pair of sentences are adjacent context sentences). Most of the existing encrypted traffic detection models are supervised, and the labeled data is often manually labeled, which requires a lot of time and manpower costs, while there is a large amount of unlabeled data in the real world. In addition, for some fields, it is difficult to obtain a sufficient amount of labeled data, making it difficult for supervised models to achieve good results in these fields. In addition, supervised models need to collect labeled data again and retrain the model when dealing with new tasks, new fields, and new scenarios, which increases the training cost and time cost, and may face the problem of insufficient model generalization ability. In this context, self-supervised learning technology has gradually become a promising alternative.
[0005] Most existing self-supervised learning-based models extract features by masking traffic data and predicting the masked parts. However, this method fails to fully explore the characteristics of network traffic in multiple dimensions such as time series correlation and behavior patterns, thus limiting the improvement of its detection performance.
[0006] In summary, how to set up proxy tasks in combination with network traffic characteristics to make full use of a large amount of unlabeled encrypted traffic to learn traffic representation and make the model have strong generalization is an important issue.
[0007] The existing methods for classifying encrypted traffic are summarized as follows:
[0008] (1) Machine learning-based methods
[0009] Machine learning-based methods mainly extract some features according to experts' prior knowledge and use machine learning algorithms for malicious traffic identification. Generally, these features are designed based on the behavior of the session. The way of classification according to behavior is mainly based on the following idea: different attacks usually have their own unique characteristics [1]. Taking the DoS attack as an example, it has two sub-categories, bandwidth consumption and resource consumption attacks, and each category has many attack methods. Taking the SYN Flood attack of bandwidth consumption attack as an example, it mainly takes advantage of the vulnerability of TCP handshake. The attacker sends a large number of SYN packets to the target server. Usually, the attacker will use a spoofed IP address, and then the server responds to each connection request and leaves the port open for a period of time to receive the corresponding packets. If the attacker sends too many SYN packets, each new SYN packet will cause the server to open the port for a period of time. Once all ports are in use, the server cannot work properly. Usually, to detect this type of attack, it can be judged by the duration and the flag field, with only SYN but no further communication.
[0010] The implementation of the machine learning-based method is relatively simple. Only the features need to be specified, and then classic algorithms can be used for processing after extraction. However, there are also big problems at the same time. First of all, in terms of feature selection, obviously there is no general feature group that can distinguish the current task. Different features need to be designed for different tasks, and since the types of malicious traffic in this field are constantly evolving, the features may also need to be continuously added and modified.
[0011] (2) Supervised deep learning-based methods
[0012] Deep-Packet[2] classifies the traffic from the Packet Level, processes each Packet into a one-dimensional vector. Considering the differences in the traffic collection environment, specifically, the data in a closed environment and the data in the open environment, i.e., the live network data, may have huge differences. The dataset collected in a closed manner is collected under limited servers and hosts. Therefore, in the processing of data packets, first, the Ethernet header is removed, the IP address is set to zero to prevent the model from being affected by irrelevant features, the UDP data header is aligned according to the TCP packet length, and finally, it is aligned according to the maximum transmission unit (MTU) of 1500 bytes. If it is not enough, padding with 0 is performed. The processed one-dimensional vectors are used by CNN and Stacked Autoencoder respectively to learn features. However, for the Packet Level, the information that a single Packet can carry is too little. We believe that based on the FlowLevel or even multiple flows, the information of the session can be better utilized. Wang opened the door to deep learning in encrypted traffic classification in 2016 with a very simple LeNet network [3]. This is the first paper to perform representation learning on the original traffic data using a deep neural network. The traffic is divided according to the five-tuple (source IP address, destination IP address, source port number, destination port number, protocol). The first 784 bytes in the flow are processed into a 2D grayscale image. For the processed data, similar to the handwritten MNIST digit recognition, the traffic is classified. In the subsequent work on end-to-end encrypted traffic based on one-dimensional CNN [4], Wang believes that network traffic is essentially a kind of time-series data, which is a one-dimensional byte stream organized in a hierarchical structure, from the bottom layer to the top layer, namely bytes, frames, sessions, and the entire traffic. Similar to the previous work, the first 784 bytes in the session are taken, except that in this work, the network uses one-dimensional CNN. The method based on supervised learning requires a large amount of labeled datasets and is not applicable in scenarios where the number of labeled data samples is small. Summary of the Invention
[0013] To solve the technical problem in the prior art of how to set proxy tasks in combination with network traffic characteristics to make full use of a large amount of unlabeled encrypted traffic to learn the representation of traffic, an embodiment of the present invention provides an encrypted traffic threat detection method and device based on multi-flow enhanced single-flow representation. The technical solution is as follows:
[0014] On the one hand, an encrypted traffic threat detection method based on multi-flow enhanced single-flow representation is provided, characterized in that the method includes:
[0015] S1. Collect traffic data samples and construct the traffic data samples into a traffic dataset;
[0016] S2. Perform data preprocessing on the traffic dataset to obtain training samples;
[0017] S3. Divide the training samples into samples for the pre-training stage and samples for the fine-tuning stage; for the samples in the pre-training stage, organize multiple streams according to quadruples to enrich the selection of positive samples in contrastive learning; for the samples in the fine-tuning stage, organize a single stream according to quintuples. For each single stream, extract the byte information of the first n data packets in the stream and the length sequence information of the first m data packets.
[0018] S4. Perform pre-training on the unlabeled dataset by setting proxy tasks to obtain a pre-trained model; among them, the proxy tasks include a contrastive learning task and a flow consistency discrimination task.
[0019] S5. Based on the pre-trained model, perform fine-tuning through a small amount of labeled malicious traffic datasets. During the training process, use the cross-loss entropy function to calculate the loss, and update the gradient according to the loss to obtain an updated model.
[0020] S6. For the updated model, perform malicious traffic detection and identification to complete the encrypted traffic threat detection based on multi-stream enhanced single-stream representation.
[0021] Optionally, in S2, perform data preprocessing on multiple pairs of traffic datasets, including:
[0022] Remove invalid traffic;
[0023] Remove fields irrelevant to traffic detection and set them to 0; among them, the irrelevant fields include: checksum, IP address, port number.
[0024] Optionally, the quadruple includes: source IP address, destination IP address, destination port number, protocol number;
[0025] The quintuple includes: source IP address, source port number, destination IP address, destination port number, protocol number.
[0026] Optionally, in S4, the contrastive learning task includes:
[0027] Split the unlabeled dataset into Traces (i.e., multiple streams) according to the quadruple {source IP, destination IP, destination port, protocol};
[0028] Construct a positive and negative sample sampling method to sample the positive and negative sample pairs required in contrastive learning;
[0029] Through the contrastive learning loss function, pull the similar streams within the multiple streams closer in the feature space, and at the same time pull the streams not in the same Trace farther apart in the vector space;
[0030] Among them, the selection of positive samples in contrastive learning is a single stream in the multiple streams, and data augmentation is performed on the single stream. The negative samples are maintained by different batches of samples through a sliding queue.
[0031] Optimize the loss by calculating with the InfoNCE loss function, reducing the distance between positive samples and increasing the distance between negative samples in the vector space.
[0032] Optionally, construct positive and negative sample sampling methods to sample the positive and negative sample pairs required in contrastive learning, including:
[0033] The pre-training data consists of triples where x and are flow samples randomly selected from the same trace, is the negative sample maintained from the dictionary queue;
[0034] Among them, for a batch of samples, similar flows in multiple flows are used as positive samples through data augmentation, and negative samples are obtained through maintaining the dictionary queue.
[0035] Optionally, in S4, the flow consistency discrimination task includes:
[0036] In the given dataset, for single-flow data, construct contexts by randomly pairing the flows;
[0037] Shuffle the data according to the preset probability p;
[0038] When the context information in the combined flows does not match, mark it as "1", indicating that the data has been confused; conversely, when the context information remains consistent, mark it as "0", indicating that the data is intact flow;
[0039] Use the cross-loss entropy function for optimization:
[0040]
[0041] where H represents the cross-entropy loss function; represents the value predicted by the model;
[0042] The overall loss function is composed of the sum of two parts of losses, jointly optimizing the learning process of the model as shown in the following formula:
[0043]
[0044] Among them, represents the loss function of the contrastive learning task.
[0045] Optionally, shuffle the data according to the preset probability p, including:
[0046] The preset probability p = 0.5; that is, half of the data is shuffled and half of the data remains unchanged according to the probability;
[0047] For a given piece of data, randomly generate a probability p_. If p_ > p, shuffle the data; otherwise, keep the data unchanged.
[0048] Optionally, in S5, based on a pre-trained model, fine-tune it with a small amount of labeled malicious traffic data sets. During the training process, use the cross-entropy loss function to calculate the loss, and update the gradient according to the loss to obtain an updated model, including:
[0049] When the feature quality obtained from pre-training is high and there is an obvious distribution shift, select linear probing to fine-tune the model; when the data set is sufficient and the graphics card resources are sufficient, fine-tune the model with all parameters.
[0050] Among them, fine-tuning includes using the cross-entropy loss function to fine-tune the model to optimize the classification performance of downstream tasks:
[0051]
[0052] where y represents the true label, represents the predicted distribution.
[0053] On the other hand, a cryptographic traffic threat detection device based on multi-flow enhanced single-flow representation is provided. This device is applied to the cryptographic traffic threat detection method based on multi-flow enhanced single-flow representation. The device includes:
[0054] A data acquisition module for collecting traffic data samples and constructing the traffic data samples into a traffic data set;
[0055] A data preprocessing module for preprocessing the traffic data set to obtain training samples;
[0056] A traffic embedding module for dividing the training samples into samples in the pre-training stage and samples in the fine-tuning stage; for the samples in the pre-training stage, organize multi-flows according to quadruples to enrich the selection of positive samples in contrastive learning; for the samples in the fine-tuning stage, organize single-flows according to quintuples. For each single-flow, extract the byte information of the first n data packets in the flow and the length sequence information of the first m data packets.
[0057] A pre-training module for pre-training in an unlabeled data set by setting proxy tasks to obtain a pre-trained model; among them, the proxy tasks include a contrastive learning task and a flow consistency discrimination task;
[0058] A fine-tuning module for fine-tuning based on the pre-trained model with a small amount of labeled malicious traffic data sets. During the training process, use the cross-entropy loss function to calculate the loss, and update the gradient according to the loss to obtain an updated model;
[0059] A threat detection module for detecting and identifying malicious traffic through an updated model, and completing encrypted traffic threat detection based on multi-flow enhanced single-flow representation.
[0060] On the other hand, there is provided an encrypted traffic threat detection device based on multi-flow enhanced single-flow representation. The encrypted traffic threat detection device based on multi-flow enhanced single-flow representation includes: a processor; a memory, on which computer-readable instructions are stored. When the computer-readable instructions are executed by the processor, any one of the methods in the above-mentioned encrypted traffic threat detection method based on multi-flow enhanced single-flow representation is implemented.
[0061] On the other hand, there is provided a computer-readable storage medium, in which at least one instruction is stored. The at least one instruction is loaded and executed by a processor to implement any one of the methods in the above-mentioned encrypted traffic threat detection method based on multi-flow enhanced single-flow representation.
[0062] The beneficial effects brought by the technical solutions provided in the embodiments of the present invention at least include:
[0063] In the embodiments of the present invention, similar flows in multi-flows are used to enhance the representation of single-flows. In addition, the present invention makes the model aware of the consistency of traffic, and the data packets in the flow have context relevance, so that the model can distinguish whether the data packets in the flow come from the same flow. The present invention splits a small amount of labeled traffic for specific tasks into flows as inputs, and uses a fine-tuning strategy to adjust the pre-trained model to make full use of the capabilities of the pre-trained model. The pre-trained model provided by the present invention, compared with supervised learning, can quickly adapt to downstream malicious detection tasks with only a small number of labeled samples during fine-tuning; the pre-trained model can learn representations independent of labels and can be transferred to different tasks, such as traffic application identification, malicious traffic detection, etc. Description of the Drawings
[0064] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0065] Figure 1 It is a schematic flowchart of an encrypted traffic threat detection method based on multi-flow enhanced single-flow representation provided by an embodiment of the present invention;
[0066] Figure 2 It is a form diagram of traffic in a network provided by an embodiment of the present invention;
[0067] Figure 3 It is a diagram of the situation of multi-flows in an application provided by an embodiment of the present invention;
[0068] Figure 4 This is the pre-training framework diagram provided by the embodiments of the present invention;
[0069] Figure 5 This is the framework diagram of the fine-tuning stage provided by the embodiments of the present invention;
[0070] Figure 6 This is the model encoding structure diagram provided by the embodiments of the present invention;
[0071] Figure 7 This is the block diagram of the encrypted traffic threat detection device based on multi-stream enhanced single-stream representation provided by the embodiments of the present invention;
[0072] Figure 8 This is the structural schematic diagram of the electronic device provided by the embodiments of the present invention. Detailed implementation manners
[0073] The following describes the technical solutions in the present invention with reference to the accompanying drawings.
[0074] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as an "example" in the present invention should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, the use of the word "example" is intended to present concepts in a specific manner. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one of the two.
[0075] In the embodiments of the present invention, sometimes subscripts such as W 1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning to be expressed is the same.
[0076] To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.
[0077] The embodiments of the present invention provide an encrypted traffic threat detection method based on multi-stream enhanced single-stream representation. This method can be implemented by an encrypted traffic threat detection device based on multi-stream enhanced single-stream representation. The encrypted traffic threat detection device based on multi-stream enhanced single-stream representation can be a terminal or a server. As Figure 1 shown in the flowchart of the encrypted traffic threat detection method based on multi-stream enhanced single-stream representation, as Figure 1 shown, the encrypted traffic threat detection method proposed by the present invention, the processing flow of this method can include the following steps:
[0078] S1. Collect traffic data samples and construct the traffic data samples into a traffic data set;
[0079] S2. Perform data preprocessing on the traffic dataset to obtain training samples;
[0080] In a feasible implementation manner, in S2, the multi - pair traffic dataset is subjected to data preprocessing, including:
[0081] Remove invalid traffic;
[0082] Remove fields irrelevant to traffic detection and set them to 0; among them, the irrelevant fields include: checksum, IP address, port number.
[0083] S3. Divide the training samples into samples for the pre - training stage and samples for the fine - tuning stage; for the samples in the pre - training stage, organize multi - flows according to quadruples to enrich the selection of positive samples in contrastive learning; for the samples in the fine - tuning stage, organize single - flows according to quintuples. For each single - flow, extract the byte information of the first n data packets and the length sequence information of the first m data packets in the flow, and perform embedding representation on the byte information and length sequence information of the traffic.
[0084] In a feasible implementation manner, the quadruple includes: source IP address, destination IP address, destination port number, protocol number;
[0085] The quintuple includes: source IP address, source port number, destination IP address, destination port number, protocol number.
[0086] In a feasible implementation manner, as Figure 2 shown, the traffic in the network mainly originates from the communication behavior between clients or between a client and a server. During the entire communication process, data packets need to be transmitted between two nodes and may pass through multiple intermediate network devices, such as routers and switches. And a single - flow is a basic unit in network traffic classification, which is used to represent a set of data packets identified by a quintuple (source IP address, source port number, destination IP address, destination port number, protocol number) with the same specific characteristics within a certain time interval. It represents the communication between an application (identified by the source port) in a source host (identified by the source IP address) and another application in another host (identified by the destination IP address) (identified by the destination port number). Here, the quintuple can be regarded as the unique identifier of network communication, ensuring that each single - flow represents a set of related communication data packets.
[0087] For multi - flows, it is a collection of single - flows, as Figure 3As shown, when a user accesses an application, multiple associated flows in space and time are generated. For example, during the process of accessing a page, there may be multiple parallel flows to accelerate the access speed of page pictures, or after a certain connection is disconnected, a connection with the application on the destination server is quickly re-established. This form of multiple flows can be defined by a quadruple {source IP address, destination IP address, destination port number, protocol number}. Each time the application re-establishes a connection, its destination port number is often determined, while the source port is randomly assigned by the system. Therefore, we use the quadruple within a certain period of time as described above to identify similar single flows.
[0088] In a feasible implementation, the present invention uses similar flows in multiple flows to enhance the representation of single flows. In addition, it can also make the model aware of the consistency of traffic, and the data packets in the flow have context relevance. For this reason, we separately designed a task to let the model distinguish whether the data packets in the flow come from the same flow. This solution includes two stages: pre-training and fine-tuning. As Figure 4 shown, there are two tasks in the pre-training stage. The first task is a contrastive learning task, which allows the model to learn the similarity between flows. Specifically, the unlabeled dataset is split into Traces (i.e., multiple flows) according to the quadruple {source IP, destination IP, destination port, protocol}, and the similar flows within the multiple flows are pulled closer in the feature space through a contrastive learning loss function, while the flows not in the same Trace are pulled apart in the vector space. The second task is a binary classification task to let the model learn the consistency of traffic. Specifically, randomly scramble the context of the flow and let the model determine whether the flow is scrambled. In the fine-tuning stage, this chapter splits a small amount of labeled traffic of specific tasks into flows as inputs and uses three fine-tuning strategies to adjust the pre-trained model to make full use of the capabilities of the pre-trained model. There are a total of three modules: a traffic embedding module, a pre-training module, and a fine-tuning module. Specifically as follows:
[0089] Traffic embedding module:
[0090] For a specific TLS flow, the first n data packets in the flow may contain some plaintext information of the TLS handshake process. After the TLS handshake negotiates the key, the payload part in the data packet has been encrypted. However, sequence features such as the data packet sequence length can provide additional information. To optimize the computational cost, one strategy is to combine the payload information of the first n data packets of the flow with the sequence information of the accompanying flow, rather than processing the entire original flow traffic. This method can more efficiently utilize the original traffic information and reduce the computational overhead associated with processing a large amount of data. The overall structure of this process is as Figure 6 shown.
[0091] Traffic data processing: For the information of the original traffic data, a multi-level traffic representation method is adopted, which combines the first 5 data packets in the flow matrix. The header data of each data packet is represented by 2 rows of header data, a total of 80 bytes, and 6 rows are allocated for the payload part to store its payload content. Subsequently, the Traffic Transformer model is used to encode the original traffic data.
[0092] Packet length sequence feature extraction: Extract the packet length sequence of the flow. For example, {64, 128, 1040, 512...}. Then, through the parameter matrix perform word embedding on this sequence to transform the discrete input sequence into a high-dimensional vector representation. In addition, this method incorporates positional encoding into each Token in the sequence to represent its position information in the sequence. After that, multiple Transformer modules are used to encode this sequence.
[0093] Feature vector fusion module: After obtaining the embeddings of the original traffic data and the packet length sequence, this chapter combines these two embedding vectors through concatenation. Then, the combined vector is input into a linear layer to generate the final traffic embedding representation.
[0094] S4. Perform pre-training on the unlabeled dataset by setting proxy tasks to obtain a pre-trained model; among them, the proxy tasks include contrast learning tasks and flow consistency discrimination tasks. As Figure 4 shown.
[0095] In a feasible implementation, a detailed analysis is carried out on how to enhance the flow embedding through trace using contrast learning. The inspiration for contrast learning comes from the human ability to recognize the commonalities in similar things and the differences between different entities. Its goal can be described as training an encoder to make the enhanced versions of the same sample closer in the embedding space, while separating the embeddings of different samples from each other. To achieve this goal, this scheme designs a positive and negative sample pair sampling strategy and alleviates the label bias introduced by data distribution and label information through contrast learning.
[0096] In a feasible implementation, in S4, the contrast learning task includes:
[0097] Split the unlabeled dataset into Traces, that is, multiple flows, according to the quadruple {source IP, destination IP, destination port, protocol};
[0098] Construct a positive and negative sample sampling method to sample the positive and negative sample pairs required in contrast learning;
[0099] Pull the similar flows within the multi-flow closer in the feature space through the contrast learning loss function, while separating the flows not in the same Trace in the vector space;
[0100] Among them, the selection of positive samples in contrastive learning is a single stream in the multi-stream, and data augmentation is performed on the single stream, while negative samples are samples from different batches maintained through a sliding queue;
[0101] The loss is optimized by calculating the loss through the InfoNCE loss function, pulling the distances of positive samples closer and the distances of negative samples farther in the vector space.
[0102] In a feasible implementation, a positive and negative sample sampling method is constructed to sample the positive and negative sample pairs required in contrastive learning, including:
[0103] The pre-training data consists of triples where x and are stream samples randomly selected from the same trace, is a negative sample maintained from the dictionary queue;
[0104] Among them, for a batch of samples, similar streams in the multi-stream are used as positive samples through data augmentation, and negative samples are obtained through maintaining the dictionary queue. After each batch of data is trained, it enters the queue, and the batch of data that entered the queue earliest, that is, the data with lower freshness, is popped out of the queue.
[0105] In a feasible implementation, this solution adopts a self-supervised learning (Momentum Contrast, MoCo) momentum contrast learning method to extract robust stream features from traces. MoCo optimizes the model on a dynamically updated negative sample set by modeling contrastive learning as a dictionary query task. Its core mechanism is to store negative samples through a dynamic dictionary maintained by a queue, thereby ensuring the comprehensiveness and diversity of negative samples while enabling efficient calculation. During the training process, the dynamic dictionary queue updates the current batch of data and removes the earliest batch of negative samples to form a dynamic, finite-sized negative sample pool. In addition, MoCo adopts a momentum encoder to maintain the consistency of training. The advantage of this momentum update mechanism is that it does not require an additional encoder to be trained separately for negative samples, but is gradually updated through the main encoder, thereby reducing the computational burden and enhancing the stability of negative sample representations. The momentum encoder mechanism in MoCo has been proven in practice to significantly improve the generalization ability of the model and be adaptable to complex downstream task scenarios.
[0106] In a feasible implementation, when the present application uses the MOCO method and the momentum encoder, it combines specific application scenarios and makes adaptive modifications to them. For example, in contrastive learning, it is necessary to construct positive and negative samples, especially the construction of positive samples. In computer vision, data augmentation (such as image flipping, Gaussian noise) can be used to augment pictures, but it cannot be directly applied to the traffic field. Therefore, the present invention uses multiple streams to mine similar traffic in the dataset as positive samples.
[0107] In a feasible implementation, the projection layer and the contrast loss function: The projection layer realizes non-linear projection through a multi-layer perceptron with a hidden layer. During the pre-training process, it is usually more important to retain the features in the representation before the projection layer that are closely related to the pre-training task objective, so as to establish a connection between the representation and the contrast loss. For the contrast loss used to measure the similarity of sample pairs in the representation space, this chapter uses the Information Noise Contrastive Estimation (InfoNCE) loss function, and its formula is as follows:
[0108] ;
[0109] Among them, is the temperature hyperparameter, used to control the shape of the logits distribution, is the contrast loss of the given triple The size of the set
[0110] is equal to the size of the dictionary queue.
[0111] In a feasible implementation, in S4, the flow consistency discrimination task includes:
[0112] In the given dataset, for single-flow data, the flow is randomly paired to construct context;
[0113] Shuffle the data according to the preset probability p;
[0114] When the context information in the combined flow does not match, mark it as "1", indicating that the data has been confused; conversely, when the context information remains consistent, mark it as "0", indicating that the data is a complete flow;
[0115]
[0116] This method essentially constructs a binary classification task, aiming to improve the model's learning ability of flow consistency. Use the cross-entropy loss function to optimize the pre-trained model:
[0117] Among them, H represents the cross-entropy loss function, represents the value predicted by the model;
[0117] The overall loss function is composed of the sum of two parts of losses, and the joint optimization model has a learning process as shown in the following formula:
[0118]
[0119] where represents the loss function of the contrastive learning task.
[0120] In a feasible implementation, shuffling the data according to a preset probability p includes:
[0121] The preset probability p = 0.5; that is, half of the data is shuffled with probability, and half of the data remains unchanged;
[0122] For a given data, a probability p_ is randomly generated. If p_ > p, it is shuffled, otherwise the data remains unchanged.
[0123] In a feasible implementation, for example, for a given flow, it consists of some ordered data packets, [data packet 1_1, data packet 2_1, data packet 3_1,..., data packet n_1], where n in data packet n_1 represents the order of the data packet in the original flow, and 1 represents the data packet in source flow 1. If it is shuffled, it will be mixed with data packets from other flows, such as [data packet 1_1, data packet 2_5, data packet 3_5,... data packet n_5]. At this time, the composition of the data packets in this flow is mixed with those from flow 1 and flow 5 and is shuffled. Otherwise, if the data packets in the flow come from the same source, it is not shuffled.
[0124] S5. Based on the pre-trained model, fine-tune it with a small amount of labeled malicious traffic data set. During the training process, use the cross-entropy loss function to calculate the loss, and update the gradient according to the loss to obtain the updated model.
[0125] In a feasible implementation, in S5, based on the pre-trained model, fine-tune it with a small amount of labeled malicious traffic data set. During the training process, use the cross-entropy loss function to calculate the loss, and update the gradient according to the loss to obtain the updated model, including:
[0126] When the feature quality obtained by pre-training is high and there is an obvious distribution shift, select linear probing to fine-tune the model; when the data set is sufficient and the graphics card resources are sufficient, fine-tune the model with all parameters;
[0127] where fine-tuning includes using the cross-entropy loss function to fine-tune the model to optimize the classification performance of the downstream task:
[0128]
[0129] where y represents the true label, represents the predicted distribution. As Figure 5 shown is the framework in the fine-tuning stage.
[0130] S6. Through the updated model, perform malicious traffic detection and identification to complete the encrypted traffic threat detection based on multi-flow enhanced single-flow representation.
[0131] The present invention proposes to use similar flows in multi-flows to enhance the representation of a single flow. In addition, the present invention enables the model to be aware of the consistency of traffic, and the data packets in the flow have context relevance, allowing the model to distinguish whether the data packets in the flow come from the same flow. The present invention splits a small amount of labeled traffic for specific tasks into flows as inputs and uses a fine-tuning strategy to adjust the pre-trained model to make full use of the capabilities of the pre-trained model.
[0132] Figure 7 is a block diagram of an encrypted traffic threat detection device 300 shown according to an exemplary embodiment. The device 300 is used for the encrypted traffic threat detection method based on multi-flow enhanced single-flow representation. Referring to Figure 7 this, the device includes a data acquisition module 310, a data preprocessing module 320, a sample organization module 330, a pre-training module 340, a fine-tuning module 350, and a threat detection module 360. Among them:
[0133] The data acquisition module 310 is configured to collect traffic data samples and construct a traffic data set from the traffic data samples;
[0134] The data preprocessing module 320 is configured to perform data preprocessing on the traffic data set to obtain training samples;
[0135] The traffic embedding module 330 is configured to divide the training samples into samples in the pre-training stage and samples in the fine-tuning stage; for the samples in the pre-training stage, organize multi-flows according to quadruples to enrich the selection of positive samples in contrastive learning; for the samples in the fine-tuning stage, organize single-flows according to quintuples, and for each single-flow, extract the byte information of the first n data packets in the flow and the length sequence information of the first m data packets;
[0136] The pre-training module 340 is configured to perform pre-training in an unlabeled data set by setting proxy tasks to obtain a pre-trained model; wherein, the proxy tasks include a contrastive learning task and a flow consistency discrimination task;
[0137] The fine-tuning module 350 is configured to perform fine-tuning based on the pre-trained model through a small amount of labeled malicious traffic data set, calculate the loss using the cross-entropy loss function during training, and update the gradient according to the loss to obtain an updated model;
[0138] The threat detection module 360 is used to detect and identify malicious traffic through the updated model, and complete the encrypted traffic threat detection based on multi-stream enhanced single-stream characterization.
[0139] This paper proposes an Adaptive Interaction Feature Fusion (AIFF) module, which dynamically adjusts feature weights through an attention allocation mechanism, achieves efficient balance and fusion of short-term and long-term temporal dependencies, and thus significantly improves the overall performance of temporal action localization tasks.
[0140] Optionally, multiple pairs of traffic data sets are subjected to data preprocessing, including:
[0141] Remove invalid traffic;
[0142] Remove the fields irrelevant to traffic detection and set them to 0; the irrelevant fields include: checksum, IP address, and port number.
[0143] Optionally, a four-tuple includes: source IP address, destination IP address, destination port number, and protocol number;
[0144] The quintuple includes: source IP address, source port number, destination IP address, destination port number, and protocol number.
[0145] Optionally, contrastive learning tasks, including:
[0146] The unlabeled data set is split into traces or multiple streams according to the four-tuple {source IP, destination IP, destination port, protocol};
[0147] Construct a positive and negative sample sampling method to sample the positive and negative sample pairs required for contrastive learning;
[0148] By contrasting the learning loss function, similar flows in multiple flows are brought closer in the feature space, while flows that are not in the same trace are pulled apart in the vector space.
[0149] Among them, the positive samples in contrastive learning are selected from a single stream among multiple streams, and data enhancement is performed on the single stream, while the negative samples are samples from different batches maintained through a sliding queue;
[0150] The InfoNCE loss function is used to calculate loss optimization, shortening the distance between positive samples and increasing the distance between negative samples in the vector space.
[0151] Optionally, construct a positive and negative sample sampling method to sample the positive and negative sample pairs required for contrastive learning, including:
[0152] The pre-training data consists of triples Composition, where x and is a flow sample randomly selected from the same trace, is a negative sample maintained in the dictionary queue;
[0153] Among them, for a batch of samples, similar flows in multiple flows are used as positive samples through data augmentation, and negative samples are obtained by maintaining the dictionary queue.
[0154] Optionally, the flow consistency discrimination task includes:
[0155] In the given dataset, for single-flow data, the context is constructed by randomly pairing the flows;
[0156] Shuffle the data according to the preset probability p;
[0157] When the context information in the combined flows does not match, it is marked as "1", indicating that the data has been confused; on the contrary, when the context information remains consistent, it is marked as "0", indicating that the data is a complete flow;
[0158] Use the cross-entropy loss function for optimization:
[0159]
[0160] Among them, H represents the cross-entropy loss function; represents the value predicted by the model;
[0161] The overall loss function is composed of the sum of two parts of losses, jointly optimizing the learning process of the model as shown in the following formula:
[0162]
[0163] Among them, represents the loss function of the contrastive learning task.
[0164] Optionally, shuffling the data according to the preset probability p includes:
[0165] The preset probability p = 0.5; that is, half of the data is shuffled and half of the data remains unchanged according to the probability;
[0166] For a given piece of data, a probability p_ is randomly generated. If p_ > p, it is shuffled, otherwise the data remains unchanged.
[0167] Optionally, the fine-tuning module 340 is used to fine-tune the model by linear probing when the feature quality obtained by pre-training is high and there is an obvious distribution shift; when the dataset is sufficient and the graphics card resources are sufficient, the model is fine-tuned with all parameters;
[0168] Among them, fine-tuning includes fine-tuning the model using the cross-entropy loss function to optimize the classification performance of downstream tasks:
[0169]
[0170] where y represents the true label, represents the predicted distribution.
[0171] The present invention uses similar flows in multiple flows to enhance the representation of a single flow. In addition, the present invention enables the model to be aware of the consistency of traffic, and the data packets in the flow have context relevance, enabling the model to distinguish whether the data packets in the flow come from the same flow. The present invention splits a small amount of labeled traffic for specific tasks into flows as input and uses a fine-tuning strategy to adjust the pre-trained model to make full use of the capabilities of the pre-trained model.
[0172] Figure 8 is a schematic structural diagram of an encrypted traffic threat detection device based on enhancing the representation of a single flow by multiple flows, as Figure 8 shown, the encrypted traffic threat detection device based on enhancing the representation of a single flow by multiple flows may include the above-mentioned Figure 7 shown encrypted traffic threat detection device based on enhancing the representation of a single flow by multiple flows. Optionally, the encrypted traffic threat detection device 410 based on enhancing the representation of a single flow by multiple flows may include a first processor 2001.
[0173] Optionally, the encrypted traffic threat detection device 410 based on enhancing the representation of a single flow by multiple flows may further include a memory 2002 and a transceiver 2003.
[0174] Among them, the first processor 2001 is connected to the memory 2002 and the transceiver 2003, such as through a communication bus.
[0175] Next, in combination with Figure 8 each component of the encrypted traffic threat detection device 410 based on enhancing the representation of a single flow by multiple flows will be specifically introduced:
[0176] Among them, the first processor 2001 is the control center of the encrypted traffic threat detection device 410 based on multi-stream enhanced single-stream representation, which can be a single processor or a collective term for multiple processing elements. For example, the first processor 2001 is one or more central processing units (CPUs), or can be an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention. For example: one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs).
[0177] Optionally, the first processor 2001 can execute various functions of the encrypted traffic threat detection device 410 based on multi-stream enhanced single-stream representation by running or executing software programs stored in the memory 2002 and calling data stored in the memory 2002.
[0178] In a specific implementation, as an embodiment, the first processor 2001 may include one or more CPUs, such as Figure 8 CPU0 and CPU1 shown in
[0179] In a specific implementation, as an embodiment, the encrypted traffic threat detection device 410 based on multi-stream enhanced single-stream representation may also include multiple processors, such as Figure 8 the first processor 2001 and the second processor 2004 shown in
[0180] Among them, the memory 2002 is used to store software programs for executing the solution of the present invention and is controlled by the first processor 2001 for execution. The specific implementation manner can refer to the above method embodiments and will not be elaborated here.
[0181] Optionally, the memory 2002 may be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, or may also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 2002 may be integrated with the first processor 2001 or may exist independently and be coupled to the first processor 2001 through an interface circuit ( Figure 8 not shown) of the encrypted traffic threat detection device 410 based on multi-flow enhanced single-flow representation. The embodiments of the present invention do not make specific limitations on this.
[0182] The transceiver 2003 is used to communicate with a network device or with a terminal device.
[0183] Optionally, the transceiver 2003 may include a receiver and a transmitter ( Figure 8 not shown separately). Among them, the receiver is used to implement the receiving function, and the transmitter is used to implement the sending function.
[0184] Optionally, the transceiver 2003 may be integrated with the first processor 2001 or may exist independently and be coupled to the first processor 2001 through an interface circuit ( Figure 8 not shown) of the encrypted traffic threat detection device 410 based on multi-flow enhanced single-flow representation. The embodiments of the present invention do not make specific limitations on this.
[0185] It should be noted that Figure 8 the structure of the encrypted traffic threat detection device 410 based on multi-flow enhanced single-flow representation shown does not constitute a limitation on the router. The actual knowledge structure recognition device may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0186] In addition, the technical effects of the encrypted traffic threat detection device 410 based on multi-flow enhanced single-flow representation may refer to the technical effects of the encrypted traffic threat detection method based on multi-flow enhanced single-flow representation described in the above method embodiments, and will not be elaborated here.
[0187] It should be understood that the first processor 2001 in the embodiments of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0188] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0189] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable sensors. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.
[0190] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. Additionally, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood by referring to the context.
[0191] It should be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above processes do not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0192] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0193] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0194] In addition, in each embodiment of the present invention, each functional unit may be integrated in a processing unit, may exist physically alone for each unit, or two or more units may be integrated in one unit.
[0195] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention.
[0196] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A method for detecting encrypted traffic threats based on multi-stream enhanced single-stream characterization, characterized in that: The method comprises: S1. Collect flow data samples and construct the flow data samples into a flow data set; S2. performing data preprocessing on the traffic data set to obtain training samples; S3, dividing the training samples into samples of the pre-training stage and samples of the fine-tuning stage; for the samples of the pre-training stage, organizing multiple streams according to four-tuples to enrich the selection of positive samples in comparative learning; for the samples of the fine-tuning stage, organizing single streams according to five-tuples, for each single stream, extracting the byte information of the first n data packets in the stream and the length sequence information of the first m data packets; S4. Pre-training is performed by setting a proxy task in an unlabeled dataset to obtain a pre-trained model; wherein the proxy task includes a contrastive learning task and a flow consistency discrimination task; In S4, the contrastive learning task includes: The unlabeled data set is split into traces or multiple streams according to the four-tuple {source IP, destination IP, destination port, protocol}; Construct a positive and negative sample sampling method to sample the positive and negative sample pairs required for contrastive learning; Through contrastive learning loss function, similar flows in multiple flows are brought closer in feature space, while flows that are not in the same trace are pulled apart in vector space; Among them, the positive samples in contrastive learning are selected from a single stream among multiple streams, and data enhancement is performed on the single stream, while the negative samples are samples from different batches maintained through a sliding queue; The InfoNCE loss function is used to calculate loss optimization, which shortens the distance between positive samples and increases the distance between negative samples in the vector space. Flow consistency judgment tasks include: In a given dataset, for a single stream of data, the context is constructed by randomly combining the streams in pairs; Shuffle the data according to the preset probability p; When the context information in the combined stream does not match, it is marked as "1", indicating that the data has been obfuscated; conversely, when the context information remains consistent, it is marked as "0", indicating that the data is a complete stream; Use the cross-loss entropy function for optimization: ; Where H represents the cross entropy loss function; represents the value predicted by the model; y represents the true label; The overall loss function is the sum of the two losses, and the joint optimization model The learning process is as follows: ; in, represents the loss function of the contrastive learning task; x and is a flow sample randomly selected from the same trace. are negative samples maintained from the dictionary queue; S5. Based on the pre-trained model, fine-tune it through a small amount of labeled malicious traffic data set, calculate the loss using the cross-loss entropy function during the training process, perform gradient update according to the loss, and obtain an updated model; S6. Through the updated model, malicious traffic detection and identification are performed to complete encrypted traffic threat detection based on multi-stream enhanced single-stream characterization.
2. The encrypted traffic threat detection method based on multi-stream enhanced single-stream characterization according to claim 1 is characterized in that: In S2, the traffic data set is preprocessed, including: Remove invalid traffic; Remove the fields irrelevant to traffic detection and set them to 0; the irrelevant fields include: checksum, IP address, and port number.
3. The encrypted traffic threat detection method based on multi-stream enhanced single-stream characterization according to claim 2 is characterized in that: The four-tuple includes: source IP address, destination IP address, destination port number, and protocol number; The quintuple includes: source IP address, source port number, destination IP address, destination port number, and protocol number.
4. The encrypted traffic threat detection method based on multi-stream enhanced single-stream characterization according to claim 3 is characterized in that: The method for constructing positive and negative sample sampling to sample the positive and negative sample pairs required in contrastive learning includes: The pre-training data consists of triples Composition, where x and is a flow sample randomly selected from the same trace. are negative samples maintained from the dictionary queue; Among them, for a batch of samples, similar streams in multiple streams are taken as positive samples through data enhancement, and negative samples are obtained through dictionary queue maintenance.
5. The encrypted traffic threat detection method based on multi-stream enhanced single-stream characterization according to claim 4 is characterized in that: The shuffling of the data according to the preset probability p includes: The default probability p=0.5; that is, half of the data is shuffled and half of the data remains the same; For a given data, randomly generate a probability p_, if p_>p then shuffle it, otherwise keep the data unchanged.
6. The encrypted traffic threat detection method based on multi-stream enhanced single-stream characterization according to claim 5 is characterized in that: In S5, based on the pre-trained model, fine-tuning is performed using a small amount of labeled malicious traffic data sets, and during the training process, a cross-loss entropy function is used to calculate the loss, and a gradient update is performed according to the loss to obtain an updated model, including: When the quality of features obtained by pre-training is high and there is an obvious distribution shift, linear detection is selected to fine-tune the model. When the data set is sufficient and the graphics card resources are sufficient, all parameters are used to fine-tune the model. The fine-tuning includes fine-tuning the model using the cross entropy loss function to optimize the classification performance of downstream tasks: ; Where y represents the true label, represents the prediction distribution.
7. An encrypted traffic threat detection device based on multi-stream enhanced single-stream characterization, the encrypted traffic threat detection device based on multi-stream enhanced single-stream characterization is used to implement the encrypted traffic threat detection method based on multi-stream enhanced single-stream characterization as claimed in any one of claims 1-6, characterized in that: The device comprises: A data collection module, used for collecting flow data samples and constructing the flow data samples into a flow data set; A data preprocessing module, used to perform data preprocessing on the traffic data set to obtain training samples; A traffic embedding module is used to divide the training samples into samples of the pre-training stage and samples of the fine-tuning stage; for the samples of the pre-training stage, multiple streams are organized according to four-tuples to enrich the selection of positive samples in comparative learning; for the samples of the fine-tuning stage, single streams are organized according to five-tuples, and for each single stream, the byte information of the first n data packets in the stream and the length sequence information of the first m data packets are extracted; A pre-training module, used to perform pre-training in an unlabeled dataset by setting proxy tasks to obtain a pre-training model; wherein the proxy tasks include contrastive learning tasks and flow consistency discrimination tasks; A fine-tuning module, which is used to perform fine-tuning based on the pre-trained model through a small amount of labeled malicious traffic data sets, calculate the loss using a cross-loss entropy function during the training process, perform gradient update according to the loss, and obtain an updated model; The threat detection module is used to detect and identify malicious traffic based on the updated model, and complete encrypted traffic threat detection based on multi-stream enhanced single-stream characterization.
8. An encrypted traffic threat detection device based on multi-stream enhanced single-stream characterization, characterized in that: The encrypted traffic threat detection device based on multi-stream enhanced single-stream characterization includes: A processor; a memory, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the encrypted traffic threat detection method based on multi-stream enhanced single-stream characterization as described in any one of claims 1-6 is implemented.
Citation Information
Patent Citations
Vulnerability detection method and device for web application and computer readable storage medium
CN114021051A
Multi-modal reasoning method and device based on large language model and knowledge graph
CN118193684A