A method and system for detecting malicious traffic of unknown category

By combining contrast learning and few-sample learning methods, extracting and analyzing the deep characteristics of network traffic, the existing intrusion detection systems are solved, and the problem of complex attacks and data set imbalance is achieved, and the rapid and accurate detection of malicious traffic of unknown categories is achieved.

CN119232502BActive Publication Date: 2025-05-06NANJING UNIV OF INFORMATION SCI & TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411755303.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-03
Publication Date
2025-05-06
Estimated Expiration
2044-12-03

AI Technical Summary

Technical Problem

Existing intrusion detection systems are difficult to extract deep features, cannot effectively deal with complex attack methods, and face dataset class imbalance and zero-day attack challenges.

Method used

The combination of contrast learning and few-sample learning is adopted to pre-process traffic data, feature extraction and data enhancement, features are extracted using contrast learning encoder, and new attack modes are quickly learned through the few-sample learning model to perform malicious traffic detection.

Benefits of technology

It realizes fast and accurate detection of malicious traffic in unknown categories, breaks through the limitations of traditional rules and signature libraries, significantly improves the ability to detect unknown attacks, reduces false positives and missed reports, and maintains efficient detection performance in different network environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119232502B_ABST
    Figure CN119232502B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for detecting malicious traffic of unknown category, which relates to the field of network security technology, and comprises the following steps: acquiring traffic data, preprocessing the traffic data, extracting features from the preprocessed traffic data, obtaining a feature sequence, inputting the feature sequence into a pre-established contrastive learning encoder for training, and obtaining a trained contrastive learning encoder; receiving sample traffic data, inputting the sample traffic data into a trained contrastive learning encoder, outputting an encoded feature vector, inputting the encoded feature vector into a pre-established few-sample learning model, obtaining a sample category prototype, acquiring a feature vector of a traffic sample to be detected, and performing similarity calculation between the feature vector of the traffic sample to be detected and the sample category prototype to determine whether the traffic data is malicious traffic; and identifying, preventing and responding to potential security threats or unauthorized access by monitoring and analyzing network traffic and system activities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of network security, and in particular to a method and system for detecting malicious traffic of unknown category. Background Art

[0002] With the continuous growth of the number of electronic devices and the increasing complexity of the network environment, network security issues are becoming more and more serious, causing huge losses to the network economy. In order to effectively avoid network attacks, building an intrusion detection system (IDS) has become the main means. Existing intrusion detection systems are divided into two categories: host-based and network-based. The former mainly detects by monitoring log information, while the latter analyzes network traffic to determine whether there is intrusion behavior. Although intrusion detection systems based on machine learning (ML) are widely used, it is difficult to extract deep features and cannot cope with increasingly complex attack methods. Deep learning (DL) has gradually become an important part of intrusion detection systems due to its powerful feature extraction capabilities, but it still faces problems such as imbalanced data sets. As malicious attackers' strategies continue to upgrade, the boundary between malicious traffic and normal traffic becomes increasingly blurred, and new types of malicious traffic continue to emerge. Existing detection methods face the challenge of zero-day attacks. Summary of the invention

[0003] In order to solve the deficiencies mentioned in the above background technology, the purpose of the present invention is to provide a method and system for detecting malicious traffic of unknown category.

[0004] In a first aspect, the purpose of the present invention can be achieved by the following technical solution: a method for detecting malicious traffic of unknown category, the method comprising the following steps:

[0005] Acquire traffic data, preprocess the traffic data, extract features from the preprocessed traffic data to obtain a feature sequence, input the feature sequence into a pre-established contrastive learning encoder for training, and obtain a trained contrastive learning encoder;

[0006] Receive sample traffic data, input the sample traffic data into a trained contrastive learning encoder, output the encoded feature vector, input the encoded feature vector into a pre-established few-sample learning model, derive the sample category prototype, obtain the feature vector of the traffic sample to be detected, calculate the similarity between the feature vector of the traffic sample to be detected and the sample category prototype, and determine whether the traffic data is malicious traffic based on the similarity calculation result.

[0007] In combination with the first aspect, in certain implementations of the first aspect, the method further includes: the preprocessing of the traffic data includes segmenting, cleaning, and normalizing the traffic data.

[0008] In combination with the first aspect, in some implementations of the first aspect, the method further includes: the process of extracting features from the preprocessed traffic data includes:

[0009] Deep neural networks are used to automatically extract high-dimensional features, and traffic is modeled through multi-layer convolutional networks or recurrent networks. The resulting feature vectors include traffic behavior patterns in the time dimension and microscopic features at the packet level.

[0010] In combination with the first aspect, in some implementations of the first aspect, the method further includes: performing data enhancement when inputting the feature vector into a pre-established contrastive learning encoder, and the calculation process of the data enhancement is as follows:

[0011] Given a sequence of network stream packets , which contains Network flow packets , for network flow packets Elements are masked. Assigned to an array , array Contains 1 to indivual Mask, Indicates that the indivual Mask contains masks with 0 at random positions, and the remaining The element at position 1 is kept unchanged, and the masking formula is defined as:

[0012] (1)

[0013] Add Mask The operation is expressed as:

[0014] (2).

[0015] In combination with the first aspect, in some implementations of the first aspect, the method further includes: after the data enhancement, assigning the category label of the original sample to the enhanced sample, so that the enhanced sample and the sample of the same category form a positive sample pair, and the enhanced sample and the sample of a different category form a negative sample pair;

[0016] The pre-established contrastive learning encoder is back-propagated and optimized using the contrastive loss function, which is defined as follows:

[0017] (3)

[0018] in, is the sequence of anchor samples, For the The contrast loss of anchor samples is For the quantity Contrastive loss of anchor samples The total value of As the characteristics of the anchor sample, the characteristics of the samples of the same category Constitute a positive sample pair, and Represent the positive sample pair set of anchor point samples and all sample pair sets respectively, is the number of samples in the positive sample set, is the positive sample set The traversal of each sample in For all sample pairs The traversal of samples in the contrast loss function; the sum in the denominator is the sum of all samples in the set Features Features of samples of the same category as the anchor sample The calculation is performed on the numerator, which represents the characteristics of the anchor point sample. Characteristics of samples of the same category Similarity between; parameters As a temperature coefficient, it is used to adjust the model to pay more attention to difficult samples; is an exponential function.

[0019] In combination with the first aspect, in some implementations of the first aspect, the method further includes: the process of inputting the encoded feature vector into a pre-established few-sample learning model includes:

[0020] By clustering, select the samples closest to the cluster center to build the support set, and randomly select samples as the query set:

[0021] Perform K-means clustering on the training set data, and select several samples closest to the cluster center as representative samples to form the support set of the prototype network. In each training iteration episode, randomly select several samples from the remaining samples as the query set. For each category , calculate its corresponding prototype center ; The calculation formula of the prototype center is as follows:

[0022] (4)

[0023] in, To support the concentration category The number of samples, middle For category Support set The traversal of samples in For sample The corresponding category, For sample Extracted feature vectors.

[0024] For each sample in the query set , its eigenvector needs to be calculated With category Prototype Center The Euclidean distance between , the distance calculation formula is as follows:

[0025] (5)

[0026] The query sample Classification is based on its Euclidean distance The smallest prototype center category , To traverse the categories of all prototype centers, the classification formula is as follows:

[0027] (6)

[0028] If there is a sample whose distance from the prototype center of all categories is greater than the preset threshold, the sample will be judged to belong to the new category and the prototype center will be updated accordingly.

[0029] In combination with the first aspect, in some implementations of the first aspect, the method further includes: the formula for calculating the similarity between the feature vector of the traffic sample to be detected and the sample category prototype is as follows:

[0030] (7)

[0031] in For a given input sample , output category for The probability of Represents exponential function , used to convert distance into probability; For sample The extracted feature vectors, For Category The prototype center, For all categories of traversal, A prototype center for all categories, It is a sample The eigenvector and prototype center of The distance between For sample The distance between the feature vector of and the center of all category prototypes.

[0032] In combination with the first aspect, in some implementations of the first aspect, the method further includes: the query set Each sample Its true category , calculate the cross entropy loss of all samples and take the average to get the loss function :

[0033] (8)

[0034] is the number of all traffic samples contained in the query set, is the traversal of samples in the query set, For sample The real category, For sample The extracted feature vectors, For Category The prototype center, For all categories of traversal, A prototype center for all categories, It is a sample The feature vector and its true category Prototype Center The distance between For sample The distance between the feature vector and the center of all category prototypes.

[0035] In a second aspect, in order to achieve the above-mentioned purpose, the present invention discloses a system for detecting malicious traffic of unknown category, comprising:

[0036] A contrastive learning module is used to obtain traffic data, preprocess the traffic data, extract features from the preprocessed traffic data to obtain a feature sequence, input the feature sequence into a pre-established contrastive learning encoder for training, and obtain a trained contrastive learning encoder;

[0037] The malicious traffic detection module is used to receive sample traffic data, input the sample traffic data into a trained contrastive learning encoder, output the encoded feature vector, input the encoded feature vector into a pre-established few-sample learning model, obtain the sample category prototype, obtain the feature vector of the traffic sample to be detected, calculate the similarity between the feature vector of the traffic sample to be detected and the sample category prototype, and determine whether the traffic data is malicious traffic based on the similarity calculation result.

[0038] In another aspect of the present invention, in order to achieve the above-mentioned purpose, a terminal device is disclosed, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor, wherein the memory stores a computer program capable of running on the processor, and when the processor loads and executes the computer program, a method for detecting malicious traffic of an unknown category as described above is adopted.

[0039] Beneficial effects of the present invention:

[0040] The present invention combines contrastive learning with few-sample learning, so that the system can quickly adapt and perform accurate detection when facing zero-day attacks and new threats. This method breaks through the limitations of traditional reliance on rules and signature libraries, and significantly improves the ability to detect unknown attacks; the features extracted by contrastive learning have stronger discrimination, and combined with few-sample learning to quickly learn new attack patterns, it significantly reduces the situation where normal traffic is misjudged as an attack (false positive) and attack traffic is not detected (missed negative). Using a few-sample learning model, it can still quickly learn attack features and apply them to actual detection tasks when there is only a small amount of labeled data. This solves the problem of high data labeling costs and scarce labeled data; through the self-supervision mechanism of contrastive learning, the system can learn more generalized features from massive network traffic, so that it can maintain efficient detection performance in different network environments, different attack types and complex and changing traffic patterns. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0042] Figure 1 It is a schematic flow chart of the method of the present invention;

[0043] Figure 2 It is a schematic diagram of the overall framework of the method of the present invention;

[0044] Figure 3 It is a schematic diagram of the comparative learning pre-training structure of the present invention;

[0045] Figure 4 Schematic diagram of the data enhancement method of the present invention;

[0046] Figure 5 It is a schematic diagram of constructing positive and negative sample pairs in the present invention;

[0047] Figure 6 It is a schematic diagram of the classification structure of few-sample learning of the present invention;

[0048] Figure 7 It is a schematic diagram of the system structure of the present invention. DETAILED DESCRIPTION

[0049] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0050] Embodiment 1:

[0051] The following is an introduction to the relevant terms involved in the embodiments of the present application:

[0052] Intrusion detection: Intrusion detection is a reasonable supplement to firewalls. It helps the system deal with network attacks, expands the security management capabilities of system administrators (including security auditing, monitoring, attack identification and response), and improves the integrity of the information security infrastructure. It collects information from several key points in the computer network system and analyzes the information to see if there are any violations of security policies and signs of attacks in the network. Intrusion detection is considered to be the second security gate after the firewall. It can detect the network without affecting network performance, thereby providing real-time protection against internal attacks, external attacks and misoperation.

[0053] like Figure 1 As shown, a method for detecting malicious traffic of unknown category includes the following steps:

[0054] S101: acquiring flow data, preprocessing the flow data, extracting features from the preprocessed flow data to obtain a feature sequence, inputting the feature sequence into a pre-established contrastive learning encoder for training, and obtaining a trained contrastive learning encoder;

[0055] The method aims to solve the problems of imbalanced malicious traffic categories and blurred traffic category boundaries, and to achieve effective detection of unknown malicious traffic. The method consists of two parts: contrastive learning pre-training and few-sample learning strategy fine-tuning. The overall method framework is as follows Figure 2 As shown in the figure, the traffic data is first preprocessed, including data cleaning and normalization, to ensure the quality and consistency of the data. Then, positive and negative samples are obtained through data enhancement. After feature extraction, the feature extractor is input into the contrastive learning task to train the feature extractor. The trained feature extractor is used for feature extraction in the prototype network classifier.

[0056] Preprocessing of traffic data includes segmenting, cleaning and normalizing traffic data. Specifically, traffic data is first parsed, including extracting five-tuple information (source IP, destination IP, source port, destination port, protocol) and packet timing information from network data packets. At the same time, traffic data is stream segmented to ensure that it can be applied in a real-time environment. In order to reduce the impact of noise data, feature normalization and irrelevant information filtering are also performed.

[0057] The process of extracting features from preprocessed traffic data includes:

[0058] A deep neural network is used to automatically extract high-dimensional features, and traffic is modeled through a multi-layer convolutional network or a recurrent network. The resulting feature vector includes traffic behavior patterns in the time dimension and microscopic features at the packet level. In addition, in order to ensure the generalization ability of the model, an attention mechanism may be added to enable the model to automatically focus on the key parts of the traffic and improve the accuracy of feature extraction.

[0059] By introducing contrastive learning to enhance the distinction between different types of traffic, the model can better learn the differences between different types of traffic, thereby improving the accuracy and robustness of the intrusion detection model. The contrastive learning pre-training process consists of data enhancement, encoder and contrastive learning tasks. Its structure is as follows: Figure 3 As shown in the figure, the traffic data is enhanced to obtain positive and negative samples, which are input into the encoder together with the anchor samples for feature extraction, and the contrast learning task is performed to obtain the contrast loss which is back-propagated to the encoder for training. The purpose is to amplify the distance between the features of samples of different categories and reduce the distance between the features of the same samples.

[0060] When constructing contrastive learning tasks, samples are usually classified by category labels. Samples of the same category form positive sample pairs, while samples of different categories form negative sample pairs. To address the problem of less sample data in some categories of malicious traffic, a data augmentation method of adding random masks is used to generate extended samples with the same semantics as the original traffic samples. These extended samples form positive sample pairs with samples of the same category as the original traffic, and form negative sample pairs with samples of other categories. This method of expanding the data set provides a more reliable foundation for subsequent contrastive learning tasks. Data augmentation methods such as Figure 4 As shown, enhanced samples are constructed by adding random masks to the original traffic data.

[0061] Data augmentation is performed when the feature vector is input into a pre-established contrastive learning encoder. The data augmentation calculation process is as follows:

[0062] Given a sequence of network stream packets , which contains Network flow packets , for network flow packets Elements are masked. Assigned to an array , array Contains 1 to indivual Mask, Indicates that the indivual Mask contains masks with 0 at random positions, and the remaining The element at position 1 is kept unchanged, and the masking formula is defined as:

[0063] (1)

[0064] Add Mask The operation is expressed as:

[0065] (2).

[0066] The core of contrastive learning is to construct positive and negative sample pairs. For network traffic, positive sample pairs can be traffic from the same category (for example, different instances of malicious traffic), while negative sample pairs are traffic from different categories (for example, the comparison between malicious and normal traffic). In contrastive learning, the loss function usually uses contrastive loss. In this way, the features learned by the model have a higher similarity in samples of the same category, while the similarity between samples of different categories is lower.

[0067] The following is the theoretical support for constructing contrastive learning tasks by adding random mask data augmentation methods:

[0068] (1) Although the data has been processed by random masking, some features and information of the original data are still retained, which makes the analysis of the masked data still valid. Therefore, the random masking method can guarantee the validity of traffic data to a certain extent.

[0069] (2) In a complex and changeable network environment, network fluctuations may cause a small amount of packet loss in network traffic data. By adding a random mask to generate traffic data, we can effectively simulate the actual network environment and reflect the actual packet loss phenomenon.

[0070] (3) In network intrusion detection tasks, the detection model should have a certain degree of robustness. Even if there are occasional missing packets in the traffic sequence, a robust intrusion detection model should still be able to accurately identify and detect potential intrusion behaviors.

[0071] After the data augmentation operation, the category label of the original sample is assigned to the augmented sample. In this way, the augmented sample forms a positive sample pair with samples of the same category, and a negative sample pair with samples of different categories, such as Figure 5 As shown in Figure 2, anchor samples and enhanced samples form positive sample pairs, and enhanced samples and samples of other categories form negative sample pairs.

[0072] To improve the classification ability of the model, the encoder maps the samples to the feature space in order to perform contrastive learning tasks. The core idea of ​​contrastive learning is to use the contrastive loss function to back-propagate the encoder so that the model can effectively distinguish samples of different categories. Specifically, the goal of contrastive loss is to improve the discriminative performance of the model by shortening the distance between positive sample pairs in the feature space and increasing the distance between negative sample pairs. The definition of the contrastive loss function is as follows:

[0073] (3)

[0074] in, is the sequence of anchor samples, For the The contrast loss of anchor samples is For the quantity Contrastive loss of anchor samples The total value of As the characteristics of the anchor sample, the characteristics of the samples of the same category Constitute a positive sample pair, and a negative sample pair with samples of different categories. and Represent the positive sample pair set of anchor point samples and all sample pair sets respectively, is the number of samples in the positive sample set, is the positive sample set The traversal of each sample in For all sample pairs The sum in the denominator of the contrast loss function is the sum of all samples in the set. Features Features of samples of the same category as the anchor sample The calculation is performed on the numerator, which represents the characteristics of the anchor point sample. Characteristics of samples of the same category To optimize formula (3), we need to maximize the dot product in the numerator and minimize the value of the denominator. As a temperature coefficient, it aims to guide the model to pay more attention to negative samples that are difficult to distinguish. The goal of contrast loss is to improve the classification performance of the model by optimizing the feature representation so that the distance between positive sample pairs in the feature space is reduced and the distance between negative sample pairs is increased.

[0075] S102: Receive sample traffic data, input the sample traffic data into a trained contrastive learning encoder, output the encoded feature vector, input the encoded feature vector into a pre-established few-sample learning model, derive the sample category prototype, obtain the feature vector of the traffic sample to be detected, calculate the similarity between the feature vector of the traffic sample to be detected and the sample category prototype, and determine whether the traffic data is malicious traffic based on the similarity calculation result.

[0076] By proposing an improved classification method based on the few-shot prototype network, the sample closest to the cluster center is selected through clustering to build a support set, and samples are randomly selected as the query set. In the model fine-tuning stage, the encoder obtained from the contrastive learning pre-training is used to perform the few-shot learning task, and the prototype network is trained to achieve traffic classification and unknown malicious traffic detection. The few-shot learning classification structure is as follows Figure 6 As shown, it supports inputting concentrated samples into the encoder trained by contrastive learning for feature extraction to obtain feature vectors, constructing a prototype network in the feature space with the feature vectors of the samples in the training set, and then inputting the samples in the query set into the prototype network to obtain the cross entropy loss, and back-propagating to optimize the feature extraction encoder to obtain the prototype network classifier to classify and identify the traffic.

[0077] In order to improve the performance of few-shot learning methods in intrusion detection, a clustering-based strategy is introduced to select the support set. First, K-means clustering is performed on the training set data, and several samples closest to the cluster center are selected as representative samples to form the support set of the prototype network. In each training iteration, several samples are randomly selected from the remaining samples as the query set to calculate the loss and update the model parameters. For each category , calculate its corresponding prototype center The calculation formula of the prototype center is as follows:

[0078] (4)

[0079] in, To support the concentration category The number of samples, is the encoder for extracting features, For category Support set The traversal of samples in For sample The corresponding category, For sample Extracted feature vectors.

[0080] For each sample in the query set , its eigenvector needs to be calculated With category Prototype Center The Euclidean distance between them is calculated as follows:

[0081] (5)

[0082] Then, the query sample Classified as the category of the prototype center closest to it , To traverse the categories of all prototype centers, the classification formula is as follows:

[0083] (6)

[0084] If there is a sample whose distance from the prototype center of all categories is greater than the preset threshold, which is the maximum distance of the farthest sample from the center of each category among all known categories, then the sample will be judged to belong to a new category and the prototype center will be updated accordingly. This ensures that the model can adapt to newly emerging categories, thereby improving its flexibility and adaptability. Suppose a traffic dataset contains three known categories A, B, and C. Each category has a prototype center, and the distance from all samples in each category to its category center has been calculated. For categories A, B, and C, the farthest distance from the prototype center of the samples in the category is 3, 4, and 5 respectively. The thresholds of samples A, B, and C are set to 3, 4, and 5 respectively. Suppose there is a new sample whose distance from the prototype center of A, B, and C is 5, 5, and 6 respectively, which are all greater than the preset threshold of each category. This sample will be judged to belong to a new category, and the model will be updated to include this new category and its prototype center.

[0085] In the prototype network, the model calculates the probability that the query sample belongs to each category based on the Euclidean distance. Specifically, the smaller the distance between the sample and the center of the category prototype, the higher the probability that the sample belongs to the category. To this end, the softmax function can be used to convert the distance into probability. The specific calculation process is as follows:

[0086] (7)

[0087] in For a given input sample , output category for The probability of Represents exponential function , used to convert the distance to non-negative values ​​and scale them to help form the probability distribution; For sample The extracted feature vectors, For Category The prototype center, For all categories of traversal, For Category The prototype center, It is a sample The eigenvector and prototype center of The distance between For sample The distance between the feature vector of and the center of all category prototypes.

[0088] For query sets Each sample Its true category , calculate the cross entropy loss of all samples and take the average to get the loss function :

[0089] (8)

[0090] is the number of all traffic samples contained in the query set, For query set The traversal of samples in For sample The real category, For sample The extracted feature vectors, For Category The prototype center, For all categories of traversal, For Category The prototype center, It is a sample The feature vector and its true category Prototype Center The distance between For sample The distance between the feature vector of and the center of all category prototypes.

[0091] The loss function guides the update of model parameters during the training process, so that the model can perform few-sample classification tasks more accurately and effectively detect new attack traffic.

[0092] Embodiment 2: In the second aspect, as Figure 7 As shown, in order to achieve the above-mentioned purpose, the present invention discloses a system for detecting unknown malicious traffic, including:

[0093] The contrastive learning module 11 is used to obtain traffic data, pre-process the traffic data, extract features from the pre-processed traffic data to obtain a feature sequence, and input the feature sequence into a pre-established contrastive learning encoder for training to obtain a trained contrastive learning encoder;

[0094] The malicious traffic detection module 12 is used to receive sample traffic data, input the sample traffic data into a trained contrast learning encoder, output the encoded feature vector, input the encoded feature vector into a pre-established few-sample learning model, obtain the sample category prototype, obtain the feature vector of the traffic sample to be detected, calculate the similarity between the feature vector of the traffic sample to be detected and the sample category prototype, and determine whether the traffic data is malicious traffic based on the similarity calculation result.

[0095] Based on the same inventive concept, the present invention also provides a computer device, which includes: one or more processors, and a memory for storing one or more computer programs; the program includes program instructions, and the processor is used to execute the program instructions stored in the memory. The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, which is used to implement one or more instructions, specifically for loading and executing one or more instructions in a computer storage medium to implement the above method.

[0096] It should be further explained that, based on the same inventive concept, the present invention also provides a computer storage medium, on which a computer program is stored, and the computer program is executed by the processor to execute the above method. The storage medium can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electrical, magnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples (non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, a computer-readable storage medium can be any tangible medium containing or storing a program, which can be used by an instruction execution system, device or device or used in combination with it.

[0097] In the description of this specification, the description with reference to the terms "one embodiment", "example", "specific example", etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0098] The above shows and describes the basic principles, main features and advantages of the present disclosure. Those skilled in the art should understand that the present disclosure is not limited by the above embodiments, and the above embodiments and descriptions are only for explaining the principles of the present disclosure. Without departing from the spirit and scope of the present disclosure, the present disclosure may have various changes and improvements, and these changes and improvements fall within the scope of the present disclosure to be protected.

Claims

1. A method for detecting malicious traffic of unknown category, characterized in that: The method comprises the following steps: Acquire traffic data, preprocess the traffic data, extract features from the preprocessed traffic data to obtain a feature sequence, input the feature sequence into a pre-established contrastive learning encoder for training, and obtain a trained contrastive learning encoder; Receive sample traffic data, input the sample traffic data into a trained contrastive learning encoder, output the encoded feature vector, input the encoded feature vector into a pre-established few-sample learning model, obtain the sample category prototype, obtain the feature vector of the traffic sample to be detected, calculate the similarity between the feature vector of the traffic sample to be detected and the sample category prototype, and determine whether the traffic data is malicious traffic based on the similarity calculation result; The process of inputting the encoded feature vector into the pre-established few-shot learning model includes: By clustering, select the samples closest to the cluster center to build the support set, and randomly select samples as the query set: Perform K-means clustering on the training set data, and select several samples closest to the cluster center as representative samples to form the support set S of the prototype network. In each training iteration episode, randomly select several samples from the remaining samples as the query set. For each category c, calculate its corresponding prototype center p c ; The calculation formula of the prototype center is as follows: Among them, |S c | is the number of samples of category c in the support set, x i is the support set S of category c c The traversal of samples in y i For sample x i The corresponding category, f φ (x i ) is the sample x i Extracted feature vectors; For each sample x in the query set, its feature vector f needs to be calculated φ (x) and the prototype center p of category c c The Euclidean distance d(f φ (x),p c ), the distance calculation formula is as follows: d(f φ (x),p c )=||f φ (x)-p c || 2 (2) The query sample x is classified into the category y' where the prototype center with the smallest Euclidean distance is located, and c is the category of all prototype centers. The classification formula is as follows: If the distance between a sample and the prototype center of all categories is greater than the preset threshold, the sample will be judged as belonging to a new category and the prototype center will be updated accordingly; The formula for calculating the similarity between the feature vector of the traffic sample to be detected and the sample category prototype is as follows: Where p(y=c|x) is the probability that the output category y is the category c given the input sample x; exp() represents the exponential function; f φ (x) is the feature vector extracted from sample x, p c is the prototype center of category c, c 'For all categories of traversal, p c' is the prototype center of category c', d(f φ (x),p c ) is the characteristic vector of sample x and the prototype center p c The distance between φ (x),p c' ) is the distance between the feature vector of sample x and the center of all category prototypes; For each sample x and its true category y in the query set Q, the cross entropy loss of all samples is calculated and averaged to obtain the loss function L: |Q| is the number of all traffic samples contained in the query set, x is the sample data in the query set Q, y is the true category of sample x, and f φ (x) is the extracted feature vector, p y is the prototype center of category y, c' is the traversal of all categories, p c ' is the prototype center of category c', d(f φ (x),p y ) is the prototype center p of the feature vector of sample x and the true category y y The distance between φ (x),p c' ) is the distance between the feature vector of sample x and the center of all category prototypes.

2. According to claim 1, a method for detecting malicious traffic of unknown category, characterized in that: The preprocessing of the flow data includes segmenting, cleaning and normalizing the flow data.

3. According to claim 1, a method for detecting malicious traffic of unknown category, characterized in that: The process of extracting features from the preprocessed traffic data includes: Deep neural networks are used to automatically extract high-dimensional features, and traffic is modeled through multi-layer convolutional networks or recurrent networks. The resulting feature vector includes traffic behavior patterns in the time dimension and microscopic features at the packet level.

4. The method for detecting malicious traffic of unknown category according to claim 1, characterized in that: The data enhancement is performed when the feature vector is input into the pre-established contrastive learning encoder. The calculation process of the data enhancement is as follows: Given a network flow packet sequence flow = [p1, p2, ..., p m ], which contains m network flow packets p, and performs masking operation on k elements in the network flow packets. Mask is assigned to an array [mask1,mask2,…,mask m ], the array contains 1 to m masks, where the ith mask contains k random positions with 0, and the remaining (mk) positions are 1, keeping the original data unchanged. The mask adding formula is defined as: Mask=[mask1,mask2,...,mask m ] Add mask operation flow masked It is expressed as: flow masked =flow*Mask (7)。 5. The method for detecting malicious traffic of unknown category according to claim 4, characterized in that: After the data is enhanced, the category label of the original sample is assigned to the enhanced sample, so that the enhanced sample and the sample of the same category form a positive sample pair, and the enhanced sample and the sample of a different category form a negative sample pair; The contrastive loss function L is used to perform back propagation optimization on the pre-established contrastive learning encoder. contrastive is defined as follows: Among them, i is the serial number of the anchor sample, is the contrast loss of the i-th anchor point sample, L contrastive is the contrast loss of the number of anchor point samples i The total value of z i As the feature of the anchor sample, it is similar to the feature z of the sample of the same category. p Constitute a positive sample pair, P (i) and A (i) Represent the positive sample pair set of anchor samples and the total sample pair set, respectively, |p (i) | is the number of samples in the positive sample pair set, and sample P is the positive sample pair set P (i) The traversal of each sample in, sample a is the set of all sample pairs A (i) The traversal of samples in the contrast loss function; the sum in the denominator is the feature z of sample a in all sample pairs a Features z of samples of the same category as the anchor sample p The calculation is performed on the numerator, which represents the characteristics of the anchor sample and the characteristics of the samples of the same category z p The similarity between them; parameter τ∈R + is the temperature coefficient; exp() is the exponential function.

6. A system for detecting malicious traffic of unknown category, adopting a method for detecting malicious traffic of unknown category as claimed in any one of claims 1 to 5, characterized in that: include: A contrastive learning module is used to obtain traffic data, preprocess the traffic data, extract features from the preprocessed traffic data to obtain a feature sequence, input the feature sequence into a pre-established contrastive learning encoder for training, and obtain a trained contrastive learning encoder; The malicious traffic detection module is used to receive sample traffic data, input the sample traffic data into a trained contrastive learning encoder, output an encoded feature vector, input the encoded feature vector into a pre-established few-sample learning model, obtain a sample category prototype, obtain a feature vector of the traffic sample to be detected, perform similarity calculation between the feature vector of the traffic sample to be detected and the sample category prototype, and determine whether the traffic data is malicious traffic based on the similarity calculation result; The process of inputting the encoded feature vector into the pre-established few-shot learning model includes: By clustering, select the samples closest to the cluster center to build the support set, and randomly select samples as the query set: Perform K-means clustering on the training set data, and select several samples closest to the cluster center as representative samples to form the support set S of the prototype network. In each training iteration episode, randomly select several samples from the remaining samples as the query set. For each category c, calculate its corresponding prototype center p c ; The calculation formula of the prototype center is as follows: Among them, |S c | is the number of samples of category c in the support set, x i is the support set S of category c c The traversal of samples in y i For sample x i The corresponding category, f φ (x i ) is the sample x i Extracted feature vectors; For each sample x in the query set, its feature vector f needs to be calculated φ (x) and the prototype center p of category c c The Euclidean distance d(f φ (x),p c ), the distance calculation formula is as follows: d(f φ (x),p c )=||f φ (x)-p c || 2 (10) The query sample x is classified into the category y' where the prototype center with the smallest Euclidean distance is located, and c is the category of all prototype centers. The classification formula is as follows: If the distance between a sample and the prototype center of all categories is greater than the preset threshold, the sample will be judged as belonging to a new category and the prototype center will be updated accordingly; The formula for calculating the similarity between the feature vector of the traffic sample to be detected and the sample category prototype is as follows: Where p(y=c|x) is the probability that the output category y is the category c given the input sample x; exp() represents the exponential function; f φ (x) is the feature vector extracted from sample x, p c is the prototype center of category c, c ' For all categories of traversal, p c' is the prototype center of category c', d(f φ (x),p c ) is the characteristic vector of sample x and the prototype center p c The distance between φ (x),p c' ) is the distance between the feature vector of sample x and the center of all category prototypes; For each sample x and its true category y in the query set Q, the cross entropy loss of all samples is calculated and averaged to obtain the loss function L: |Q| is the number of all traffic samples contained in the query set, x is the sample data in the query set Q, y is the true category of sample x, and f φ (x) is the extracted feature vector, p y is the prototype center of category y, c' is the traversal of all categories, p c ' is the prototype center of category c', d(f φ (x),p y ) is the prototype center p of the feature vector of sample x and the true category y y The distance between φ (x),p c' ) is the distance between the feature vector of sample x and the center of all category prototypes.

7. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that: The memory stores a computer program that can be run on the processor. When the processor loads and executes the computer program, a method for detecting malicious traffic of unknown category described in any one of claims 1 to 5 is adopted.

Citation Information

Patent Citations

  • Malicious traffic detection method based on mask automatic encoder pre-training

    CN118400195A