An Encrypted Malicious Traffic Detection Method Based on Multi-Scale Spatiotemporal Interaction Graph Network

By building a multi-scale spatiotemporal interactive graph network, extracting multi-scale features, and using the Deep SVDD model for malicious traffic detection, the problem of difficult to identify encrypted malicious traffic in the existing technology is solved, and a higher malicious traffic recognition accuracy and detection efficiency are achieved.

CN119583154BActive Publication Date: 2025-06-20CHINA UNIV OF MINING & TECH (BEIJING)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411699305.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-26
Publication Date
2025-06-20
Estimated Expiration
2044-11-26

AI Technical Summary

Technical Problem

It is difficult for the prior art to effectively identify and detect encrypted malicious traffic. Traditional traffic detection methods cannot encapsulate traffic data packets through encryption protocols, resulting in inaccuracy and inefficiency in identifying malicious traffic.

Method used

The encrypted malicious traffic detection method based on a multi-scale spatiotemporal interaction graph network is adopted. By constructing in-stream traffic interaction graphs and inter-stream spatiotemporal maps, multi-scale features are extracted, and a single classification model of normal traffic and malicious traffic is trained using the Deep SVDD single classification objective function, and the classification results of the two models are combined for prediction.

Benefits of technology

It improves the accuracy of malicious traffic identification, can more effectively identify and distinguish normal traffic from malicious traffic, and enhances the detection ability of unknown types of network traffic.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119583154B_ABST
    Figure CN119583154B_ABST
Patent Text Reader

Abstract

The present invention discloses an encrypted malicious traffic detection method based on a multi-scale spatio-temporal interaction graph network, including: constructing an intra-flow traffic interaction graph network based on the interaction relationship between data packets between the client and the server in the flow; constructing an intra-flow traffic interaction graph feature extractor to obtain the local feature information of the flow; constructing an inter-flow spatio-temporal graph; constructing an inter-flow spatio-temporal graph feature extractor to obtain the spatio-temporal feature information of the flow; obtaining multi-scale fusion features according to a multi-scale feature fusion module; obtaining a normal traffic deep single-classification model; obtaining a malicious traffic deep single-classification model; acquiring the traffic data to be measured and inputting them into the normal traffic deep single-classification model and the malicious traffic deep single-classification model respectively to obtain classification results; and detecting the traffic category based on the classification results. The present invention improves the accuracy of malicious traffic identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of network security technology, and particularly relates to an encrypted malicious traffic detection method based on a multi-scale spatio-temporal interaction graph network. Background Art

[0002] In the field of network security, encrypted malicious traffic detection aims to identify and prevent malicious access behaviors in the network, thereby protecting computer networks and their resources from illegal attacks. However, at the present stage, network traffic usually uses encryption protocols such as SSL / TLS to encrypt the traffic, preventing users from directly accessing traffic data packets, resulting in the inability of traditional traffic detection methods to effectively identify encrypted malicious traffic.

[0003] Traditional traffic detection methods usually adopt machine learning methods, which mainly include: anomaly detection methods, classification detection methods, and hybrid detection methods. Anomaly detection methods construct a normal traffic detection model by using a large number of normal network access traffic characteristics. For traffic that deviates from the normal traffic clustering center and exceeds the model-set threshold, it is determined as abnormal traffic (or malicious traffic). Qin et al. constructed a traffic detection model based on the K-means algorithm by using network traffic characteristics such as destination address, destination port, source address, packet size, and traffic duration entropy, and were able to effectively detect DDoS attack traffic. Zolotukhin et al. used the DBSCAN (Density-Based Spatial Clustering of Applications with Noise) algorithm to construct a normal traffic detection model. If the target traffic exceeds the detection model threshold, it is determined as DDoS attack traffic. In summary, the anomaly detection method based on the clustering algorithm can not only update the detection model online, but also enhance the ability to detect unknown traffic. However, if an attacker mimics the normal user network access behavior, the above methods may have the problem of clustering attack-related traffic into the normal traffic type.

[0004] Since abnormal network traffic also has significant characteristic information, malicious traffic detection can be regarded as a traffic classification problem. A network traffic classifier is constructed based on normal traffic and abnormal traffic, and then the classifier is used to identify whether the target traffic is abnormal traffic. Ajaeiya et al. proposed to use the statistical data of network traffic as a feature description to construct a traffic classification model based on the Baggedtrees algorithm. Xu et al. used information entropy, packet size, packet frequency and connection time to describe the characteristics of attack traffic, constructed a traffic classification model based on the random forest algorithm, and used it to detect whether the network node is under malicious traffic network attack. However, the training of classification-based detection methods is highly dependent on traffic feature description, but the feature description of malicious traffic will change with the development of encryption technology, making it difficult for this method to accurately identify unknown types of network traffic.

[0005] The hybrid detection method takes advantage of the technical advantages of the above two detection methods and effectively improves the performance of network anomaly detection. The method is usually divided into two steps. The first step is classification-based detection, and the second step is anomaly-based detection. Traffic samples classified as "normal" in the first step will be further processed in the second step and re-evaluated using anomaly or outlier detection algorithms to identify potential abnormal behaviors. IoTArgos detects suspicious behaviors by extracting 11 key features, including the number of packets every 5 minutes, and combining supervised learning and unsupervised learning algorithms. Supervised machine learning algorithms are used to filter attack traffic subsets, while unsupervised machine learning algorithms are used to identify samples that were misclassified as "normal" in the previous stage. Compared with single anomaly detection or classification detection methods, the hybrid detection method not only performs well in the classification accuracy of known attacks, but can also effectively detect unknown attack behaviors.

[0006] Deep learning methods are widely used in encrypted malicious traffic detection due to their powerful feature extraction capabilities. They can directly extract features from raw traffic. Zeng et al. proposed an end-to-end anomaly detection framework based on deep learning. Compared with traditional machine learning methods, deep learning algorithms require less manual intervention. Multi-layer neural networks can combine shallow features to form complex and abstract features, which are suitable for detecting hidden attack behaviors. Zolotukhin et al. combined clustering algorithms with SAE (Stack Auto Encoder) algorithms to build a DDoS attack detection model. Eight features, such as the percentage of packets with different TCP flags, were extracted in each time interval as inputs to the clustering and SAE models. The clustering algorithm is used to build a normal user behavior model and mark the behavior that deviates from this model as a DDoS attack; SAE is used to detect attacks that imitate normal user browsing behaviors, which are usually not effectively identified by clustering algorithms.

[0007] Knowledge-based methods determine whether network traffic is malicious through predefined rules. These rules are usually based on ports, protocol headers, and manually extracted features. David and Thomas proposed a fingerprint method that uses fast entropy to calculate the traffic count for each connection. However, the reliance of this method on traffic counts limits its accuracy in distinguishing DDoS attacks. Summary of the Invention

[0008] To solve the above technical problems, the present invention proposes an encrypted malicious traffic detection method based on a multi-scale spatio-temporal interaction graph network. The traffic data is stored and predicted in the form of a graph. Traffic features are extracted through a multi-scale feature fusion module. Two single-class models for normal traffic and malicious traffic are respectively trained using the Deep SVDD single-class objective function, and the two models are used together to predict the traffic type, improving the accuracy of malicious traffic identification.

[0009] To achieve the above object, the present invention provides an encrypted malicious traffic detection method based on a multi-scale spatio-temporal interaction graph network, including:

[0010] Obtain traffic source data, and process the traffic source data to obtain flows;

[0011] Based on the interaction relationship between packets between the client and the server in the flow, construct an in-flow traffic interaction graph network; construct an in-flow traffic interaction graph feature extractor based on the in-flow traffic interaction graph network to obtain the local feature information of the flow;

[0012] Based on the temporal and spatial relationships between flows, construct an inter-flow spatio-temporal graph; construct an inter-flow spatio-temporal graph feature extractor based on the inter-flow spatio-temporal graph to obtain the spatio-temporal feature information of the flow;

[0013] Extract different-scale features based on the local feature information of the flow and the spatio-temporal feature information of the flow, and obtain multi-scale fusion features according to the multi-scale feature fusion module;

[0014] Use normal traffic as positive samples and malicious traffic as negative samples, and train based on the Deep SVDD single-class objective function by combining the multi-scale fusion features with the normal traffic dataset and the malicious traffic dataset to obtain a normal traffic deep single-class model;

[0015] Use normal traffic as negative samples and malicious traffic as positive samples, and train based on the Deep SVDD single-class objective function by combining the multi-scale fusion features with the normal traffic dataset and the malicious traffic dataset to obtain a malicious traffic deep single-class model;

[0016] Obtain the traffic data to be measured, and input it into the normal traffic depth single-classification model and the malicious traffic depth single-classification model respectively to obtain classification results; detect the traffic category based on the classification results.

[0017] Optionally, constructing an intra-flow traffic interaction graph network includes:

[0018] Construct an intra-flow interaction graph for each flow, where the data packets in the flow are used as nodes, and edge connections are established between consecutive data packets within the same consecutive data packet sequence transmitted in the same direction, and the first and last data packets in the two consecutive data packet sequences transmitted in the same direction are connected respectively.

[0019] Optionally, constructing an inter-flow spatio-temporal graph includes:

[0020] Each flow serves as a node, and edge connections are established between the current flow and the top k flows with the closest start time to it, and between other flows with the same source IP address or destination IP address; five different types of edges are constructed according to the IP address, namely, both the source IP address and the destination IP address are the same, only the source IP address is the same, only the destination IP address is the same, the source IP address is the same as the destination IP address of another flow, and the destination IP address is the same as the source IP address of another flow; among them, the first three types are bidirectional edges, and the last two types are unidirectional edges.

[0021] Optionally, constructing an intra-flow traffic interaction graph feature extractor includes:

[0022] The intra-flow traffic interaction graph feature extractor is composed of an embedding layer, a positional encoding layer, three layers of GAT, and a global attention pooling layer;

[0023] The first layer is the embedding layer, which is used to convert the packet size into a vector representation through an embedding matrix;

[0024] Through the positional encoding layer, relative and absolute positional encodings are added to obtain positional information;

[0025] The second, third, and fourth layers respectively use GAT for graph convolution operations, and graph features are extracted through a global attention mechanism after each layer of convolution;

[0026] The features obtained from each layer of convolution are concatenated to fuse feature information at different levels;

[0027] The fifth layer obtains the local feature information of the flow through global attention pooling.

[0028] Optionally, constructing an inter-flow spatio-temporal graph feature extractor includes:

[0029] The inter-flow spatio-temporal graph feature extractor is composed of a positional encoding layer, three layers of RGAT, and a global attention pooling layer;

[0030] Use the local feature information of the flow as the initial node features in the spatio-temporal graph, and obtain the position information through the position encoding layer by combining relative and absolute position encodings;

[0031] In the second, third, and fourth layers, RGAT is used to perform relational attention convolution on five different types of edges to obtain the spatio-temporal feature information of the flow nodes;

[0032] The fifth layer obtains the spatio-temporal feature information of the flow through the global attention pooling layer.

[0033] Optionally, obtaining the multi-scale fusion features according to the multi-scale feature fusion module includes:

[0034] The multi-scale feature fusion module consists of a bilinear transformation layer and a weighted fusion module:

[0035] Concatenate the flow local features and the flow spatio-temporal features extracted based on the local feature information of the flow and the spatio-temporal feature information of the flow respectively, and calculate the weights through the bilinear transformation layer;

[0036] Based on the weighted fusion module, fuse the flow local features and the flow spatio-temporal features through the weights to obtain the multi-scale fusion features.

[0037] Optionally, training the single-classification objective function based on Deep SVDD based on the multi-scale fusion features in combination with the normal traffic dataset and the malicious traffic dataset includes:

[0038] Iteratively optimize the minimum enclosing hypersphere based on the soft-boundary objective function and feedback to optimize the flow local feature extractor and the flow spatio-temporal feature extractor to obtain the optimal Deep SVDD single-classification model.

[0039] Optionally, detecting the traffic category based on the classification result includes:

[0040] If the classification result of the malicious traffic single-classification model is the positive class, it is determined as malicious traffic;

[0041] If the classification result of the normal traffic single-classification model is the positive class and the classification result of the malicious traffic single-classification model is the negative class, it is determined as normal traffic;

[0042] If the classification results of both models are negative classes, it is determined as unknown traffic.

[0043] Technical effects of the present invention: The present invention discloses an encrypted malicious traffic detection method based on a multi-scale spatio-temporal interaction graph network. Traffic source data is obtained, flows are delimited according to the five-tuple rule, and information such as packet size, direction, and capture time is extracted for each flow to construct a traffic detection data set. An intra-flow interaction graph is constructed to extract local features of the flow, and an inter-flow spatio-temporal graph is constructed to extract global features of the flow. The fusion features obtained using the multi-scale feature fusion module are used to train a normal traffic classification model and a malicious traffic classification model based on Deep SVDD. The traffic to be tested is input into the two trained classification models respectively, and the classification results of the two models are jointly used to predict the category of the traffic to be tested, improving the accuracy of malicious traffic identification. Description of the Drawings

[0044] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:

[0045] Figure 1 It is a schematic flowchart of a method for detecting encrypted malicious traffic based on a multi-scale spatio-temporal interaction graph network according to an embodiment of the present invention;

[0046] Figure 2 It is a schematic diagram of the intra-flow interaction graph in an embodiment of the present invention;

[0047] Figure 3 It is a schematic diagram of the multi-scale encrypted malicious traffic feature extractor in an embodiment of the present invention;

[0048] Figure 4 It is a schematic diagram of a single-classification model for malicious traffic and a single-classification model for normal traffic based on Deep SVDD in an embodiment of the present invention. Detailed Embodiments

[0049] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will refer to the drawings and combine the embodiments to detail this application.

[0050] It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0051] As Figure 1 shown, an encrypted malicious traffic detection method based on a multi-scale spatio-temporal interaction graph network is provided in this embodiment, including:

[0052] Obtain traffic source data, and process the traffic source data to obtain flows;

[0053] Based on the interaction relationship of data packets between the client and the server in the flow, construct an in-flow traffic interaction graph network; based on the in-flow traffic interaction graph network, construct an in-flow traffic interaction graph feature extractor to obtain the local feature information of the flow;

[0054] Based on the temporal and spatial relationships between flows, construct an inter-flow spatio-temporal graph; based on the inter-flow spatio-temporal graph, construct an inter-flow spatio-temporal graph feature extractor to obtain the spatio-temporal feature information of the flow;

[0055] Extract different-scale features based on the local feature information of the flow and the spatio-temporal feature information of the flow, and obtain multi-scale fusion features according to the multi-scale feature fusion module;

[0056] Use normal traffic as positive samples and malicious traffic as negative samples. Based on the multi-scale fusion features, combine the normal traffic dataset and the malicious traffic dataset, and train based on the Deep SVDD single-classification objective function to obtain a normal traffic deep single-classification model;

[0057] Use normal traffic as negative samples and malicious traffic as positive samples. Based on the multi-scale fusion features, combine the normal traffic dataset and the malicious traffic dataset, and train based on the Deep SVDD single-classification objective function to obtain a malicious traffic deep single-classification model;

[0058] Obtain the traffic data to be measured, input it into the normal traffic deep single-classification model and the malicious traffic deep single-classification model respectively to obtain classification results; detect the traffic category based on the classification results.

[0059] Specifically, in this embodiment, first obtain a traffic dataset, construct an in-flow traffic interaction network, construct an in-flow interaction graph feature extractor to obtain the local feature information of the flow; construct an inter-flow spatio-temporal graph, construct an inter-flow spatio-temporal graph feature extractor to obtain the spatio-temporal feature information of the flow; obtain multi-scale fusion features based on the feature fusion module; input normal traffic as positive samples and malicious traffic as negative samples, and train based on the Deep SVDD single-classification objective function to obtain a trained normal traffic deep single-classification model; input normal traffic as negative samples and malicious traffic as positive samples, and train based on the Deep SVDD single-classification objective function to obtain a trained malicious traffic deep single-classification model; input the traffic data to be measured into the trained models, use the classification results of the two models to jointly predict the traffic category, and finally output the classification result of the traffic. This example proposes a method for detecting malicious traffic by combining graph networks, realizing the accurate detection of malicious traffic in web access.

[0060] Furthermore, obtain traffic source data, and process the traffic source data to obtain flows including:

[0061] Construct traffic source data, obtain network traffic data packets, divide the traffic data packets using the five-tuple (source IP address, source IP address port, destination IP address, destination IP address port, protocol) rule, define the data packets with the same five-tuple as a flow, extract information such as packet size, direction, and capture time for each flow, and construct a traffic detection data set, including a normal traffic data set and a malicious traffic data set;

[0062] Further, as Figure 2 shown, constructing an in-flow traffic interaction graph network includes:

[0063] Divide the traffic data packets for all data packets using the five-tuple (source IP address, source IP address port, destination IP address, destination IP address port, protocol) rule, define the data packets with the same five-tuple as a flow, construct an in-flow interaction graph for each flow, where the data packets in the flow are used as nodes, establish edge connections between consecutive data packets within the same consecutive data packet sequence transmitted in the same direction, and connect the first and last data packets in two consecutive data packet sequences transmitted in the same direction respectively.

[0064] Further, constructing an inter-flow spatio-temporal graph includes:

[0065] Each flow is used as a node, and edge connections are established between the current flow and the first k flows closest to its start time, and between other flows with the same source IP address or destination IP address; five different types of edges are constructed according to the IP address, namely both the source IP address and the destination IP address are the same, only the source IP address is the same, only the destination IP address is the same, the source IP address is the same as the destination IP address of another flow, and the destination IP address is the same as the source IP address of another flow; among them, the first three types are bidirectional edges, and the last two types are unidirectional edges.

[0066] Specifically, in this embodiment, divide the data set and construct a graph network. Use CICAndMal2017 as the data set, which contains a total of 2126 samples, and each sample corresponds to an instance of an Android application installed and executed on a mobile phone. For each instance, capture all network traffic during its execution. 100 samples of each type are sampled for encrypted malicious traffic classification training. The remaining samples are sampled at a rate of 5% in a stratified manner for verification, and the remaining samples are used for testing.

[0067] Further, as Figure 3 shown, constructing an in-flow traffic interaction graph feature extractor includes:

[0068] Given a set of in-flow interaction graphs The local feature extractor aims to learn the vector representation F of the local features of the flow L .

[0069] The in-flow traffic interaction graph feature extractor consists of an embedding layer, a positional encoding layer, three layers of GAT, and a global attention pooling layer;

[0070] The first layer is the embedding layer, which is used to convert the packet size X = [x1, x2,..., x n into a vector representation H = E[x] through the embedding matrix E ∈ R V×d ; H ∈ R n×d ;

[0071] Among them, X = [x1, x2,..., x n represents the sequence of packet sizes of length n in a flow, E ∈ R V×d represents the embedding matrix, which converts the input packet size into an embedding vector, V represents the number of different packet sizes in X, and d represents the dimension after embedding. H = E[x], H ∈ R n×d is the result after embedding of X = [x1, x2,..., x n .

[0072] Through the positional encoding layer, relative and absolute positional encodings are added to obtain positional information:

[0073]

[0074] Among them, pos represents the relative or absolute position of the packet in the flow, i represents a certain dimension of the vector, and d model represents the size of the embedding dimension.

[0075] Use the order of the packets in the flow as pos to obtain the relative positional encoding, and use the time difference between each packet and the first packet in the flow as pos to obtain the absolute positional encoding. Add the two positional encodings together as the final positional encoding. Combining relative and absolute positional encodings can capture both the positional order and actual time information in the sequence, enabling the model to better understand the positional relationship between packets and better adapt to the actual time difference for packets with irregular time intervals.

[0076] The second, third, and fourth layers respectively use GAT for graph convolution operations, and extract graph features through the global attention mechanism after each layer of convolution;

[0077] Concatenate the features obtained from each layer of convolution for fusing feature information at different levels:

[0078]

[0079] Among them, l represents the current GAT layer number, is the attention score between node i and node j in the k-th attention head, W kDenote the learnable weight matrix of the k-th attention head, N i Denote the neighbors of node i:

[0080]

[0081]

[0082] F L = concat(F L1 , F L2 , F L3 ),

[0083] where f gate is a linear transformation used to calculate the attention weights of each node and normalize them through softmax, and f feat is a linear transformation used to convert the features of each node into a new feature representation;

[0084] The fifth layer obtains the local feature information of the flow through global attention pooling.

[0085] Furthermore, constructing an inter-flow spatio-temporal graph feature extractor includes:

[0086] Given a spatio-temporal graph G ST , the goal of the spatio-temporal feature extractor is to learn the vector representation F G of the local features of the flow.

[0087] The inter-flow spatio-temporal graph feature extractor consists of a position encoding layer, three layers of RGAT, and a global attention pooling layer;

[0088] Using the local feature information of the flow as the initial node features in the spatio-temporal graph, obtain the position information through the position encoding layer by combining relative and absolute position encodings:

[0089]

[0090] where pos represents the relative or absolute position of the flow, i represents a certain dimension of the vector, and d model represents the size of the embedding dimension.

[0091] Specifically, use the order of the flow in all the captured traffic as pos to obtain the relative position encoding, and use the time difference between each flow and the start time of the first flow as pos to obtain the absolute position encoding. Add the two position encodings together as the final position encoding. Combining relative and absolute position encodings can capture both the position order and the actual time information in the sequence, enabling the model to better understand the position relationship between flows and better adapt to the actual time differences for flows with irregular time intervals.

[0092] The second, third, and fourth layers respectively use RGAT to perform relational attention convolution on five different types of edges to obtain the spatio-temporal feature information of the flow nodes;

[0093] Use three layers of RGAT to extract the spatio-temporal features of the flow nodes:

[0094]

[0095] Among them, is the vector obtained by calculating the node attention of the th layer, is the attention score between node i and node j in the th layer and the kth attention head. is the vector obtained by calculating the relational attention of the th layer, is the attention score between node i and node j in the th layer and relation m, is the learnable weight matrix of the th layer and relation m. r ij is the feature of different types of edges, W m1 , b m1 , W m2 , b m2 , are the weight matrix and bias term, and perform a linear transformation on r ij to calculate the attention score of relation m.

[0096] The fifth layer obtains the spatio-temporal feature information of the flow through the global attention pooling layer.

[0097] Furthermore, according to the multi-scale feature fusion module, the multi-scale fusion features include:

[0098] The multi-scale feature fusion module consists of a bilinear transformation layer and a weighted fusion module:

[0099] Based on the local feature information of the flow and the spatio-temporal feature information of the flow, the flow local feature and the flow spatio-temporal feature are respectively extracted and concatenated, and the weights are calculated through the bilinear transformation layer;

[0100] Based on the weighted fusion module, the flow local feature and the flow spatio-temporal feature are fused through the weights to obtain the multi-scale fusion feature.

[0101] Specifically, the feature fusion module calculates the weight λ based on the flow local feature and the flow spatio-temporal feature:

[0102] λ = bilinear(F L , F G , W),

[0103] Among them, F L is the flow local feature, F G is the flow spatio-temporal feature, and W is the weight matrix in the bilinear layer. The flow local feature F L and the flow spatio-temporal feature FG ,

[0104] F combined F = (1 - λ)·F L + λ·F G .

[0105] Furthermore, based on the multi-scale fusion features, combining the normal traffic dataset and the malicious traffic dataset, the training of the single-classification objective function based on Deep SVDD includes:

[0106] Iteratively optimizing the minimum enclosing hypersphere based on the soft-boundary objective function and feeding back to optimize the flow local feature extractor and the flow spatio-temporal feature extractor to obtain the optimal Deep SVDD single-classification model.

[0107] Specifically, in this example, normal traffic is input as the positive sample and malicious traffic is input as the negative sample, and training is performed based on the DeepSVDD single-classification objective function to obtain a trained normal traffic deep single-classification model; normal traffic is input as the negative sample and malicious traffic is input as the positive sample, and training is performed based on the Deep SVDD single-classification objective function to obtain a trained normal traffic deep single-classification model, as Figure 4 shown:

[0108] Input space:

[0109] Output space:

[0110] Mapping neural network from X to F: There are L ∈ N hidden layers with weights W = {W 1 ,..., W L}, and W l represents the weight of the l ∈ [1,..., L] layer. The goal of Deep SVDD is to learn the network parameters to find the minimum volume enclosing circle with radius R and center c in the output space F.

[0111]

[0112] In the above formula, minimizing R 2 can achieve the minimization of the hypersphere. The second term is used to penalize the points outside the hypersphere, and the hyperparameter vn ∈ (0, 1] controls the volume of the sphere and the degree of relaxation of the boundary. The last term is the regularization loss for the training parameters.

[0113] Calculate the distance between the output of the network and the center point c on the training data. If it is less than the threshold R, it is considered a positive class, otherwise it is other classes:

[0114]

[0115] Furthermore, detecting the traffic category based on the classification result includes:

[0116] If the classification result of the malicious traffic single classification model is the positive class, it is determined as malicious traffic;

[0117] If the classification result of the normal traffic single classification model is the positive class and the classification result of the malicious traffic single classification model is the negative class, it is determined as normal traffic;

[0118] If the classification results of both models are the negative class, it is determined as unknown traffic.

[0119] The present invention discloses an encrypted malicious traffic detection method based on a multi-scale spatio-temporal interaction graph network. Traffic source data is obtained, flows are delimited according to the five-tuple rule, and packet size, direction, and capture time information are extracted in units of flows to construct a traffic detection data set. An intra-flow interaction graph is constructed to extract local features of the flow and an inter-flow spatio-temporal graph is constructed to extract global features of the flow. The fusion features obtained by using the multi-scale feature fusion module are used to train a normal traffic classification model and a malicious traffic classification model based on Deep SVDD. The traffic to be detected is respectively input into the two trained classification models, and the classification results of the two models are jointly used to predict the category of the traffic to be detected, improving the accuracy of malicious traffic identification.

[0120] The above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for detecting encrypted malicious traffic based on a multi-scale spatiotemporal interactive graph network, characterized in that: include: Acquire traffic source data, and process based on the traffic source data to obtain a flow; Based on the interactive relationship of data packets between the client and the server in the flow, an intra-flow traffic interaction graph network is constructed; based on the intra-flow traffic interaction graph network, an intra-flow traffic interaction graph feature extractor is constructed to obtain local feature information of the flow; Based on the temporal and spatial relationship between the flows, a spatiotemporal graph between flows is constructed; based on the spatiotemporal graph between flows, a spatiotemporal graph feature extractor between flows is constructed to obtain spatiotemporal feature information of the flows; Extracting different scale features based on the local feature information of the flow and the spatiotemporal feature information of the flow, and obtaining multi-scale fusion features according to a multi-scale feature fusion module; Normal traffic is used as positive samples and malicious traffic is used as negative samples. Based on the multi-scale fusion features, the normal traffic dataset and the malicious traffic dataset are combined to train the Deep SVDD single classification objective function to obtain a normal traffic deep single classification model. Normal traffic is used as negative samples and malicious traffic is used as positive samples. Based on the multi-scale fusion features, the normal traffic dataset and the malicious traffic dataset are combined to train the Deep SVDD single classification objective function to obtain a malicious traffic deep single classification model. Obtain the traffic data to be tested, and input the normal traffic deep single classification model and the malicious traffic deep single classification model respectively to obtain classification results; A traffic class is detected based on the classification result.

2. The encrypted malicious traffic detection method based on a multi-scale spatiotemporal interactive graph network as claimed in claim 1 is characterized in that: Constructing the intra-flow traffic interaction graph network includes: An intra-flow interaction graph is constructed for each flow, in which the packets in the flow are used as nodes, and edges are established between consecutive packets in the same sequence of consecutive packets transmitted in the same direction, and the first and last packets in two sequences of consecutive packets transmitted in the same direction are connected respectively.

3. The encrypted malicious traffic detection method based on a multi-scale spatiotemporal interactive graph network as claimed in claim 1 is characterized in that: Constructing the inter-flow spatiotemporal graph includes: Each flow is taken as a node, and edge connections are established between the current flow and the first k flows closest to its start time, and between other flows with the same source IP address or destination IP address. Five different types of edges are constructed based on the IP address, namely, the source IP address and the destination IP address are the same, only the source IP address is the same, only the destination IP address is the same, the source IP address is the same as the destination IP address of another flow, and the destination IP address is the same as the source IP address of another flow. Among them, the first three types are bidirectional edges, and the last two types are unidirectional edges.

4. The encrypted malicious traffic detection method based on a multi-scale spatiotemporal interactive graph network as claimed in claim 1 is characterized in that: Constructing the in-flow traffic interaction graph feature extractor includes: The in-flow traffic interaction graph feature extractor consists of an embedding layer, a position encoding layer, a three-layer GAT and a global attention pooling layer; The first layer is the embedding layer, which is used to convert the packet size into a vector representation through the embedding matrix; Through the position encoding layer, relative and absolute position encoding are added to obtain position information; The second, third, and fourth layers use GAT to perform graph convolution operations, and extract graph features through the global attention mechanism after each convolution layer; The features obtained from each layer of convolution are concatenated to fuse feature information at different levels; The fifth layer obtains the local feature information of the flow through global attention pooling.

5. The encrypted malicious traffic detection method based on a multi-scale spatiotemporal interactive graph network as claimed in claim 1 is characterized in that: Building an inter-stream spatiotemporal graph feature extractor includes: The inter-stream spatiotemporal graph feature extractor consists of a position encoding layer, a three-layer RGAT, and a global attention pooling layer; Use the local feature information of the flow as the initial node feature in the spatiotemporal graph, and obtain the position information by combining relative and absolute position encoding through the position encoding layer; The second, third, and fourth layers use RGAT to perform relational attention convolution on five different types of edges to obtain the spatiotemporal feature information of the flow nodes; The fifth layer obtains the spatiotemporal feature information of the flow through the global attention pooling layer.

6. The encrypted malicious traffic detection method based on a multi-scale spatiotemporal interactive graph network as claimed in claim 1 is characterized in that: The multi-scale fusion features obtained according to the multi-scale feature fusion module include: The multi-scale feature fusion module consists of a bilinear transformation layer and a weighted fusion module: Based on the local feature information of the stream and the spatiotemporal feature information of the stream, respectively extracting the local features of the stream and the spatiotemporal features of the stream for splicing, and calculating the weights through a bilinear transformation layer; Based on the weighted fusion module, the flow local features and the flow spatiotemporal features are fused by weight to obtain multi-scale fusion features.

7. The encrypted malicious traffic detection method based on a multi-scale spatiotemporal interactive graph network as claimed in claim 1 is characterized in that: Based on the multi-scale fusion features combined with the normal traffic data set and the malicious traffic data set, the training of the DeepSVDD-based single classification objective function includes: Based on the soft-boundary objective function, the minimum enclosing hypersphere is iteratively optimized and the stream local feature extractor and stream spatiotemporal feature extractor are feedback optimized to obtain the optimal Deep SVDD single classification model.

8. The method for detecting encrypted malicious traffic based on a multi-scale spatiotemporal interactive graph network as claimed in claim 1, characterized in that: Detecting traffic categories based on the classification results includes: If the classification result of the malicious traffic single classification model is positive, it is judged as malicious traffic; If the classification result of the normal traffic single classification model is positive, and the classification result of the malicious traffic single classification model is negative, then it is judged as normal traffic; If the classification results of the two models are both negative, it is judged as unknown traffic.

Citation Information

Patent Citations

  • Malicious encrypted traffic detection method based on graph convolutional network

    CN115174169A

  • Malicious traffic detection method based on integral space-time diagram convolutional neural network fused with space-time attention

    CN117579290A