An unknown traffic clustering method based on a graph propagation model
Through the method based on the graph propagation model, a directed graph is created and high-dimensional features are extracted using graph convolutional neural networks, which solves the problem of low clustering of unknown traffic in the existing technology, and achieves more efficient unknown traffic recognition.
Patent Information
- Application Number
- CN202411062761.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-05
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2044-08-05
AI Technical Summary
The existing traffic clustering methods cannot effectively identify unknown traffic, and clustering accuracy depends on feature selection.
Using a graph propagation model-based method, we use directed graphs, extract high-dimensional features using graph convolutional neural networks, and build clustered neural networks in complex networks to maximize spectral modules to perform unknown traffic clustering.
It improves the clustering accuracy of unknown traffic, enhances the effectiveness of node feature representation, and achieves more efficient unknown traffic recognition.
Smart Images

Figure CN119004153B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of network security, and particularly relates to a method for clustering unknown traffic based on a graph propagation model. Background Art
[0002] Network traffic analysis refers to obtaining traffic at the network gateway and conducting analysis. Among them, traffic clustering analysis is the basis for applications such as threat traffic detection and unknown traffic identification. Currently, there are mainly two directions for traffic clustering analysis. One is to conduct traffic identification and clustering based on port, protocol, payload characteristics, etc. The clustering accuracy is high, but it cannot cluster unknown traffic. The other is to conduct traffic clustering based on traffic behavior statistical characteristics, which can make up for the defect that unknown traffic cannot be clustered, but the clustering accuracy depends on feature selection. Using neural network methods can improve the effectiveness of feature selection. Summary of the Invention
[0003] (1) Technical Problems to be Solved
[0004] The technical problem to be solved by the present invention is how to provide a method for clustering unknown traffic based on a graph propagation model to solve...
[0005] (2) Technical Solutions
[0006] To solve the above technical problems, the present invention proposes a method for clustering unknown traffic based on a graph propagation model. The method includes the following steps:
[0007] S1. Traffic acquisition and preprocessing
[0008] Step S11. Flow acquisition and processing
[0009] Mirror the traffic of the network backbone node router and divert the mirrored traffic to the analysis server. Build a virtual machine on the analysis server and use DPDK to obtain high-speed real-time traffic data. Merge the data streams, delete the known service traffic data and noise data, and extract traffic features according to the selected traffic data;
[0010] Step S12. Create a flow graph
[0011] Create a directed graph G=(V, E), where the node V is the IP address and E is the edge set; if there is packet interaction between two nodes, add an edge between the two nodes, and the direction of the edge is from the source IP address to the destination IP address. Obtain the traffic features of the edge and perform normalization processing on the obtained traffic features;
[0012] Step S13. Complex network determination
[0013] Obtain the adjacency matrix A based on the directed graph G, initialize the distance matrix As, update the distance matrix using the Floyd algorithm, calculate the average shortest path L according to the distance matrix As, and determine whether the directed graph G is a complex network based on the average shortest path L;
[0014] S2. Unknown traffic clustering method
[0015] S21. Traffic feature fusion
[0016] Use the message passing mechanism to transfer the eigenvector of the edge to the feature of the node to enhance the effectiveness of the node representation;
[0017] S22. High-dimensional feature extraction
[0018] Use the graph convolutional neural network to extract high-dimensional features;
[0019] S23. Unknown traffic clustering
[0020] On the premise that the graph structure is determined to be a complex network, construct a clustering neural network structure. The clustering neural network structure returns the class probability of each node, and define the loss function as maximizing the spectral modularity. (III) Beneficial effects
[0021] The present invention proposes an unknown traffic clustering method based on a graph propagation model. The present invention specifically uses the method of graph neural network to carry out traffic clustering research. The innovation lies in using the graph propagation model to integrate traffic features into node features, using the graph convolutional neural network to extract high-dimensional features, and aiming at modularity optimization to improve the accuracy of unknown traffic clustering. Description of the drawings
[0022] Figure 1 It is a flow chart of the unknown traffic clustering method based on the graph propagation model of the present invention. Detailed implementation manners
[0023] To make the objectives, contents and advantages of the present invention clearer, the following further describes in detail the specific implementation manners of the present invention with reference to the drawings and embodiments.
[0024] The present invention proposes an unknown traffic clustering method based on a graph propagation model, and the method includes the following steps:
[0025] S1. Traffic acquisition and preprocessing
[0026] Step S11. Flow acquisition and processing
[0027] Mirror the traffic of the network backbone node router, draw the mirrored traffic to the analysis server, set up a virtual machine on the analysis server, and use DPDK to obtain high-speed real-time traffic data, merge the data streams, delete the known service traffic data and noise data, and extract traffic features according to the selected traffic data;
[0028] Step S12, Create a flow graph
[0029] Create a directed graph G = (V, E), where the nodes V are IP addresses and E is the set of edges; if there is packet interaction between two nodes, add an edge between these two nodes, and the direction of the edge is from the source IP address to the destination IP address. Obtain the traffic characteristics of the edge and perform normalization processing on the obtained traffic characteristics;
[0030] Step S13, Complex network determination
[0031] Obtain the adjacency matrix A based on the directed graph G, and initialize the distance matrix As. Use the Floyd algorithm to update the distance matrix. According to the distance matrix As, calculate the average shortest path L, and determine whether the directed graph G is a complex network based on the average shortest path L;
[0032] S2. Unknown traffic clustering method
[0033] S21, Traffic characterization fusion
[0034] Use the message passing mechanism to transfer the feature vector of the edge to the feature of the node to enhance the effectiveness of the node representation;
[0035] S22, High-dimensional feature extraction
[0036] Use the graph convolutional neural network to extract high-dimensional features;
[0037] S23, Unknown traffic clustering
[0038] On the premise that the measured graph structure is a complex network, construct a clustering neural network structure. The clustering neural network structure returns the class probability of each node, and define the loss function as maximizing the spectral modularity.
[0039] Example 1:
[0040] S1. Traffic acquisition and preprocessing.
[0041] Step S11, Flow acquisition and processing.
[0042] Mirror the traffic of the network backbone node router and divert the mirrored traffic to the analysis server. Set up a virtual machine on the analysis server and use DPDK to obtain high-speed real-time traffic data. Merge the packets with the same protocol, source IP, destination IP, source port, and destination port within a fixed time threshold into a data stream.
[0043] Delete the known service traffic data and noise data: specifically, screen the traffic data with the number of packets sent > 5 and the application type that cannot be recognized by the open-source tool, and further delete the known service traffic according to the service whitelist.
[0044] Finally, according to the filtered traffic data, traffic characteristics are extracted, including the number of packets sent, the number of packets received, the interval between packet sending times, the average size of packets sent, the maximum size of packets sent, the interval between packet receiving times, the average size of packets received, the maximum size of packets received, the number of two-way syn packets, the number of two-way fin packets, the number of two-way ack packets, and so on.
[0045] Step S12: Create a flow graph.
[0046] Create a directed graph G=(V, E), where the nodes V are IP addresses and E is the set of edges. If there is packet interaction between two nodes, an edge is added between these two nodes, and the direction of the edge is from the source IP address to the destination IP address. The method for obtaining the traffic characteristics of the edge is as follows: Based on the traffic characteristics extracted in step S1.1, select the traffic characteristics with the source IP and destination IP being Vi and Vj respectively; for the case where there may be multiple flows between Vi and Vj, perform a merging process, and the traffic characteristics take the average value of multiple flows; finally, the obtained traffic characteristics are normalized to prevent the situation where individual characteristic values are too large or too small.
[0047] Step S13: Complex network determination.
[0048] A complex network refers to a network with some or all of the properties of self-organization, self-similarity, attractors, small world, and scale-free. A small world means that although the network is large, there is a relatively short path between any two nodes, that is, the average path length grows logarithmically with the network scale. The modularity optimization method can be applied to complex networks.
[0049] Obtain the adjacency matrix A from the directed graph G and initialize the distance matrix As. The adjacency matrix refers to a two-dimensional matrix storing the adjacent relationship between vertices. If two points are reachable in the adjacency matrix A, the distance value between these two points in the distance matrix As is set to 1; if two points in the adjacency matrix A are not reachable, the distance value between these two points in the distance matrix As is set to infinity; the distance value between the same point in the distance matrix As is 0. Use the Floyd algorithm to find the shortest path between any two points and update the distance matrix. The Floyd algorithm is a dynamic programming method, and its core idea is: As[i,j]=min{As[i,k]+As[k,j], As[i,j]}, where As[i,j] is the distance value between points i and j in the matrix As.
[0050] According to the distance matrix As, calculate the average shortest path L. L is the average value of the sum of distances between any two points in the distance matrix As. The calculation method is: L = s / (n * (n - 1)), where s is the sum of distances between any two points in the distance matrix As (ignoring unreachable paths), and n is the number of reachable nodes in the distance matrix As. m is the total number of nodes in graph G. If L ≈ ln(m), it is considered that the network has the small-world property, so it can be identified as a complex network, and the modularity optimization method can be used to carry out clustering analysis.
[0051] S2. Unknown traffic clustering method
[0052] S21. Traffic feature fusion:
[0053] The graph propagation model is the message passing mechanism, which updates the features of each node based on the graph structure using aggregation operations. The present invention uses the message passing mechanism to transfer the feature vector of the edge to the feature of the node to enhance the effectiveness of node representation and facilitate subsequent node clustering operations. The node feature after fusing the edge feature is:
[0054] H = f[B in X, B out X]
[0055] where the [,] operation refers to the concatenation of two vectors, and the concatenated vector is input to f(x); f(x) = q(xw + b), where q, w, and b are trainable parameters, H is the node feature vector, and X is the edge feature vector.
[0056]
[0057] where V is the set of nodes, E is the set of edges, v k is any node k, and e j is any edge j; the association relationship between the node and the edge is constructed through the B in , B out matrices. If a node has an incoming edge (i.e., the node is the destination IP of the edge), then the relevant element item of B in is set to 1, otherwise 0; similarly, if a node has an outgoing edge (i.e., the node is the source IP of the edge), then the relevant element item of B out is set to 1, otherwise 0. Finally, through the dot product operation of B in , B out and the edge feature vector X, the traffic feature is fused into the node feature.
[0058] S22. High-dimensional feature extraction:
[0059] The graph convolutional neural network (GCN) can be expressed as:
[0060]
[0061] Among them, A is the adjacency matrix, I is the identity matrix, is the degree matrix, W is the weight parameter, and σ is the activation function.
[0062] The present invention extracts high-dimensional features through a graph convolutional neural network.
[0063] S23. Clustering of unknown traffic:
[0064] On the premise that the measured graph structure is a complex network, the loss function is defined as maximizing the spectral modularity, where the spectral modularity is defined as: d is the degree vector, m is the number of nodes, Tr() is the rank, and C is the clustering probability matrix.
[0065] The clustering probability matrix C obtained by the present invention through constructing a clustering neural network structure is:
[0066] C = softmax(GCN(A, H) + Dense(H))
[0067] Among them, GCN() refers to the above-mentioned graph convolutional neural network, and Dense() refers to the fully connected layer, and finally returns the class probability of each node.
[0068] The present invention specifically adopts the method of a graph neural network to carry out traffic clustering research. The innovation lies in integrating traffic features into node features by using a graph propagation model, extracting high-dimensional features by using a graph convolutional neural network, and aiming at modularity optimization to improve the accuracy of unknown traffic clustering.
[0069] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and deformations can still be made, and these improvements and deformations should also be regarded as the protection scope of the present invention.
Claims
1. A method for clustering unknown traffic based on a graph propagation model, characterized in that The method includes the following steps: S1. Traffic acquisition and preprocessing Step S11. Flow acquisition and processing Mirror the traffic of the network backbone node router, and divert the mirrored traffic to the analysis server. Set up a virtual machine on the analysis server, and use DPDK to obtain high-speed real-time traffic data, merge the data streams, delete the known service traffic data and noise data, and extract traffic features according to the selected traffic data; Step S12. Create a flow graph Create a directed graph G=(V, E), where the node V is the IP address and E is the edge set; if there is packet interaction between two nodes, add an edge between the two nodes, and the direction of the edge is from the source IP address to the destination IP address, obtain the traffic feature of the edge, and perform normalization processing on the obtained traffic feature; Step S13. Complex network determination Obtain the adjacency matrix A based on the directed graph G, and initialize the distance matrix As. Use the Floyd algorithm to update the distance matrix. According to the distance matrix As, calculate the average shortest path L, and judge whether the directed graph G is a complex network according to the average shortest path L; S2. Unknown traffic clustering method S21. Traffic feature fusion Use the message passing mechanism to transfer the feature vector of the edge to the feature of the node to enhance the effectiveness of the node representation; S22. High-dimensional feature extraction Use the graph convolutional neural network to extract high-dimensional features; S23. Unknown traffic clustering On the premise that the graph structure is determined to be a complex network, construct a clustering neural network structure. The clustering neural network structure returns the class probability of each node, and define the loss function as maximizing the spectral modularity; Among them, In the S11, the traffic features include: the number of sent packets, the number of received packets, the sending time interval, the average sending size, the maximum sending packet size, the receiving time interval, the average receiving size, the maximum receiving packet size, the number of bidirectional syn packets, the number of bidirectional fin packets, and the number of bidirectional ack packets; In the S13, obtaining the adjacency matrix A based on the directed graph G and initializing the distance matrix As, and using the Floyd algorithm to update the distance matrix includes: obtaining the adjacency matrix A from the directed graph G and initializing the distance matrix As; the adjacency matrix refers to a two-dimensional matrix storing the adjacent relationship between vertices. If two points are reachable in the adjacency matrix A, the distance value between the two points in the distance matrix As is set to 1; if two points in the adjacency matrix A are unreachable, the distance value between the two points in the distance matrix As is set to infinity; the distance value between the same point in the distance matrix As is 0; use the Floyd algorithm to find the shortest path between any two points and update the distance matrix. The Floyd algorithm is a dynamic programming method, and the core idea is: As[i,j]=min{As[i,k]+As[k,j],As[i,j]}, where As[i,j] is the distance value between points i and j in the matrix As.
2. The method for clustering unknown traffic based on a graph propagation model according to claim 1, wherein In the S11, the data packets with the same protocol, source IP, destination IP, source port, and destination port within a fixed time threshold are merged into a data stream.
3. The unknown traffic clustering method based on the graph propagation model according to claim 1, characterized in that, In S11, deleting the known service traffic data and noise data specifically means: screening the traffic data with the number of sent packets > 5 and the application type of which cannot be recognized by the open-source tool, and further deleting the known service traffic according to the service whitelist.
4. The method for clustering unknown traffic based on a graph propagation model according to claim 1, wherein In S12, the method for obtaining the traffic characteristics of edges is as follows: based on the traffic characteristics extracted in step S1.1, select the traffic characteristics with the source IP and destination IP being Vi and Vj respectively; if there are multiple traffic flows between Vi and Vj, merge them, and the traffic characteristics take the average value of the multiple traffic flows.
5. The unknown traffic clustering method based on the graph propagation model according to claim 1, wherein In S13, according to the distance matrix As, calculating the average shortest path L, and judging whether the directed graph G is a complex network according to the average shortest path L specifically includes: according to the distance matrix As, calculating the average shortest path L, where L is the average value of the sum of the distances between any two points in the distance matrix As, and the calculation method is: L = s / (n*(n - 1)), where s is the sum of the distances between any two points in the distance matrix As, n is the number of reachable nodes in the distance matrix As, m is the total number of nodes in the graph G, if L ≈ ln(m), then it is considered that the network has the small-world property, so it can be determined as a complex network.
6. The unknown traffic clustering method based on the graph propagation model according to any one of claims 1-5, characterized in that, In S21, the node characteristics after fusing the edge characteristics are: H = f[B in X, B out X] Among them, the [,] operation refers to the concatenation of two vectors, and the concatenated vector is input into f(x); f(x) = q(xw + b), where q, w, and b are trainable parameters, H is the node feature vector, and X is the edge feature vector; where V is the set of nodes, E is the set of edges, v k is any node k, and e j is any edge j; through B in , B out matrix to construct the association relationship between nodes and edges. If a node has an incoming edge, the relevant element item of B in is set to 1, otherwise 0; similarly, if a node has an outgoing edge, the relevant element item of B out is set to 1, otherwise 0; finally, through the dot product operation of B in , B out and the edge feature vector X, the traffic feature is fused into the node feature.
7. The method for clustering unknown traffic based on a graph propagation model according to claim 6, wherein The graph convolutional neural network GCN in S22 is expressed as: Among them, A is the adjacency matrix, I is the identity matrix, is the degree matrix, W is the weight parameter, and σ is the activation function.
8. The method for clustering unknown traffic based on a graph propagation model according to claim 7, wherein, In S23, The spectral modularity is defined as: where d is the degree vector, m is the number of nodes, Tr() is the rank, and C is the clustering probability matrix; The clustering probability matrix C obtained by constructing the clustering neural network structure is: C = softmax(GCN(A, H) + Dense(H)) Among them, GCN() refers to the above graph convolutional neural network, Dense() refers to the fully connected layer, and finally returns the class probability of each node.
Citation Information
Patent Citations
Multi-type electric vehicle load prediction method and system based on charging behavior
CN114580789A
Electric vehicle ordered charging and discharging method based on V2G technology
CN116362496A