Multi-view-angle-based malicious traffic detection method and device

By building a package time correlation model and a package interaction topology model, obtaining time features and interaction features and fusing them for graph neural network training, the limitations of existing network intrusion detection systems in complex attacks are solved, and detection efficiency and accuracy are improved.

CN120389876APending Publication Date: 2025-07-29MAINTENANCE BRANCH OF STATE GRID FUJIAN ELECTRIC POWER +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510313266.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

Existing network intrusion detection systems have limitations in handling complex attacks, especially the high demand for large-scale high-quality data and high computing resources, and deep learning models have degraded detection performance when dealing with new or variant attacks.

Method used

By extracting data packets in the network flow training set, a packet time correlation model and packet interaction topology model are built, time features and interaction features are obtained, and fused for graph neural network model training to perform malicious traffic detection.

Benefits of technology

It significantly improves the efficiency and accuracy of abnormal detection, and can more carefully detect real-time traffic data in the network space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120389876A_ABST
    Figure CN120389876A_ABST
Patent Text Reader

Abstract

The invention discloses a malicious traffic detection method and device based on multiple views, and the method comprises the steps: extracting data packets and data packet information in a network traffic training set, and obtaining a data packet set; respectively constructing a packet time correlation model and a packet interaction topology model according to the data packet set; time features are obtained according to the packet time correlation model, and interaction features are obtained according to the packet interaction topology model; fusing the time features with the interaction features to obtain comprehensive feature representation; the comprehensive feature representation is used for graph neural network model training, and a trained intrusion detection model is obtained; and detecting to-be-detected network traffic data through the intrusion detection model to obtain a malicious traffic detection result. Intrusion detection of malicious traffic is carried out through the graph neural network by integrating two perspectives of time correlation and interactive topology, the efficiency and accuracy of anomaly detection are improved, and more detailed detection of real-time traffic data in a network space is facilitated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network security technology, and in particular, to a multi-perspective based malicious traffic detection method and device. Background Art

[0002] With the rapid development of information and communication technologies, ensuring the security of interconnected information has become increasingly important. Existing network intrusion detection systems (NIDS) widely apply deep learning technologies to identify malicious traffic. Deep learning models can automatically extract features from raw traffic data and establish complex mapping relationships from input to output, reducing the dependence on manual feature engineering and making system design simpler. However, deep learning models have very high requirements for large-scale high-quality data, and the training process is time-consuming and requires high computing resources. Secondly, deep learning models may face generalization problems and may not be able to maintain high detection performance when dealing with new or variant attacks. Therefore, there may be limitations in dealing with complex attacks. Summary of the Invention

[0003] The technical problem to be solved by the present invention is: to provide a multi-perspective based malicious traffic detection method and device to achieve the detection of malicious traffic existing in network traffic.

[0004] To solve the above technical problem, the technical solution adopted by the present invention is:

[0005] A multi-perspective based malicious traffic detection method, comprising:

[0006] Extracting data packets and data packet information in a network flow training set to obtain a data packet set;

[0007] Respectively constructing a packet time correlation model and a packet interaction topology model according to the data packet set;

[0008] Obtaining time features according to the packet time correlation model and obtaining interaction features according to the packet interaction topology model;

[0009] Fusing the time features and the interaction features to obtain a comprehensive feature representation;

[0010] Using the comprehensive feature representation for training a graph neural network model to obtain a trained intrusion detection model;

[0011] Detecting the network traffic data to be detected through the intrusion detection model to obtain a malicious traffic detection result.

[0012] To solve the above technical problem, another technical solution adopted by the present invention is:

[0013] A malicious traffic detection device based on multiple perspectives, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, each step in the above-mentioned malicious traffic detection method based on multiple perspectives is implemented.

[0014] The beneficial effects of the present invention are as follows: After extracting a data packet set from the network flow training set, a packet time correlation model and a packet interaction topology model are respectively constructed according to the data packet set. Then, time features and interaction features are respectively obtained according to the time correlation model and the packet interaction topology model, and the time features and the interaction features are fused and used for the training of the graph neural network model, so that the graph neural network can obtain the complex relationships and topological information between traffic flows, and intrusion detection of malicious traffic is carried out from the two perspectives of time correlation and interaction topology, thereby significantly improving the efficiency and accuracy of anomaly detection and helping to perform more detailed detection on real-time traffic data in the cyber space. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 It is a flowchart of the steps of a malicious traffic detection method based on multiple perspectives in an embodiment of the present invention;

[0016] Figure 2 It is another flowchart of the steps of a malicious traffic detection method based on multiple perspectives in an embodiment of the present invention;

[0017] Figure 3 It is a schematic diagram of the construction of a packet time correlation graph of a malicious traffic detection method based on multiple perspectives in an embodiment of the present invention;

[0018] Figure 4 It is a schematic diagram of the construction of a packet tree-like topology graph of a malicious traffic detection method based on multiple perspectives in an embodiment of the present invention;

[0019] Figure 5 It is a schematic diagram of a graph neural network model of a malicious traffic detection method based on multiple perspectives in an embodiment of the present invention;

[0020] Figure 6 It is a schematic diagram of the structure of a malicious traffic detection device based on multiple perspectives in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] To describe the technical content, the achieved objectives and the effects of the present invention in detail, the following is described in conjunction with the embodiments and with reference to the accompanying drawings.

[0022] When existing detection methods process network traffic data, they often treat traffic as independent entities, only considering the topological structure information between network traffic flows and lacking awareness of the temporal correlation of network traffic. Due to the neglect of the correlation between traffic flows and deeper structural information, there may be limitations in dealing with complex attacks.

[0023] To address the above technical problems, the present technical solution proposes the following technical measures:

[0024] A malicious traffic detection method based on multiple perspectives, comprising:

[0025] Extract the data packets and data packet information in the network flow training set to obtain a data packet set;

[0026] Construct a packet time correlation model and a packet interaction topology model respectively according to the data packet set;

[0027] Obtain time features according to the packet time correlation model and interaction features according to the packet interaction topology model;

[0028] Fuse the time features and the interaction features to obtain a comprehensive feature representation;

[0029] Use the comprehensive feature representation for the training of a graph neural network model to obtain a trained intrusion detection model;

[0030] Detect the network traffic data to be detected through the intrusion detection model to obtain a malicious traffic detection result.

[0031] As can be seen from the above description, the beneficial effects of the present invention are as follows: After extracting the data packets and data packet signals from the network flow training set as a data packet set, a packet time correlation model and a packet interaction topology model are constructed respectively according to the data packet set, and then time features and interaction features are obtained respectively according to the time correlation model and the packet interaction topology model, and the time features and the interaction features are fused and used for the training of the graph neural network model, so that the graph neural network can obtain the complex relationships and topological information between traffic flows, and perform intrusion detection of malicious traffic from the two perspectives of time correlation and interaction topology, thereby significantly improving the efficiency and accuracy of anomaly detection and helping to perform more detailed detection on real-time traffic data in the network space.

[0032] Further, the extraction of the data packets in the network flow training set includes:

[0033] Extract the data packets of each network flow in the network flow training set through a data packet extraction tool, and use all the data packets extracted from the target network flow as the target data packet flow;

[0034] Delete the invalid data in each data packet;

[0035] Eliminate the target data packet flows with the number of data packets less than the preset number to obtain the data packet set

[0036] As can be seen from the above description, by deleting the invalid data with missing or abnormal data and eliminating the target data packet flows with the number of data packets less than the preset number, it is possible to prevent the model from learning biased information.

[0037] Further, the data packet information includes payload information;

[0038] The extraction of the data packet information in the network flow training set further includes:

[0039] Store the payload information at a preset byte position within the data packet.

[0040] As can be seen from the above description, by storing the payload information at a preset byte position within the data packet, it is possible to quickly obtain the payload information of the data packet during subsequent data packet processing, improving the detection efficiency.

[0041] Further, constructing the packet time association model includes:

[0042] Take each of the data packets as a node and use the data packet information as the node feature;

[0043] Obtain the timestamp of each data packet, and based on the K-nearest neighbor strategy of time proximity, determine a preset number of neighboring nodes closest to its timestamp for each node;

[0044] Connect the target node to all its adjacent neighboring nodes to form a directed edge;

[0045] Calculate the time difference between the target node and all its adjacent neighboring nodes, determine the direction of the directed edge according to the time difference, and use the time difference as the weight of the directed edge;

[0046] Obtain the packet time association model according to all the directed edges.

[0047] As can be seen from the above description, taking each data packet as a node, determining some neighboring nodes closest to its timestamp for each node, and establishing a directed edge according to the time difference between the target node and the neighboring nodes, it is possible to accurately describe the time relationship between each node.

[0048] Further, obtaining the time feature according to the packet time association model includes:

[0049] Construct a packet time correlation processing model, where the packet time correlation processing model includes a time graph embedding layer and a global aggregation layer, and the time graph embedding layer contains a time feature mapping function and a batch normalization function;

[0050] The packet time feature mapping function is used to learn node features based on the timestamps in the packet time correlation model to obtain the packet sequence relationship;

[0051] The batch normalization function is used to adaptively adjust the distribution of the input node features during training, so that the node features remain consistent when transmitted between different time graph embedding layers;

[0052] The time graph embedding layer aggregates the node features of neighbor nodes according to the packet sequence relationship to update the node features of the target node;

[0053] The global aggregation layer performs weighted aggregation on all the node features to obtain the time features.

[0054] As can be seen from the above description, by sequentially processing the input packet time correlation model through multiple time graph embedding layers, the sequence relationship between each data packet in the packet time correlation model can be effectively extracted, and then the global aggregation layer performs weighted aggregation on all node features to form time features, so that the time relationship between each data packet in the network traffic can be accurately described based on the time features.

[0055] Furthermore, constructing a packet interaction topology model includes:

[0056] The data packet information includes the transmission direction and the packet size;

[0057] Take each data packet as a node, and use the transmission direction and the packet size as node features;

[0058] Set the continuous and co-directional data packets to the same level;

[0059] Fully connect all the data packets within each level to all the data packets within the previous level to form inter-level edges;

[0060] Obtain the packet interaction topology model according to all the inter-level edges.

[0061] As described above, by using data packets as nodes, setting consecutive and co-directional data packets at the same level, and then fully connecting the data packets within adjacent layers, a packet tree-like topological structure is formed. The dense subgraph formed by the full connection method enhances the message passing ability between specific layers, can capture local high-density interaction scenarios. In detecting covert channels or encrypted malicious traffic, local dense interactions such as abnormal high-frequency short packet interactions are key signals, and the full connection design is more sensitive to this, can accurately reflect the packet interaction process during network traffic transmission, and detect abnormal data packets among them.

[0062] Further, the obtaining interaction features according to the packet interaction topology model includes:

[0063] Construct a packet interaction topology processing model, where the packet interaction topology processing model includes a multi-scale graph convolutional attention network layer and an average pooling layer;

[0064] Obtain the neighborhood information of the packet interaction topology model, and normalize the neighborhood information;

[0065] Input the node features in the packet interaction topology model and the normalized neighborhood information into the multi-scale graph convolutional attention network layer;

[0066] The multi-scale graph convolutional attention network layer aggregates the node features of the nodes adjacent to the target node according to the neighborhood information, and performs adaptive weighting to obtain the node features corresponding to the target node;

[0067] Aggregate all the node features through the average pooling layer to obtain the interaction features.

[0068] As described above, by sequentially processing the packet interaction topology model through the multi-scale graph convolutional attention network layer and the average pooling layer, the packet interaction topology model can be effectively transformed into interaction features, so that the transmission relationship between each data packet in the network traffic can be accurately described based on the interaction features.

[0069] Further, the fusing the time feature and the interaction feature to obtain a comprehensive feature representation includes:

[0070] Obtain a first weight value of the time feature and a second weight value of the interaction feature;

[0071] Perform weighted fusion calculation according to the time feature, the first weight value, the interaction feature, and the second weight value to obtain the comprehensive feature representation.

[0072] As described above, by assigning different weight values to the time feature and the interaction feature and then performing weighted fusion, the corresponding weight values can be adjusted for different detection scenarios, thereby improving the detection effect of malicious traffic in different scenarios.

[0073] Further, the using the comprehensive feature representation for training a graph neural network model to obtain a trained intrusion detection model includes:

[0074] The graph neural network model includes a fully connected layer and a cross-entropy loss function;

[0075] Performing a linear transformation on the comprehensive feature representation through the fully connected layer, and using the Softmax function to predict the network traffic category:

[0076] Y i = Softmax(Z);

[0077] The cross-entropy loss function includes:

[0078]

[0079] where N represents the total number of samples, K represents the number of categories in the samples, y i,k is the true value of the i-th sample for the k-th class label, is the probability that the model predicts the i-th sample belongs to the k-th class; Z is the comprehensive feature representation.

[0080] As described above, the comprehensive feature is obtained by fusing the time feature and the interaction feature through the packet time correlation model and the packet interaction topology model. By introducing the cross-entropy loss function, the parameters of the model can be optimized, so that the model outputs more accurate prediction results.

[0081] Another embodiment of the present invention provides a malicious traffic detection device based on multiple perspectives, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements each step in the above-mentioned malicious traffic detection method based on multiple perspectives.

[0082] The malicious traffic detection method and device provided by the present invention can be applied to the scenario of detecting malicious traffic in network traffic, which will be described below through specific embodiments:

[0083] Embodiment 1

[0084] Please refer to Figure 1 , a malicious traffic detection method based on multiple perspectives, including:

[0085] S1. Extract the data packets and data packet information in the network flow training set to obtain a data packet set. Specifically:

[0086] S11. Use a data packet extraction tool to extract the data packets of each network flow in the network flow training set, and regard all the data packets extracted from the target network flow as the target data packet stream. For example, use tools such as flowcontainer tool or SplitCap to extract the data packets, extract the data packet length sequence information and the payload information of the data packets, and extract N data packets from each network flow as the representation of the network flow. In this way, the amount of data to be processed can be greatly reduced, and the efficiency of real-time detection and analysis can be improved.

[0087] S12. Delete the invalid data in each data packet. All network traffic data packets contain a network five-tuple: source IP address, source port, destination IP address, destination port, and protocol. When judging the invalid data, first check the field integrity of the five-tuple, and then check the time continuity of the connection of the same five-tuple, so as to delete the missing and abnormal values in the data.

[0088] S13. Eliminate the target data packet streams with the number of data packets less than the preset number to obtain the data packet set. For example, eliminate the network flow data with the number of data packets less than 3 to avoid the model learning biased information. Through the processing of steps S12 - S13, empty packets, retransmitted packets, and network flows with too few data packets can be removed, thus masking the biased data packet fields.

[0089] S14. Process the payload information of the data packets and store the payload information in the preset byte position in the data packets. As the optimal implementation method, retain the first preset number of bytes in the data packet as the payload representation of the data packet, because the information it contains is the most representative and can ensure the detection accuracy to the greatest extent; the number of retained bytes can be adjusted according to the actual application, but does not exceed 128; for example, in this embodiment, each data packet retains the first 32 bytes as the payload representation of the data packet to improve the detection efficiency. In other optional ways, the number of retained bytes and the position can also be determined by selecting position markers and dynamic adjustment methods, but the implementation process is complex and will have a greater impact on the detection efficiency and cost. At the same time, normalize the information in the network traffic data, such as data packet length, direction, and data packet payload content, so as to generate data available for model training; that is, regard the obtained information as the features in the graph data, input and use the graph processing method to generate graph data available for the model.

[0090] S2. Construct a packet time correlation model and a packet interaction topology model respectively according to the data packet set. Specifically:

[0091] Please refer to Figure 3 , and constructing a packet time correlation model includes the following steps:

[0092] S21a. Take each of the said data packets as a node, and take the data packet information as the node feature. That is, map the extracted data packets to nodes, and after normalizing the original bytes in each data packet, store them in the nodes as node features; for example, regard each data packet in the network as an independent node, denoted as p i , where V = {p0, p1,..., p i}, V is used to represent the set of all nodes, use the payload bytes of the data packet as the feature representation of the node, and set the length of the feature vector to 1500 to match the capacity of the maximum transmission unit (MTU, Maximum Transmission Unit); MTU defines the maximum data unit size that the data link layer can transmit. In the subsequent process of analyzing network data, setting the dimension consistent with the MTU size can make the features learned by the model more naturally mapped to the real network scenario and improve the interpretability of the model. As Figure 3 shown, assuming that a flow has 5 data packets, then the figure contains 5 nodes, which are respectively represented as P0, P1, P2, P3, P4.

[0093] S22a. Obtain the timestamp of each of the said data packets, and based on the K-nearest neighbor strategy of time proximity, determine a preset number of neighboring nodes closest to its timestamp for each node. For example, establish an edge connection relationship between nodes according to the timestamp information carried by the data packets, and denote the timestamp as t i ; adopt the K-nearest neighbor (KNN by TimeProximity) strategy based on time proximity to determine the n neighbors closest to its time for each node, that is, by calculating the time difference Δt i between p j and p ij = t j - t i , quantify the time proximity degree between nodes, for each node p i , sort according to Δt ij from small to large, and select the n neighbors closest in time to form its neighbor set P i ={p i1 , p i2 ,..., p in}. As Figure 3As shown in the figure, the timestamps corresponding to the 5 nodes are respectively denoted as t0, t1, t2, t3, and t4; assuming n = 3, that is, the 3 neighbors closest in time are selected for each node; assuming the timestamps are t0 = 1s, t1 = 2s, t2 = 5s, t3 = 6s, t4 = 10s in sequence, which reflects the arrival order of the reaction data packets; according to the time interval, the specific connection relationships are as follows: The neighbors of node P0 are P1, P2, P3, the neighbors of node P1 are P0, P2, P3, the neighbors of node P2 are P0, P1, P3, the neighbors of node P3 are P1, P2, P4, and the neighbors of node P4 are P1, P2, P3.

[0094] S23a. Connect all the neighboring nodes adjacent to the target node to form directed edges; that is, establish directed edges between each neighbor in p i and P i . As shown in Figure 3 , the 5 nodes P0, P1, P2, P3, and P4 are correspondingly connected according to the above relationships.

[0095] S24a. Calculate the time differences between the target node and all the neighboring nodes adjacent to it, determine the directions of the directed edges according to the time differences, and use the time differences as the weights of the directed edges. That is, the directions of the edges are determined by the sequence of timestamps. For example, when t i < t j , the edge points from p i to p j ; and use the time difference Δt ij as the weight of the edge, denoted as ω ij = Δt ij . Each edge is represented as e ij =(p i , p j , ω ij ).

[0096] S25a. Obtain the packet time correlation model according to all the directed edges; that is, the final edge set E is aggregated from the KNN connection relationships of all nodes, and finally a packet time correlation graph is obtained. The edge set E is represented as:

[0097] E ={(P0, P1, 1s), (P0, P2, 4s), (P0, P3, 5s), (P1, P2, 4s)

[0098] , (P1, P3, 4s), (P2, P3, 1s), (P3, P4, 4s)};

[0099] Since in this process, the time sequence and time interval information of the data packets are embedded into the attributes of the edges, it can reflect the time dynamic characteristics of the traffic.

[0100] Please refer toFigure 4 , constructing a packet interaction topology model includes:

[0101] S21b. The packet information includes a transmission direction and a packet size. Since the transmission of network traffic is bidirectional, that is, it is transmitted between a client and a server to achieve information interaction; for this bidirectional characteristic, in this embodiment, the data packet sent from the client to the server is defined as a forward data packet, and the network data packet sent from the server to the client is a reverse data packet, and the direction of the forward data packet is converted into a value of 0, and the direction of the reverse data packet is converted into a value of -1 to model the interaction relationship of the traffic.

[0102] S22b. Each of the data packets is used as a node, and the transmission direction and the packet size are used as node features. Since the packet interaction topology model and the packet time correlation model are two independent models (figures), two separate mappings are required. For example, the data packet is denoted as a node v i , where i = 1..k, and the node set is represented by V; and the transmission direction and the data packet size of the data packet are used as node features; the node features are represented as F = (D, L), where D represents the direction, and the value is 0 or -1; L represents the normalized data packet length, and its value is between 0 and 1. As Figure 4 shown, assuming that a complete network flow is represented by 14 data packets, there are 14 nodes in the figure; these nodes are represented as {52, -52, 40, 174, -40, -162, -40, 198, 40, -160, -189, 1585, 40, 40}.

[0103] S23b. The consecutive and same-direction data packets are set to the same level; that is, the concept of a layer is introduced during the connection of nodes, and the consecutive same-direction data packets are defined as one layer. As Figure 4 shown, we respectively obtain:

[52] , [-52], [40, 174], [-40, -162, -40], [198, 40], [-160, -189], [1585, 40, 40], a total of 7 layers.

[0104] S24b. All the data packets within each level are fully connected to all the data packets within the previous level to form inter-layer edges; for example, between the third layer and the fourth layer, the nodes of the two layers are fully connected, then 40 is respectively connected to three data packets, -40, -162, and -40; similarly, 174 is respectively connected to three data packets, -40, -162, and -40, and so on, finally forming a tree-like structure.

[0105] S25b. Obtain the packet interaction topology model based on all the inter-layer edges; that is, data packets are connected layer by layer according to the inter-layer relationship, and finally a packet tree-like topology graph is formed. Among them, in the related art, there is a graph construction method in which the nodes inside each burst are connected, and the two end points are connected between each burst. This method is called the traffic interaction graph (TIG, Traffic Interaction Graph); in the network scenarios of the real world, network fluctuations are a common phenomenon, and network fluctuations will cause out-of-order or even missing transmission of data packets; therefore, connecting the nodes inside each burst will cause potential out-of-order or missing information to be transmitted to the model, resulting in the model learning negative information and reducing the detection accuracy. Therefore, the graph construction method of TIG connecting the nodes inside the burst will cause detection errors. The packet tree-like topology graph in this embodiment removes this internal connection method of the burst to avoid the error-prone information caused by network fluctuations in the real world, so as to improve the robustness of the entire system.

[0106] At the same time, the graph construction method of TIG connects the two end points between each burst. This method may not be sufficient to capture all interaction details for complex traffic such as multi-threaded downloads and P2P communications. In the packet tree-like topology graph, the dense subgraph formed by the full connection method enhances the message passing ability between specific layers and can capture local high-density interaction scenarios; when detecting covert channels or encrypted malicious traffic, local dense interactions may be key signals such as abnormal high-frequency short packet interactions, and the full connection design is more sensitive to this; thus improving the detection accuracy.

[0107] Please refer to Figure 5 , S3. Obtain time features according to the packet time correlation model and obtain interaction features according to the packet interaction topology model. That is, the intrusion detection model based on the graph neural network includes two parts. The first part is the packet time correlation graph processing model. The input is the relevant features of the packet time correlation graph, such as including the topological structure of the graph, node features, edge features, and global features, and the output is the comprehensive graph representation of the packet time correlation graph, that is, time features. Specifically:

[0108] (1) Obtaining time features according to the packet time correlation model includes:

[0109] S31a. Construct a packet time correlation processing model, where the packet time correlation processing model includes three layers of spatio-temporal graph embedding layers (STGE layer) and a global feature aggregation (GFAF) layer; and each spatio-temporal graph embedding layer contains a time feature mapping function and a batch normalization function. Among them, the packet time feature mapping function is used to learn node features based on the timestamps in the packet time correlation model to capture the temporal relationship between data packets, that is, to obtain the temporal relationship of data packets. The batch normalization function (BatchNorm) is used to adaptively adjust the distribution of the input node features during training to keep the node features consistent when passing between different spatio-temporal graph embedding layers.

[0110] S32a. The spatio-temporal graph embedding layer updates the node features of the target node by aggregating the node features of neighbor nodes according to the temporal relationship of the data packets. That is, each spatio-temporal graph embedding layer adopts a neighborhood aggregation mechanism, which is based on the spatio-temporal characteristics of the graph and updates the node representation by aggregating the features of neighbor nodes; during the node update process of each layer, the features of the node are updated by aggregating the feature information of adjacent nodes, and the specific formula is as follows:

[0111]

[0112] Among them, represents the feature vector of node v in the l-th layer i . represents the set of neighbor nodes of node v i , represents the feature vector of neighbor node u in the (l - 1)-th layer, AGG v represents the node aggregation operation, W and b are learning weights and biases, and φ is a non-linear activation function.

[0113] S33a. The global feature aggregation layer performs weighted aggregation on all the node features to obtain the time features, specifically:

[0114]

[0115] Among them, is the final feature of node v in the L-th layer i , and |V| is the total number of nodes in the graph.

[0116] (2) The second part is the packet interaction topology processing model. The interaction features obtained according to the packet interaction topology model include:

[0117] S31b. Construct a packet interaction topology processing model, where the packet interaction topology processing model includes two layers of multi-scale graph convolutional attention network layers and one layer of average pooling layer; it is used to take the neighborhood information and node features of the packet tree topology graph as inputs and obtain the representation of the entire packet tree topology graph as the output.

[0118] S32b. Obtain the neighborhood information of the packet interaction topology model and perform normalization processing on the neighborhood information:

[0119]

[0120] Among them, represents adding a self-loop to the adjacency matrix, is the degree matrix of, and there is is the normalized data.

[0121] S33b. Input the node features in the packet interaction topology model and the normalized neighborhood information into the multi-scale graph convolutional attention network layer; among them, the multi-scale graph convolutional attention network layer aggregates the node features of the nodes adjacent to the target node according to the neighborhood information and performs adaptive weighting, and can aggregate the information of each neighbor node to obtain the node features corresponding to the target node. The specific formula is as follows:

[0122]

[0123] Among them, represents the convolutional operation based on scale k in the l-th layer, is the weight matrix corresponding to the scale, h v represents the node features learned by the model, σ(·) is a non-linear activation function, and K represents the number of convolutional scales; through the combined application of multi-scale convolution and graph attention mechanism, each layer effectively aggregates and updates the node features; the final features of the nodes are implicitly represented and learned through the graph autoencoder layer, and the encoding process is as follows:

[0124]

[0125] S34b. The global aggregation layer performs weighted aggregation on all the node features to obtain the time features; that is, after feature propagation and update, all the node features are aggregated through the average pooling layer to generate the final graph representation feature Z G :

[0126] Z G = global_mean_pool(z v ).

[0127] S4. Integrate the time feature and the interaction feature to obtain a comprehensive feature representation. Specifically: obtain the first weight value of the time feature and the second weight value of the interaction feature. Here, the first weight value is λ, and the second weight value is (1 - λ). Perform weighted fusion calculation based on the time feature, the first weight value, the interaction feature, and the second weight value to obtain the comprehensive feature representation. The formula is as follows:

[0128] Z = λZ G +(1 - λ)Z T ;

[0129] where λ is a learnable parameter.

[0130] S5. Use the comprehensive feature representation for the training of the graph neural network model to obtain a trained intrusion detection model. Specifically:

[0131] The graph neural network model includes a fully connected layer and a cross - entropy loss function;

[0132] Perform a linear transformation on the comprehensive feature representation through the fully connected layer and use the Softmax function to predict the network traffic category:

[0133] Y i = Softmax(Z);

[0134] The cross - entropy loss function includes:

[0135]

[0136] where N represents the total number of samples, K represents the number of categories in the samples, y i,k is the true value of the i - th sample in the k - th class label, is the probability that the model predicts the i - th sample belongs to the k - th class; Z is the comprehensive feature representation. The cross - entropy loss function is used in the training stage of the model. The model is trained through supervised learning to minimize the difference between the predicted value and the network traffic category (label) in the training set during the training process to achieve the most accurate network traffic detection.

[0137] S6. Detect the network traffic data to be detected through the intrusion detection model to obtain the malicious traffic detection result.

[0138] Embodiment 2

[0139] Please refer to Figure 6, A malicious traffic detection device based on multiple perspectives, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements each step in a malicious traffic detection method based on multiple perspectives as described in Embodiment 1.

[0140] In summary, the malicious traffic detection method and device provided by the present invention, after processing network traffic data, obtain a packet time correlation graph and a packet tree topology graph according to the characteristics of the data packets in the network traffic, and respectively learn the graph representation features of the two graph data through a graph neural network. After fusing the two features, it realizes the efficient prediction of the network traffic category to detect malicious traffic existing in the network traffic. By using the graph neural network technology combined with multiple perspectives, it better combines the time information and topological structure information existing in the network traffic data, thereby realizing a more detailed analysis of the network traffic data from multiple perspectives, which helps to conduct a more detailed detection of real-time traffic data in the cyberspace.

[0141] The above are only the embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in the relevant technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. A malicious traffic detection method based on multiple perspectives, characterized in that, Including: Extracting data packets and data packet information from the network flow training set to obtain a data packet set; Respectively constructing a packet time correlation model and a packet interaction topology model according to the data packet set; Obtaining time features according to the packet time correlation model and obtaining interaction features according to the packet interaction topology model; Fusing the time features and the interaction features to obtain a comprehensive feature representation; Using the comprehensive feature representation for training a graph neural network model to obtain a trained intrusion detection model; Detecting the network traffic data to be detected through the intrusion detection model to obtain a malicious traffic detection result.

2. The malicious traffic detection method based on multiple perspectives according to claim 1, characterized in that, The extracting of data packets from the network flow training set includes: Extracting the data packets of each network flow in the network flow training set through a data packet extraction tool, and taking all the data packets extracted from the target network flow as the target data packet flow; Deleting invalid data in each data packet; Eliminating the target data packet flows with the number of data packets less than a preset number to obtain the data packet set.

3. A multi-perspective based malicious traffic detection method according to claim 2, characterized in that, The data packet information includes payload information; The extracting of data packet information from the network flow training set further includes: Storing the payload information at a preset byte position within the data packet.

4. A multi-perspective based malicious traffic detection method according to claim 1, characterized in that, Constructing a packet time correlation model includes: Taking each data packet as a node and taking the data packet information as the node feature; Obtaining the timestamp of each data packet, and determining a preset number of neighboring nodes closest to its timestamp for each node based on the K-nearest neighbor strategy of time proximity; Connecting the target node with all its adjacent neighboring nodes to form a directed edge; Calculating the time difference between the target node and all its adjacent neighboring nodes, determining the direction of the directed edge according to the time difference, and taking the time difference as the weight of the directed edge; Obtaining the packet time correlation model according to all the directed edges.

5. The malicious traffic detection method based on multiple perspectives according to claim 4, characterized in that The obtaining of time features according to the packet time correlation model includes: Constructing a packet time correlation processing model, the packet time correlation processing model including a time graph embedding layer and a global aggregation layer, and the time graph embedding layer containing a time feature mapping function and a batch normalization function; The packet time feature mapping function is used to learn the node features according to the timestamps in the packet time correlation model to obtain the data packet time sequence relationship; The batch normalization function is used to adaptively adjust the distribution of the input node features during training so that the node features remain consistent when transmitted between different time graph embedding layers; The time graph embedding layer aggregates the node features of neighbor nodes to update the node features of the target node according to the data packet time sequence relationship; The global aggregation layer performs weighted aggregation on all the node features to obtain the time features.

6. A multi-perspective based malicious traffic detection method according to claim 1, characterized in that Constructing a packet interaction topology model includes: The data packet information includes the transmission direction and the packet size; Taking each data packet as a node and taking the transmission direction and the packet size as the node features; Setting the consecutive data packets in the same direction to the same level; Fully connecting all the data packets within each level with all the data packets within the upper level to form inter-level edges; Obtain the packet interaction topology model based on all the inter-layer edges.

7. A multi-perspective-based malicious traffic detection method according to claim 6, characterized in that The obtaining of the interaction features according to the packet interaction topology model includes: Construct a packet interaction topology processing model, where the packet interaction topology processing model includes a multi-scale graph convolutional attention network layer and an average pooling layer; Obtain the neighborhood information of the packet interaction topology model and normalize the neighborhood information; Input the node features in the packet interaction topology model and the normalized neighborhood information into the multi-scale graph convolutional attention network layer; The multi-scale graph convolutional attention network layer aggregates the node features of the nodes adjacent to the target node according to the neighborhood information and performs adaptive weighting to obtain the node features corresponding to the target node; Aggregate all the node features through the average pooling layer to obtain the interaction features.

8. A malicious traffic detection method based on multiple perspectives according to claim 1, characterized in that The fusing of the time features and the interaction features to obtain the comprehensive feature representation includes: Obtain the first weight value of the time features and obtain the second weight value of the interaction features; Perform weighted fusion calculation according to the time features, the first weight value, the interaction features, and the second weight value to obtain the comprehensive feature representation.

9. A malicious traffic detection method based on multiple perspectives according to claim 1, characterized in that The using of the comprehensive feature representation for training a graph neural network model to obtain a trained intrusion detection model includes: The graph neural network model includes a fully connected layer and a cross-entropy loss function; Perform a linear transformation on the comprehensive feature representation through the fully connected layer and use the Softmax function to predict the network traffic category: Y i = Softmax(Z); The cross-entropy loss function includes: where N represents the number of the total samples, K represents the number of classes in the samples, and y i,k is the true value of the i-th sample for the k-th class label, is the probability that the model predicts the i-th sample belongs to the k-th class; Z is the comprehensive feature representation.

10. A malicious traffic detection device based on multiple perspectives, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements each step in a multi-perspective malicious traffic detection method according to any one of claims 1-9.