Malicious traffic identification method and device based on multiple granularities
By preprocessing and feature fusion of network traffic data, and building a data packet association diagram, the problem of difficulty in identifying malicious traffic in encrypted communications is solved by traditional methods, and efficient and accurate malicious traffic detection is achieved.
Patent Information
- Application Number
- CN202510313507.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-29
AI Technical Summary
Existing network traffic detection methods are difficult to effectively identify malicious traffic in encrypted communications, especially when facing advanced persistent threats and challenges of encryption protocols, the detection effect of traditional methods is greatly reduced.
By preprocessing the network traffic data, the header byte information and load byte information of the data packet are extracted, and the byte-level traffic detection model is input, the header and load byte representation characteristics are generated, and the packet association diagram is carried out to construct the network flow-level traffic detection model for malicious traffic detection.
It realizes efficient and accurate identification of malicious traffic, enhances robustness and detection accuracy, and improves the generalization ability and detection effect of the model through multi-grained information fusion.
Smart Images

Figure CN120389877A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security technology, and particularly to a multi-granularity based malicious traffic identification method and device. Background Art
[0002] With the rapid development of Internet technology and the diversification of network applications, the types and complexity of network traffic are constantly increasing, making network traffic detection and identification one of the key technologies to ensure network security and service quality. Especially in preventing malicious traffic, malicious traffic not only includes traditional attack behaviors such as DDoS (Distributed Denial of Service) and SQL (Structured Query Language) injection, but also covers more concealed threats such as Advanced Persistent Threat (APT), ransomware propagation, botnet activities, etc. The detection of these malicious traffic is crucial for realizing effective network resource management, preventing network attacks, and maintaining normal network order.
[0003] With the wide application of encryption technology, especially the popularization of encryption protocols such as SSL (Secure Sockets Layer) / TLS (Transport Layer Security) in daily communication, the privacy protection level of users has been greatly improved, but it also brings unprecedented challenges to the analysis and classification of network traffic. Traditional network traffic classification methods, such as techniques based on port scanning or deep packet inspection (DPI), these traditional methods rely on predefined rules and signature libraries and are easily bypassed. Especially when facing malicious traffic using encrypted communication, their detection effect is greatly reduced and it is difficult to adapt to the increasingly complex network environment. Summary of the Invention
[0004] The technical problem to be solved by the present invention is: to provide a multi-granularity based malicious traffic identification method and device to achieve efficient and accurate malicious traffic detection and classification.
[0005] To solve the above technical problem, the technical solution adopted by the present invention is:
[0006] A multi-granularity based malicious traffic identification method, comprising:
[0007] Preprocessing the network traffic data to be analyzed to obtain network flow data; the network flow data includes data packets;
[0008] Extracting header byte information and payload byte information of each data packet in the network flow data;
[0009] Inputting the header byte information and the payload byte information into a byte-level traffic detection model respectively to obtain a header byte representation feature and a payload byte representation feature;
[0010] The header byte representation feature and the payload byte representation feature are combined to obtain a packet level representation feature;
[0011] Constructing a data packet association graph with the data packets as nodes, and using the data packet level representation features as node features of each node;
[0012] The data packet association graph is input into a network flow-level traffic detection model to obtain a malicious traffic detection result.
[0013] In order to solve the above technical problems, another technical solution adopted by the present invention is:
[0014] A multi-granularity-based malicious traffic identification device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the processor implements the various steps of the multi-granularity-based malicious traffic identification method described above.
[0015] The beneficial effects of the present invention are: after preprocessing the network traffic data to be analyzed to obtain network flow data, the network flow data is extracted to obtain header byte information and payload byte information, and then packet-level representation features are generated based on the header byte information and payload byte information, and a packet association graph is constructed with the data packets as nodes, and the packet-level representation features are used as node features of each node, and finally the malicious traffic detection results are obtained by inputting the packet association graph into the network flow-level traffic detection model, that is, data processing is performed from the byte level, packet level and network flow level in turn, combining multiple granularities of data, and realizing multi-granularity network traffic information fusion, so that the fused node features are more effective, the robustness is enhanced, and the traffic features in the input model are richer, thereby realizing efficient and accurate identification of malicious traffic. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 This is a flowchart of a method for identifying malicious traffic based on multiple granularities according to an embodiment of the present invention;
[0017] Figure 2 This is another step flow chart of a method for identifying malicious traffic based on multiple granularities in an embodiment of the present invention;
[0018] Figure 3Schematic diagram of the byte-level traffic detection model for a multi-granularity-based malicious traffic recognition method in an embodiment of the present invention;
[0019] Figure 4 Schematic diagram of the adaptive feature fusion for a multi-granularity-based malicious traffic recognition method in an embodiment of the present invention;
[0020] Figure 5 Example diagram of the packet association graph for a multi-granularity-based malicious traffic recognition method in an embodiment of the present invention;
[0021] Figure 6 Schematic diagram of the associated attention traffic classifier for a multi-granularity-based malicious traffic recognition method in an embodiment of the present invention;
[0022] Figure 7 Schematic diagram of the structure of a multi-granularity-based malicious traffic recognition device in an embodiment of the present invention. Detailed implementation manners
[0023] To illustrate the technical content, achieved objectives and effects of the present invention in detail, the following is described in conjunction with the implementation manners and accompanied by the drawings.
[0024] Please refer to Figure 1 , a multi-granularity-based malicious traffic recognition method, including:
[0025] Preprocess the network traffic data to be analyzed to obtain network flow data; the network flow data includes data packets;
[0026] Extract the header byte information and payload byte information of each data packet in the network flow data;
[0027] Input the header byte information and payload byte information into the byte-level traffic detection model respectively to obtain the header byte representation feature and the payload byte representation feature;
[0028] Fuse the header byte representation feature and the payload byte representation feature to obtain the packet-level representation feature;
[0029] Construct a packet association graph with the data packet as a node, and use the packet-level representation feature as the node feature of each node;
[0030] Input the packet association graph into the network flow-level traffic detection model to obtain the malicious traffic detection result.
[0031] As can be seen from the above description, the beneficial effects of the present invention are as follows: after preprocessing the network traffic data to be analyzed to obtain network flow data, extracting the header byte information and payload byte information from the network flow data, generating packet-level representation features based on the header byte information and payload byte information, constructing a packet association graph with packets as nodes, and using the packet-level representation features as the node features of each node. Finally, by inputting the packet association graph into the network flow-level traffic detection model, a malicious traffic detection result is obtained, that is, data processing is performed sequentially from the byte level, packet level, and network flow level, combining multiple granularities of data, realizing multi-granularity network traffic information fusion, making the fused node features more effective, enhancing robustness, and making the traffic features input into the model more abundant, thereby realizing efficient and accurate identification of malicious traffic.
[0032] Further, the preprocessing of the network traffic data to be analyzed to obtain network flow data includes:
[0033] Deleting the invalid data in the network traffic data to obtain network valid data;
[0034] Screening and extracting the packets in the network valid data to obtain network screening data;
[0035] Sorting the packets in the network screening data according to the timestamps of the packets to obtain network sorted data;
[0036] Performing data cleaning on the network sorted data to obtain network cleaned data;
[0037] Performing normalization processing on the network cleaned data to obtain normalized data;
[0038] Generating the network flow data from the normalized data through a network traffic processing method.
[0039] As can be seen from the above description, by eliminating the invalid data in the network traffic data and screening and cleaning the packets, not only the invalid data is removed, but also the amount of data to be processed is reduced, thereby improving the efficiency of subsequent data processing.
[0040] Further, the extraction of the header byte information and payload byte information of each packet in the network flow data includes:
[0041] After separating the header information and payload information in the packet, performing the same normalization processing to obtain a header byte sequence and a payload byte sequence of the same length;
[0042] The header byte sequence and payload byte sequence in each data packet are respectively embedded in a preset encoding vector to obtain a header byte matrix and a payload byte matrix respectively.
[0043] From the above description, it can be seen that by separating and normalizing the header information and payload information, the length differences between different data packets can be eliminated, ensuring the consistency of the input data, thereby avoiding model bias or training instability caused by data packet length differences, and helping to improve the generalization ability and detection accuracy of the model; and constructing the header byte matrix and payload byte matrix according to the header byte sequence and payload byte sequence respectively, so that the subsequent model can learn the representation and function of different byte patterns.
[0044] Furthermore, the step of inputting the header byte information and the payload byte information into a byte-level traffic detection model to obtain the header byte representation feature and the payload byte representation feature comprises:
[0045] The byte-level traffic detection model includes two convolutional layers with different convolution kernels, a connection layer, a fully connected layer, and an activation layer;
[0046] The target byte information is processed respectively by two convolution layers with different convolution kernels to obtain a first vector and a second vector respectively; the target byte information includes header byte information or payload byte information;
[0047] Merging the first vector and the second vector through the connection layer to obtain a merged vector;
[0048] The merged vector is processed in sequence through the fully connected layer and the activation layer to obtain the target byte information representation feature corresponding to the target byte information.
[0049] From the above description, it can be seen that after the header byte information and payload byte information of the data packet are converted into corresponding matrices, the byte-level traffic detection model extracts the header byte representation features and payload byte representation features respectively, thereby providing more accurate basic information for subsequent packet-level feature fusion and final traffic classification.
[0050] Furthermore, the header byte representation features and the payload byte representation features are fused to obtain packet-level representation features including:
[0051] Calculating a header query vector, a header key vector, and a header value vector for the header byte representation feature, and calculating a payload query vector, a payload key vector, and a payload value vector for the payload byte representation feature;
[0052] Calculating dot products of the head query vector with the head key vector and the payload key vector respectively to obtain a first weight score and a second weight score;
[0053] Integrate the first weight score and the second weight score to obtain a header weight score;
[0054] Calculate the dot product of the header weight score with the header value vector and the payload value vector respectively and sum them to obtain a weighted header feature vector;
[0055] Calculate the dot product of the payload query vector with the header key vector and the payload key vector respectively to obtain a third weight score and a fourth weight score;
[0056] Integrate the third weight score and the fourth weight score to obtain a payload weight score;
[0057] Calculate the dot product of the payload weight score with the header value vector and the payload value vector respectively and sum them to obtain a weighted payload feature vector;
[0058] Concatenate the weighted header feature vector and the weighted payload feature vector to obtain the packet-level representation feature.
[0059] As can be seen from the above description, by using the adaptive feature fusion module, by calculating the query vector, key vector, and value vector of each feature vector, and respectively obtaining the weighted header feature vector and the weighted payload feature vector based on a special calculation method, and then performing feature vector fusion, the fusion of different feature vectors is realized.
[0060] Further, the construction of the packet association graph with the packet as a node includes:
[0061] According to the transmission relationship between all the packets, set the consecutive and in the same direction packets in the same node level;
[0062] Establish edges between the first packet and the last packet in each node level and the last packet in the previous node level respectively to obtain the packet association graph.
[0063] As can be seen from the above description, by taking the packet as a node and establishing edges according to the transmission relationship of the packets, the dynamic interaction mode of the packets in the network flow can be comprehensively captured, providing detailed structured information for subsequent feature extraction and classification in the network flow-level traffic detection model.
[0064] Further, the input of the packet association graph into the network flow-level traffic detection model to obtain the malicious traffic detection result includes:
[0065] Input the packet association graph into the associated attention traffic classifier, and the associated attention traffic classifier includes three layers, and each layer outputs a graph feature vector respectively;
[0066] Concatenate the three graph feature vectors to obtain a network flow-level feature representation;
[0067] Input the network flow-level feature representation into a fully connected layer and then perform a Softmax operation to obtain the malicious traffic detection result.
[0068] As can be seen from the above description, by using a three-layer stacked correlation attention traffic classifier to extract information from the packet association graph, key graph feature vectors in the packet association graph can be effectively extracted. After concatenating the three graph feature vectors, through a fully connected layer and a Softmax operation, an accurate detection result can be obtained.
[0069] Furthermore, the output of the three graph feature vectors by inputting the packet association graph into the three-layer stacked correlation attention traffic classifier includes:
[0070] Each target layer calculates the attention coefficient of each node feature in the packet association graph through an attention mechanism;
[0071] Normalize all the attention coefficients to obtain a normalized coefficient;
[0072] Update each node feature according to the normalized coefficient to obtain an updated feature;
[0073] Perform average pooling on all the updated features to obtain a target graph feature vector.
[0074] As can be seen from the above description, since each vertex in the packet association graph represents a packet, different weights are assigned to different nodes in the packet association graph through the attention mechanism, so that in the information propagation process, the different degrees of importance of each node can be reflected, highlighting the information of key nodes, while weakening the influence of nodes with lower correlation and improving the final detection effect.
[0075] Furthermore, the calculation of the attention coefficient of each node feature in the packet association graph through the attention mechanism includes:
[0076] e nm = a(W F n , W F m );
[0077] In the formula, e nm is the attention coefficient, representing the importance of packet m to packet n; a represents the attention mechanism; W is a learnable weight matrix; F n represents the node feature of the target packet; F m represents the node feature of the adjacent packet of the target packet;
[0078] Updating each of the node features according to the normalization coefficient to obtain updated features includes:
[0079]
[0080] wherein, σ is an activation function; α nm is the normalization coefficient; is the updated feature.
[0081] As can be seen from the above description, by using the attention mechanism to assign different weights to different nodes in the neighborhood, the neighboring nodes can reflect different degrees of importance, thereby highlighting the information of key nodes and weakening the influence of nodes with lower correlation, improving the final detection effect.
[0082] Another embodiment of the present invention provides a multi-granularity-based malicious traffic recognition device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements each step in the above-mentioned multi-granularity-based malicious traffic recognition method.
[0083] The multi-granularity-based malicious traffic recognition method and device provided by the present invention can be applied to the detection of malicious traffic for network traffic information fusion, which will be described below through specific embodiments:
[0084] Embodiment 1
[0085] Please refer to Figure 1 and Figure 2 , a multi-granularity-based malicious traffic recognition method includes:
[0086] S1. Preprocess the network traffic data to be analyzed to obtain network flow data; the network flow data includes data packets, specifically:
[0087] S11. Delete the invalid data in the network traffic data to obtain network valid data; that is, delete invalid data such as missing values and outliers from the original network traffic data, so as to obtain the cleaned network valid data.
[0088] S12. Screen and extract the data packets in the network valid data to obtain network screening data;
[0089] S13. Sort the data packets in the network screening data according to the timestamps of the data packets to obtain network sorted data;
[0090] S14. Perform data cleaning on the network sorted data to obtain network cleaned data;
[0091] S15. Normalize the network cleaning data to obtain normalized data; that is, convert the data into a unified scale range through normalization operations, and finally obtain the normalized information.
[0092] S16. Generate the network flow data from the normalized data through a network traffic processing method.
[0093] S2. Extract the header byte information and payload byte information of each data packet in the network flow data. Specifically:
[0094] S21. After separating the header information and payload information in the data packet, perform the same normalization process to obtain a header byte sequence and a payload byte sequence of the same length; for example, fill in zeros to complete data packets shorter than the preset length, and truncate data packets longer than the preset length, so as to obtain a header information sequence and a payload byte sequence of the same length. Although data packets longer than the preset length are truncated, the truncated part usually contains redundant or information that does not directly affect traffic characteristics, so truncating does not affect the data.
[0095] S22. Embed the header byte sequence and the payload byte sequence in each data packet into a preset coding vector respectively to obtain a header byte matrix and a payload byte matrix. Among them, the embedding coding method is as follows:
[0096] During the embedding process, for each possible byte value (0 - 255), select a relatively small embedding dimension such as 32 or 64 to form the byte embedding coding vectors for the header and payload. At the same time, the embedding coding vectors will be automatically updated according to the backpropagation algorithm during the training process to gradually optimize the representation of the bytes. When embedding, arrange the header byte embedding coding vectors in the order from top to bottom according to the sequence in a header byte to construct the header byte embedding matrix. The following is a specific embedding example:
[0097] Suppose the header information of a certain data packet is: [205, 100, 45, 200, 15, 255, 76, 90]; this indicates that the data packet header contains 8 bytes, and the value of each byte is in the range of 0 to 255; for each possible byte value, a 32-dimensional embedding vector is assigned to it; for example, the embedding vector corresponding to the byte value 205 is [0.1, -0.2, 0.3,..., 0.5] (a total of 32 floating-point numbers); the embedding vector corresponding to the byte value 100 is [0.4, 0.1, -0.3,...,-0.2] (a total of 32 floating-point numbers); other byte values also have corresponding 32-dimensional embedding vectors; arrange the embedding vectors corresponding to each byte in the order of the bytes in the header, and construct a matrix row by row from top to bottom. The final header byte embedding matrix obtained is an 8-row and 32-column matrix, where each row corresponds to the embedding vector of a byte, that is, the constructed header byte matrix is as follows:
[0098] |0.1 -0.2 0.3...0.5| <- byte value 205
[0099] |0.4 0.1 -0.3...-0.2| <- byte value 100 ......
[0101] |...............| <- byte value 90
[0102] Similarly, embed and encode the payload byte sequence into encoding vectors in the same way, arrange them in the order of the bytes in a payload byte from top to bottom, and construct a payload byte matrix.
[0103] S3. Input the header byte information and the payload byte information into the byte-level traffic detection model respectively to obtain the header byte representation feature and the payload byte representation feature. Specifically:
[0104] Please refer to Figure 3 , the byte-level traffic detection model includes two convolutional layers with different convolutional kernels, a connection layer, a fully connected layer, and an activation layer;
[0105] S31. Process the target byte information through the two convolutional layers with different convolutional kernels respectively to obtain a first vector and a second vector; the target byte information includes the header byte information or the payload byte information; for example: the header byte matrix B obtained in step S2 h respectively go through two convolutional operations with different convolutional kernels and output the first vector and the second vector Among them, the convolutional layer consists of a convolutional operation (Conv), an activation function (ReLU), and a max pooling layer (MaxPool). The specific calculation is as follows:
[0106]
[0107] S32. Merge the first vector and the second vector through the connection layer to obtain a merged vector; for example, when merging the first vector and the second vector , that is:
[0108]
[0109] wherein, represents the merged vector of the header byte matrix; CONCAT represents the connection layer;
[0110] S33. Process the merged vector sequentially through the fully connected layer and the activation layer to obtain the target byte information representation feature corresponding to the target byte information; for example, process the merged vector of the header byte matrix:
[0111]
[0112] wherein, F h represents the header byte information representation feature; FC represents the fully connected operation.
[0113] Similarly, by performing the same operations on the payload byte matrix B p , the corresponding payload byte information representation feature can be obtained, denoted as F p .
[0114] S4. Fuse the header byte representation feature and the payload byte representation feature to obtain a packet-level representation feature. Specifically:
[0115] Please refer to Figure 4 for the schematic diagram of the processing process of the weighted header feature vector;
[0116] S41. Calculate the header query vector, header key vector, and header value vector of the header byte representation feature, and calculate the payload query vector, payload key vector, and payload value vector of the payload byte representation feature; for example, the query vector is denoted as Q i , the key vector is denoted as K i , and the value vector is denoted as V i , then:
[0117]
[0118] In the formula, W (Q / K / V) is the learnable weight matrix corresponding to the query vector, key vector, and value vector; i = [h, p], representing the header or payload.
[0119] S42. Calculate the dot products of the head query vector with the head key vector and the payload key vector respectively to obtain a first weight score and a second weight score; that is, for the head query vector Q h calculate the dot product with the head key vector K h and the payload key vector K p respectively to generate a first weight score and a second weight score
[0120] S43. Integrate the first weight score and the second weight score to obtain a head weight score; for example, integrate the first weight score and the second weight score through the Softmax function to obtain a normalized weight score
[0121] S44. Calculate the dot products of the head weight score with the head value vector and the payload value vector respectively and sum them up to obtain a weighted head feature vector That is
[0122]
[0123] S44. Calculate the dot products of the payload query vector with the head key vector and the payload key vector respectively to obtain a third weight score and a fourth weight score; that is, for the payload query vector Q p calculate the dot product with the head key vector K h and the payload key vector K p respectively to generate a first weight score and a second weight score
[0124] S45. Integrate the third weight score and the fourth weight score to obtain a payload weight score
[0125] S46. Calculate the dot products of the payload weight score with the head value vector and the payload value vector respectively and sum them up to obtain a weighted payload feature vector That is
[0126]
[0127] S47. Concatenate the weighted head feature vector and the weighted payload feature vector to obtain the packet-level representation feature. That is, concatenate the weighted head feature vector with the weighted payload feature vector to obtain the final fused feature vector F k , where k represents the kth packet
[0128]
[0129] S5. Construct a packet association graph with the said data packets as nodes, and use the packet-level representation features as the node features of each node. Specifically:
[0130] S51. According to the transmission relationships among all the said data packets, set the consecutive and co-directional data packets in the same node level; wherein, the transmission relationships among data packets include the consecutive transmission relationship in the same direction and the interactive transmission relationship in the opposite direction; the consecutive transmission relationship in the same direction means that the sender continuously sends multiple data packets, that is, multiple data packets sent by the same party; such as Figure 5 the data packets -40, -162, -40 in are consecutive and co-directional data packets. The interactive transmission relationship in the opposite direction means that the receiver responds to the received data packet, that is, it reflects the relationship between the data packet in the reverse transmission and the original data packet.
[0131] S52. Establish edges respectively between the first data packet and the last data packet in each said node level and the last data packet in the previous node level to obtain the packet association graph; as Figure 5 shown, the edges formed by the data packets continuously sent by the sender in chronological order, and the first data packet in each layer is connected to the last data packet in the previous layer, thus connecting to form a packet association graph; this construction method comprehensively captures the dynamic interaction mode of data packets in the network flow, and provides detailed structured information for subsequent feature extraction and classification in the network flow-level traffic detection model.
[0132] S6. Input the packet association graph into the network flow-level traffic detection model to obtain the malicious traffic detection result. Specifically:
[0133] Please refer to Figure 6 , the network flow-level traffic detection model is an associated attention traffic classifier and includes a three-layer structure.
[0134] S61. After inputting the packet association graph into the associated attention traffic classifier, each layer outputs a graph feature vector respectively; since each vertex in the packet association graph represents a data packet and has different importance in the entire graph structure, the associated attention traffic classifier uses the attention mechanism to assign different weights to different nodes in the neighborhood. The specific assignment method is as follows:
[0135] S611. Each target layer calculates the attention coefficient of each said node feature in the packet association graph through the attention mechanism; for example, represent the node features of the packet association graph as F k ={F1,F2,…,F N}, where N is the number of nodes, and the attention coefficient is calculated using the attention mechanism a:
[0136] e nm =a(WF n ,WF m );
[0137] Where, e nm is the attention coefficient, which indicates the importance of data packet m to data packet n; a represents the attention mechanism; W is the learnable weight matrix; F n Indicates the node characteristics of the target data packet; F m Represents the node features of the neighboring data packets of the target data packet;
[0138] S612. Normalize all the attention coefficients to obtain normalized coefficients. For example, a Softmax function is used to perform normalization to ensure that the attention coefficients are easy to compare between different data packets, that is:
[0139] α nm =Softmax(e nm );
[0140] Among them, α nm is the normalized coefficient;
[0141] S613. Update each node feature according to the normalization coefficient to obtain an updated feature:
[0142]
[0143] Where σ is the activation function; To update features, l is the level, including the 1st layer, the 2nd layer and the 3rd layer;
[0144] S614: Perform average pooling on all updated features to obtain the target graph feature vector g (1) ,Right now:
[0145]
[0146] After obtaining the graph feature vector, activation function (ReLU) and batch normalization (BatchNorm) are applied after each layer.
[0147] S62, concatenate the three graph feature vectors to obtain a network flow level feature representation; that is, obtain the three graph feature vectors g according to step S61 respectively. (1) ,g (2) ,g (3) ; Use CONCAT to splice and get the network flow level feature F final It is expressed as:
[0148] F final = CONCAT(g (1) , g (2) , g (3) ).
[0149] S63. After inputting the network flow - level feature representation into the fully - connected layer, perform the Softmax operation to obtain the malicious traffic detection result. The specific formula is expressed as:
[0150]
[0151] In the formula, W and b are the parameters of the corresponding fully - connected layer, is the predicted probability value.
[0152] Embodiment 2
[0153] Please refer to Figure 7 , a malicious traffic recognition device based on multi - granularity, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it realizes each step in a malicious traffic recognition method based on multi - granularity as described in Embodiment 1.
[0154] In summary, the malicious traffic recognition method and device provided by the present invention pre - process the obtained network traffic data to obtain network flow data, construct a data - packet header byte sequence, a payload byte sequence, and a data - packet association graph based on the network flow data, input the data - packet header byte sequence and the payload byte sequence into the byte - level traffic detection model, output the representation features of the data - packet header byte sequence and the payload byte sequence, and fuse them to obtain the data - packet - level representation feature as the node feature of the data - packet association graph, and input it into the network flow - level traffic detection model. By constructing the byte sequence and the data - packet association graph, data processing is performed from multiple data - level granularities such as byte - level, data - packet - level, and network flow - level, realizing the fusion of multi - granularity network traffic information, making the fused node features more effective, enhancing the robustness, and making the traffic features input into the model more abundant, thereby realizing the efficient and accurate recognition of malicious traffic.
[0155] The above are only the embodiments of the present invention, and do not limit the patent scope of the present invention. Any equivalent transformation made using the content of the specification and drawings of the present invention, or directly or indirectly applied in the related technical fields, shall be included in the patent protection scope of the present invention by the same token.
Claims
1. A malicious traffic recognition method based on multiple granularities, characterized in that Including: Preprocessing the network traffic data to be analyzed to obtain network flow data; The network flow data includes data packets; Extracting the header byte information and payload byte information of each data packet in the network flow data; Inputting the header byte information and payload byte information into a byte-level traffic detection model respectively to obtain a header byte representation feature and a payload byte representation feature; Fusing the header byte representation feature and the payload byte representation feature to obtain a data packet-level representation feature; Constructing a data packet association graph with the data packets as nodes and using the data packet-level representation feature as the node feature of each node; Inputting the data packet association graph into a network flow-level traffic detection model to obtain a malicious traffic detection result.
2. The malicious traffic recognition method based on multiple granularities according to claim 1, wherein The preprocessing of the network traffic data to be analyzed to obtain network flow data includes: Deleting invalid data in the network traffic data to obtain network valid data; Screening and extracting data packets in the network valid data to obtain network screened data; Sorting the data packets in the network screened data according to the timestamps of the data packets to obtain network sorted data; Performing data cleaning on the network sorted data to obtain network cleaned data; Performing normalization processing on the network cleaned data to obtain normalized data; Generating the network flow data from the normalized data through a network traffic processing method.
3. A multi-granularity-based malicious traffic identification method according to claim 1, characterized in that, The extracting of the header byte information and payload byte information of each data packet in the network flow data includes: After separating the header information and payload information in the data packet, performing the same normalization processing to obtain a header byte sequence and a payload byte sequence of the same length; Embedding the header byte sequence and the payload byte sequence in each data packet into a preset encoding vector respectively to obtain a header byte matrix and a payload byte matrix.
4. A multi-granularity-based malicious traffic identification method according to claim 1, characterized in that The inputting of the header byte information and payload byte information into a byte-level traffic detection model respectively to obtain a header byte representation feature and a payload byte representation feature includes: The byte-level traffic detection model includes two convolutional layers with different convolutional kernels, a connection layer, a fully connected layer, and an activation layer; Processing the target byte information through the two convolutional layers with different convolutional kernels respectively to obtain a first vector and a second vector; the target byte information includes header byte information or payload byte information; Merging the first vector and the second vector through the connection layer to obtain a merged vector; Processing the merged vector sequentially through the fully connected layer and the activation layer to obtain a target byte information representation feature corresponding to the target byte information.
5. A multi-granularity based malicious traffic identification method according to claim 1, characterized in that The fusing of the header byte representation feature and the payload byte representation feature to obtain a data packet-level representation feature includes: Calculating a header query vector, a header key vector, and a header value vector of the header byte representation feature, and calculating a payload query vector, a payload key vector, and a payload value vector of the payload byte representation feature; Calculating the dot product of the header query vector with the header key vector and the payload key vector respectively to obtain a first weight score and a second weight score; Integrate the first weight score and the second weight score to obtain the head weight score; Calculate the dot product of the head weight score with the head value vector and the payload value vector respectively and sum them to obtain the weighted head feature vector; Calculate the dot product of the payload query vector with the head key vector and the payload key vector respectively to obtain the third weight score and the fourth weight score; Integrate the third weight score and the fourth weight score to obtain the payload weight score; Calculate the dot product of the payload weight score with the head value vector and the payload value vector respectively and sum them to obtain the weighted payload feature vector; Concatenate the weighted head feature vector and the weighted payload feature vector to obtain the packet-level representation feature.
6. The malicious traffic recognition method based on multiple granularities according to claim 1, wherein Said constructing a packet association graph with the packet as a node includes: According to the transmission relationship between all the packets, set the consecutive and co-directional packets in the same node level; Establish edges between the first packet and the last packet in each node level and the last packet in the previous node level respectively to obtain the packet association graph.
7. A multi-granularity-based malicious traffic identification method according to claim 1, characterized in that Said inputting the packet association graph into the network flow-level traffic detection model to obtain the malicious traffic detection result includes: Input the packet association graph into the associated attention traffic classifier, and the associated attention traffic classifier includes three layers, and each layer outputs a graph feature vector respectively; Concatenate the three graph feature vectors to obtain the network flow-level feature representation; Input the network flow-level feature representation into the fully connected layer and then perform the Softmax operation to obtain the malicious traffic detection result.
8. The malicious traffic recognition method based on multiple granularities according to claim 7, characterized in that, Said inputting the packet association graph into the three-layer stacked associated attention traffic classifier and outputting three graph feature vectors includes: Each target layer calculates the attention coefficient of each node feature in the packet association graph through the attention mechanism; Normalize all the attention coefficients to obtain the normalized coefficients; Update each node feature according to the normalized coefficient to obtain the updated feature; Perform average pooling on all the updated features to obtain the target graph feature vector.
9. The malicious traffic recognition method based on multiple granularities according to claim 8, wherein Said calculating the attention coefficient of each node feature in the packet association graph through the attention mechanism includes: e nm = a(WF n , WF m ); where e nm is the attention coefficient, representing the importance of data packet m to data packet n; a represents the attention mechanism; W is the learnable weight matrix; F n represents the node features of the target data packet; F m represents the node features of the neighboring data packets of the target data packet; Said updating each node feature according to the normalized coefficient to obtain the updated feature includes: where σ is the activation function; α nm is the normalization coefficient; is the updated feature.
10. A malicious traffic recognition device based on multiple granularities, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements each step in a multi-granularity-based malicious traffic recognition method as described in any one of claims 1-9.