Encrypted traffic identification method based on multi-feature fusion

By adopting a multi-feature fusion method that combines SE-ResNet1D and autoencoder in parallel in encrypted traffic recognition, the problems of low recognition accuracy and high feature extraction cost in encrypted traffic classification are solved, and efficient and robust encrypted traffic recognition effect is achieved.

CN120030501APending Publication Date: 2025-05-23SICHUAN UNIVERSITY OF SCIENCE AND ENGINEERING

Patent Information

Application Number
CN202510492177.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The prior art faces the problems of low recognition accuracy, manual pattern extraction of experts in the field of feature extraction and high cost, and deep learning methods resulting in loss of partial feature information due to network structure characteristics.

Method used

A multi-feature fusion encrypted traffic recognition method (MFF) is proposed, using parallel combination of SE-ResNet1D and automatic encoder (AE), timing features are extracted through SE attention mechanism and ResNet1D, and global statistical features are extracted through AE, and fused into a unified feature vector for classification.

Benefits of technology

It improves the accuracy and efficiency of encrypted traffic recognition, enhances the robustness of the model, and can more accurately identify and analyze traffic data in complex scenarios, achieving classification accuracy of 99.61% and 99.99%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030501A_ABST
    Figure CN120030501A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data security transmission, and discloses an encrypted traffic identification method based on multi-feature fusion, which can improve the identification precision and improve the overall efficiency and robustness. The encrypted traffic identification method based on multi-feature fusion comprises the following steps: S1, data preprocessing; the data preprocessing comprises the steps of flow splitting, statistical information extraction, flow anonymization and flow size unification; s2, traffic time sequence characteristics are extracted through SE-ResNet1D; s3, traffic statistical features are extracted through AE; s4, traffic classification; fusing the time sequence feature and the statistical feature of the traffic to form a traffic comprehensive feature F; and sending the traffic comprehensive feature F into a full-connection neural network for classification. By adopting the encrypted traffic identification method based on multi-feature fusion, local details of the traffic data can be captured, and the overall distribution condition of the traffic data can be better understood, so that the overall performance and accuracy are improved; the traffic data can be identified and analyzed more accurately in various complex scenes.
Need to check novelty before this filing date? Find Prior Art

Claims

1. The encrypted traffic identification method based on multi-feature fusion is characterized by: The following steps are involved: S1, data preprocessing; The data preprocessing includes the steps of traffic splitting, extracting statistical information, traffic anonymization, and unifying traffic size; S2, extract the time series features of traffic through SE-ResNet1D; Combine the SE attention mechanism with ResNet to form SE-ResNet1D; The SE attention mechanism includes the steps of: compression and activation; First, in the compression step, a global average pooling operation is performed on each channel of the input feature map to compress the spatial dimension, and finally a channel description vector z of length C is obtained, where C represents the total number of channels of the input feature map, as shown in the following formula: ; Where H is the height of the feature map, W is the width of the feature map, c is the index variable, indicating the current c-th channel, and c∈{1,2,…,C}; is the number of channels; represents the eigenvalue of the i-th row, j-th column, and c-th channel in the input feature map; z is the input vector; refers to the compression function; Then, in the activation step, a multi-layer perceptron is introduced to map the compressed features, and then activated by the activation function to obtain the weight of each channel, as shown in the following formula: ; In the formula, is the weight of each channel of the output, represents the weight, is the input vector, The meaning of this letter is the Relu activation function. is the Sigmoid activation function, The weights of the first fully connected layer, The weights of the second fully connected layer; It refers to the activation function; Finally, the weighted feature map is obtained by multiplying these channel weights with the original feature map; The ResNet includes multiple residual blocks, and a SE attention mechanism module is embedded after each residual block; The response of each channel is modeled through global average pooling to generate a channel-level weight vector. The activation strength of each channel is adjusted by multiplying the weight vector with the output feature after convolution, guiding the network to focus on important channel features in the traffic. Combine SE attention mechanism with ResNet to form SE-ResNet1D; extract the time series features of traffic through SE-ResNet1D; S3, extracting traffic statistical features through AE; The AE refers to an automatic encoder; the automatic encoder includes an encoder and a decoder; the encoder and the decoder are both implemented using a fully connected neural network; The encoder passes the function , enter Convert to latent feature representation ; The following formula: ; The decoder passes the function , the potential features Convert back to an output close to the original input; the following formula: ; Enter With output The difference between is defined as Δ; specifically as follows: ; In the formula It means taking the average of all observations; L2 regularization is used to minimize the difference Δ between input and output; the specific formula is as follows: ; In the formula Indicates input, Indicates expected value, represents the regularization parameter, represents the model parameters; By inputting 26 statistical features into the autoencoder (AE) for feature extraction, higher-order statistical features are obtained, thereby extracting deeper and potential information; S4, traffic classification; The time series features and statistical features of the traffic are concatenated and fused into a unified feature vector to form the comprehensive feature F of the traffic; the comprehensive feature F of the traffic is sent to the fully connected neural network for classification.

2. The method for identifying encrypted traffic based on multi-feature fusion as claimed in claim 1, characterized in that: In step S3, the mean absolute error MAE is used as the loss function of the autoencoder, as shown in the following formula: ; In the formula represents the loss function, Indicates The predicted value of samples, Indicates The true value of the samples, Represents the total number of samples.

3. The method for identifying encrypted traffic based on multi-feature fusion as claimed in claim 1, characterized in that: In step S1, traffic is split into sessions; The original traffic consists of a series of packets of different sizes. The collection of all packets is Indicates that each data packet is represented by , as shown in the following formula: ; In the formula, Refers to the original flow The total number of packets in Represents five-tuple information, namely source IP, source port, destination IP, destination port, transport protocol, Indicates the length of the packet in bytes. Indicates the time when the packet is sent; The session is used as a way to divide traffic; the session is a set of bidirectional flows, that is, the source and destination addresses in the five-tuple can be interchanged; flow To have the same five-tuple information The set of all data packets is as shown in the following formula: 。 4. The method for identifying encrypted traffic based on multi-feature fusion as claimed in claim 1, characterized in that: In step S1, statistical information is extracted, and a total of 26 statistical features are extracted to represent various characteristics of traffic; Including flag bit features; Packet length characteristics, time characteristics, load characteristics, and protocol characteristics; The flag bit characteristics include the average number of SYN, URG, FIN, ACK, PSH and RST flags, The protocol characteristics include the average number of DNS, TCP, UDP and ICMP packets; The temporal characteristics include the duration of the window flow and the average, minimum, maximum and standard deviation time intervals between each pair of data packets; The data packet length characteristics include the average, minimum, maximum and standard deviation lengths of the data packets; The load characteristics include the average number of small load data packets, the average value, minimum value, maximum value and standard deviation of the load size; Finally, the average number of DNS packets and the total number of packets transmitted over TCP are used as supplementary features; After extracting 26 statistical features, the Z-score standardization method is used to standardize each dimension of statistical features. After normalization, the standardized eigenvalue is recorded as ; As shown in the following formula: ; In the formula, Indicates The mean of the dimensional statistical features, Indicates The standard deviation of the dimensional statistical features, Represents a minimum constant.

5. The method for identifying encrypted traffic based on multi-feature fusion as claimed in claim 1, characterized in that: In step S1, traffic anonymization includes replacing the IP address and MAC address in each session by the placeholder 0x00; and deleting empty files and duplicate files.

6. The method for identifying encrypted traffic based on multi-feature fusion as claimed in claim 1, characterized in that: In step S1, the traffic size is unified; for traffic whose original size exceeds 1024 bytes, only the first 1024 bytes in the session are intercepted; and for traffic whose original size is less than 1024 bytes, 0x00 is added to the end until the length is supplemented to 1024 bytes.

Citation Information

Patent Citations

  • Transform-based comprehensive feature network traffic classification method

    CN118740428A

  • Visual time sequence feature network implementation method based on multi-module feature fusion

    CN119206847A

Cited By

  • Anonymous network flow association method based on feature extraction and feature enhancement

    CN119906552A

  • Encrypted malicious traffic identification method based on three-channel behavior image

    CN121907621A