QUIC encrypted video traffic identification method based on attention mechanism CNN deep learning model
By correcting the data shard length of QUIC encrypted video traffic based on the attention mechanism CNN deep learning model, the problem of low recognition accuracy of QUIC encrypted video traffic is solved, and high-precision recognition and robustness enhancement are achieved.
Patent Information
- Application Number
- CN202510198050.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-21
AI Technical Summary
The prior art is difficult to accurately identify QUIC encrypted video traffic, mainly because the encryption mechanism of the QUIC protocol hides the traffic characteristics that traditional recognition methods rely on, and the dynamic stream multiplexing and packet size variability of QUIC lead to a decrease in recognition accuracy.
The CNN deep learning model based on attention mechanism is adopted, and the data shard length is corrected, combined with the structural characteristics of the QUIC protocol, and the linear regression formula is trained to restore the original characteristics of the encrypted video traffic, and the time attention mechanism and channel attention mechanism are used to enhance the robustness of the model.
It realizes high-precision identification of QUIC encrypted video traffic, improves the accuracy and reliability of traffic segmentation, and can adapt to complex environments such as network fluctuations and transmission distortions.
Smart Images

Figure CN120075526A_ABST
Abstract
Description
Technical Field
[0001] The present invention proposes a video data shard length correction processing technology, which applies a CNN deep learning model based on the attention mechanism to realize the recognition of video traffic encrypted based on the QUIC protocol, belonging to the technical fields of network traffic management and security monitoring. Background Art
[0002] In today's Internet environment, video traffic has become an important part of network data transmission. The QUIC protocol has been widely used in the field of video streaming due to its strong encryption, fast connection establishment, and low latency characteristics brought by integrating UDP and TLS1.3. However, this encryption mechanism poses great challenges to network security and traffic analysis work.
[0003] Traditional video traffic recognition technologies have many problems when facing QUIC encrypted traffic. On the one hand, the encryption of the QUIC protocol makes it difficult for traditional methods of traffic recognition based on plaintext protocol header information to obtain effective features. On the other hand, existing means of identifying encrypted video traffic, such as methods based on video fingerprint libraries, machine learning, and reversible encryption / watermarking, cannot accurately identify video traffic due to characteristics such as the dynamic stream multiplexing, variable packet size, and protocol-level obfuscation of the QUIC protocol.
[0004] (1) Video recognition method based on fingerprint library
[0005] The video recognition method based on fingerprint library realizes video content recognition by extracting the unique fingerprint features of the video (such as payload length, frame features, audio features, coding characteristics, etc.) and comparing them with the database, and is widely used in copyright protection, video search, and media monitoring. However, this method has limitations in many aspects. First, constructing and maintaining the fingerprint library requires a large amount of video data for training and storage, and the matching process requires high computing resources, increasing the complexity and cost of the recognition system. Second, this method is relatively sensitive to changes in the video stream format. If the video undergoes transmission optimization, the packet structure may change, thus affecting the consistency of the fingerprint and reducing the recognition accuracy.
[0006] Under the QUIC protocol, the limitations of this method are particularly obvious. First, the QUIC protocol uses end-to-end encryption to encrypt the transport layer information of the data packet (such as stream ID, packet size, etc.), making it impossible for traditional methods to directly obtain useful fingerprint features. Second, QUIC supports the multiplexing function, that is, a connection can carry multiple data streams at the same time, which disrupts the packet order and traffic pattern relied on by traditional traffic recognition methods. In addition, the dynamic traffic control and adaptive congestion control mechanisms of QUIC make the traffic characteristics of the same video under different network conditions may be different, resulting in a decrease in the stability and accuracy of fingerprint recognition.
[0007] (2) Machine learning-based methods
[0008] Machine learning-based methods utilize labeled datasets to train classifiers (such as support vector machines, random forests, etc.) to identify different types of video traffic. These methods typically rely on extractable features in traffic patterns, such as packet size distribution, traffic peak variations, protocol header information, etc. However, the end-to-end encryption mechanism of the QUIC protocol hides many traffic features relied on by traditional machine learning methods, significantly reducing the availability of training data. At the same time, QUIC adopts a flexible flow control mechanism, and the transmission mode of packets can be dynamically adjusted according to network conditions, resulting in unstable traffic features and affecting the generalization ability of the model. In addition, since QUIC is still in the rapid development stage, there are few existing datasets, and the difficulty of model training is large, leading to a reduction in the recognition accuracy of traditional machine learning-based methods in the QUIC environment.
[0009] (3) Reversible encryption / watermark-based methods
[0010] This method embeds special watermarks or hidden information in encrypted videos for identification during subsequent transmission or decoding. Reversible encryption or watermarking techniques usually rely on traffic retaining certain detectable features during transmission, such as embedding specific bit patterns, modifying the statistical characteristics of video frames, etc. However, QUIC adopts end-to-end encryption, and the content at the packet level is completely invisible during transmission, making it difficult for external observers to detect or extract the embedded information. In addition, QUIC is the underlying protocol of HTTP / 3, and HTTP / 3 further hides the application layer data, making watermark detection even more difficult. Since video traffic does not leak its data features in a fixed pattern during QUIC transmission, the applicability of watermark-based encrypted video recognition methods under QUIC is severely limited. Summary of the Invention
[0011] To solve the problem in the prior art that it is difficult to accurately identify QUIC encrypted video traffic, the present invention proposes a method for identifying QUIC encrypted video traffic based on an attention mechanism CNN deep learning model. This method corrects the data shard length and combines it with an attention mechanism CNN deep learning model to achieve high-precision identification of QUIC encrypted video traffic.
[0012] The present invention proposes a method for identifying QUIC encrypted video traffic based on an attention mechanism CNN deep learning model.
[0013] To achieve the object of the present invention, the specific technical steps of the solution are as follows: A method for identifying QUIC encrypted video traffic based on an attention mechanism CNN deep learning model, the method comprising the following steps:
[0014] Step (1) As a user, collect the QUIC encrypted video traffic data in a normal environment and its corresponding original unencrypted traffic data;
[0015] Step (2) Process the video traffic data collected in step (1) by identifying the client request packets to achieve the division of data transmission units;
[0016] Step (3) Use the data transmission unit length sequence of the obtained QUIC encrypted video traffic data and the data transmission unit length sequence of its corresponding original unencrypted traffic data to train a linear regression formula for correcting the length.
[0017] Step (4) As a third party, collect the QUIC encrypted video traffic data in different network environments;
[0018] Step (5) First, divide the data transmission units by identifying the client request packets to divide the data transmission units of the encrypted traffic data, and then process them through the linear regression formula for correcting the length obtained in step (3) to obtain the data transmission unit corrected length sequence;
[0019] Step (6) Feed the data transmission unit corrected length sequence into a CNN model based on the attention mechanism for training to obtain a model that can accurately identify QUIC encrypted video traffic;
[0020] Step (7) Process the QUIC encrypted video traffic data according to step (5), and use the recognition model obtained in step (6) to identify the QUIC encrypted video traffic in different network environments.
[0021] Further, in step (1), the specific process of collecting data is as follows:
[0022] (1.1) Build a collection environment for capturing QUIC encrypted video traffic data, which requires one or more PCs and several network cables, and play the required videos on each PC;
[0023] (1.2) Obtain the URLs of the videos to be recognized on the corresponding platforms and summarize them to generate a video URL list file;
[0024] (1.3) Deploy a script program on each PC to capture the transmitted encrypted video packets while automatically playing the videos;
[0025] (1.4) Record the protocol handshake key when capturing videos on each PC, decrypt the captured encrypted video packets with this key, and then extract the actual data in the encrypted traffic for subsequent analysis and processing.
[0026] Further, in the step (2), the method for dividing the QUIC video traffic data transmission unit is as follows:
[0027] (2.1) By analyzing the principle of the video stream protocol, since when QUIC is applied to HTTP / 3, to achieve forward compatibility, some features of the HTTP protocol are still retained, including the basic request-response unit. Therefore, obvious segmentation can be seen in the traffic, that is, after the client sends a request message, it receives a response from the server. Accordingly, the encrypted video stream can be segmented;
[0028] (2.2) Group the data packets exchanged between two adjacent client request messages into a data transmission unit for subsequent analysis.
[0029] Further, in the step (3), the specific process of training the linear regression formula for the correction length is as follows:
[0030] (3.1) Select protocol features related to the video stream as input data, including the data packet length and the number of transmitted data packets;
[0031] (3.2) Use the actual data of the collected encrypted video stream (including the data packet length and the number of data packets), and through linear regression methods such as the least squares method, train the parameters (U1, U2) in the formula. The formula for linear regression is as follows:
[0032] cor_len = cip_len - pk_num × U1 + U2
[0033] Let the actual video packet length of the i-th sample be cor_len i , the encrypted data packet length be cip_len i , the number of data packets be pk_num i . Calculate the predicted value according to the current parameters U1 and U2
[0034] Define the error function:
[0035]
[0036] where n is the number of samples.
[0037] (3.3) Use the gradient descent method to optimize the model parameters to minimize the difference between the length of the encrypted data stream and the length of the actual data stream, and obtain the final parameter values.
[0038] Initialize the parameter U1 (0) , U2 (0)和 The learning rate α = 0.01.
[0039] The gradient calculation formula is:
[0040]
[0041] Update the parameters according to the update rule of gradient descent:
[0042]
[0043]
[0044] Set a convergence threshold ε , calculate the change amount ΔJ of the error function after each iteration, ΔJ = J(U1 (k) , U2 (k) ) - J(U1 (k+1) , U2 (k+1) ), when ΔJ < ε or reach the maximum number of iterations, when reaching the maximum number of iterations.
[0045] Furthermore, in step (4), as a third party, collect the QUIC encrypted video stream traffic data under different network environments:
[0046] The same as steps (1.1)(1.2)(1.3).
[0047] Furthermore, in step (5), the specific method for obtaining the corrected length sequence of the QUIC encrypted video traffic data transmission unit is as follows:
[0048] (5.1) Adopt the method of step (2) to divide the data transmission units of the encrypted traffic data;
[0049] (5.2) Extract the packet length of each data transmission unit and the UDP payload length;
[0050] (5.3) Process through the linear regression formula of the corrected length obtained in step (4) to obtain the corrected length sequence of the data transmission unit.
[0051] Furthermore, in step (6), the specific method for constructing a CNN model based on the attention mechanism for identifying QUIC encrypted video traffic is as follows:
[0052] (6.1) Construct a CNN model based on the attention mechanism. Establish a convolutional neural network model with dual attention enhancement to form a CNN-STA hybrid architecture. Dynamically enhance the key channel feature expression of video segments through the channel attention module (SEBlock), and combine the temporal attention module (TemporalAttention) to capture the important segment dependence relationship in the temporal dimension, and construct a deep learning model integrating spatial-temporal dual attention mechanisms to form an enhanced feature extraction architecture suitable for video traffic recognition.
[0053] Design a hierarchical attention model architecture. The model consists of a feature convolution unit, a dual attention unit, and a classification unit:
[0054] The feature convolution unit adopts a three-level progressive structure. Each level contains a Conv1D layer, a BN layer, a channel attention module, a temporal attention module (only in the third level), a ReLU activation function, and a MaxPool downsampling. Spatiotemporal features are extracted step by step through a 3×3 convolution kernel;
[0055] The dual attention unit is composed of channel attention and temporal attention in parallel. The SEBlock uses global average pooling and a fully connected layer to achieve channel dimension recalibration. TemporalAttention assigns feature weight in the temporal dimension through the QKV attention mechanism;
[0056] The classification unit adopts a three-order fully connected layer structure, including a 256-128 neuron hierarchical dimensionality reduction, combined with a ReLU activation function and a Dropout regularization layer (dropout rate 0.5). Finally, the class probability distribution is output through a linear projection layer.
[0057] (6.2) Train the model. Train the CNN model based on the attention mechanism through supervised learning: optimize the network parameters through end-to-end supervised learning. The channel attention module learns the importance weights of feature channels, the temporal attention module establishes long-range dependencies across frames, and the dual attention mechanism collaboratively improves the model's learning of the feature of the data transmission unit correcting the length sequence of the input video traffic, realizing accurate classification of the QUIC encrypted video stream traffic.
[0058] Furthermore, in step (7), the recognition process of the QUIC encrypted video traffic in the network is specifically as follows:
[0059] (7.1) Capture the QUIC encrypted video traffic in the network, divide it into data transmission units by the method of step (5), and then process it using the linear regression formula for correcting the length to obtain a corrected length sequence.
[0060] (7.2) Input the processed corrected length sequence into the attention mechanism-based CNN model trained in step (6). According to the classification result output by the model, perform real-time recognition on the transmitted encrypted video traffic to determine whether it belongs to the video to be recognized.
[0061] Compared with the prior art, the technical solution of the present invention has the following beneficial technical effects:
[0062] (1) Traditional video traffic recognition technologies cannot effectively overcome the length distortion problems caused by encryption and stream multiplexing when dealing with QUIC encrypted video streams. Compared with existing methods, the present invention proposes a method based on data shard length correction processing, combines the structural characteristics of the QUIC protocol, considers the influence of the encrypted header, trains a linear regression formula, and then realizes the correction of the length of video data shards, and can accurately restore the original characteristics of encrypted video traffic, thereby improving the accuracy and reliability of traffic segmentation.
[0063] (2) Compared with traditional CNN models, the present invention introduces a temporal attention mechanism and a channel attention mechanism, enabling the deep learning model to be more adaptable and stable when dealing with complex environments such as network fluctuations, packet loss, and transmission distortion, and further enhancing the robustness of the model in identifying QUIC encrypted video streams. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 is a system framework diagram of a method for identifying QUIC encrypted video traffic based on data shard length correction processing;
[0065] Figure 2 is a line graph of the length sequence of QUIC encrypted traffic data transmission units of the same video under different environments;
[0066] Figure 3 is a network topology diagram of a CNN model based on the attention mechanism;
[0067] Figure 4 is the channel attention (SEBlock) and temporal attention (TemporalAttention) mechanism modules. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0068] The following will detail the technical solutions provided by the present invention in combination with specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and not to limit the scope of the present invention.
[0069] Embodiment: A method for identifying QUIC encrypted video traffic provided by the present invention has an overall system structure as Figure 1 shown, and includes the following steps:
[0070] Step (1) As a user, collect QUIC encrypted video traffic data and its corresponding original unencrypted traffic data in a normal environment;
[0071] Step (2) Process the video traffic data collected in step (1) by identifying client request packets to achieve the division of data transmission units.
[0072] Step (3) uses the obtained data transmission unit length sequence of the QUIC encrypted video traffic data and the corresponding data transmission unit length sequence of the original unencrypted traffic data to train a linear regression formula for correcting the length.
[0073] Step (4) As a third party, collect QUIC encrypted video traffic data under different network environments;
[0074] Step (5) First, divide the data transmission unit by identifying the client request packet, divide the data transmission unit of the encrypted traffic data, and then process it through the linear regression formula for correcting the length obtained in step (3) to obtain the data transmission unit corrected length sequence.
[0075] Step (6) Feed the data transmission unit corrected length sequence into a CNN model based on the attention mechanism for training to obtain a model that can accurately identify QUIC encrypted video traffic.
[0076] Step (7) After dividing the QUIC encrypted video traffic data according to step (5), perform correction processing and use the recognition model obtained in step (6) to identify the QUIC encrypted video traffic under different network environments.
[0077] In an embodiment of the present invention, in step (1), as a user, by collecting the QUIC encrypted video traffic data in a normal environment and its corresponding original unencrypted traffic data, a dataset of the YouTube video platform encrypted by QUIC can be finally obtained. The steps of collecting data are as follows:
[0078] (1.1) Play the required video on a Firefox browser of a PC in Australia to capture QUIC encrypted video traffic data;
[0079] (1.2) Select videos from the international popular video provider YouTube that adopt the QUIC encryption protocol as the experimental object, write a web crawler script according to the characteristics of the YouTube video platform, obtain the URLs of the platform videos and summarize them into a video URL list file;
[0080] (1.3) Deploy the Uibot software on the PC and write an automated script program to use Uibot Creator to complete the automated collection work of video packets. The main work of this automated script program is: read the URL file obtained in (1.2), control the browser of the PC to play the video according to the read URL, control the PC to open Wireshark to capture packets before playing the video and store the captured packets;
[0081] (1.4) Configure a browser for playing videos on the PC. After each TLS session ends, record the key for decrypting the session data in the corresponding log file. Wireshark can decrypt the captured QUIC session data stream by accessing the key log file and using the key inside.
[0082] In an example of the present invention, in step (2), the steps for dividing the QUIC encrypted video traffic data transmission unit are as follows:
[0083] (2.1) By analyzing the principle of the video stream protocol, since when QUIC is applied to HTTP / 3, to achieve forward compatibility, some characteristics of the HTTP protocol are still retained, including the basic request - response unit. Therefore, obvious segmentation can be seen in the traffic, that is, after the client sends a request message, it receives a response from the server. Accordingly, the encrypted video stream can be segmented.
[0084] (2.2) Group the data packets exchanged between two adjacent client request messages as a data transmission unit for subsequent analysis.
[0085] In an embodiment of the present invention, in step (3), the specific process of training the linear regression formula for the correction length is as follows:
[0086] First, select protocol features related to the video stream as input data. The main features include:
[0087] · Packet length (cip_len): Represents the total length of the data packet after encryption by the QUIC protocol.
[0088] · Number of transmitted data packets (pk_num): The number of data packets in each data transmission unit.
[0089] The goal of the linear regression model is to predict the actual video stream length (output value) based on the characteristics of the encrypted stream (input features). In the present invention, the formula for linear regression is as follows:
[0090] cor_len = cip_len - pk_num × U1 + U2
[0091] Where cor_len is the corrected actual video stream length; cip_len is the length of the encrypted data packet (Ciphertext Length); pk_num is the number of data packets in the data transmission unit (DTU); U1, U2 are the parameters of the regression model, and these parameters need to be determined through training.
[0092] Use the actual data of the collected encrypted video stream (including the packet length and the number of data packets) as input features, and the target value is the actual video data length. Through linear regression methods such as the least squares method, the model will minimize the difference between the length of the encrypted data stream and the length of the actual data stream by optimizing U1 and U2.
[0093] Finally, we get: U1 = -28.0301, U2 = 10.9564
[0094] In one embodiment of the present invention, in step (4), as a third party, the specific process of collecting the QUIC encrypted video stream traffic data in different network environments is as follows:
[0095] (4.1) Play the required video on the Firefox browsers of a PC in Australia and a PC in Japan, and capture the QUIC encrypted video traffic data;
[0096] (4.2) Select the videos of the internationally popular video provider YouTube that uses the QUIC encryption protocol as the experimental objects, write a web crawler script according to the characteristics of the YouTube video platform, obtain the URLs of the platform videos and summarize them into a URL file;
[0097] (4.3) Set different network environments, deploy the Uibot software on the PC and write an automated script program to use Uibot Creator to complete the automated collection of video packets. The main tasks of this automated script program are: read the URL file obtained in (1.2), control the browser of the PC to play the video according to the read URL, control the PC to open Wireshark to capture packets before playing the video and store the captured packets;
[0098] Table 1 Dataset captured in step (4)
[0099]
[0100]
[0101] In one embodiment of the present invention, in step (5), process the QUIC encrypted video traffic captured in step (4) to obtain its corresponding data transmission unit correction length sequence. The specific process is as follows:
[0102] (5.1) Divide the data transmission units of the encrypted traffic data by the same method as in step (2);
[0103] (5.2) Extract the packet length, the number of transmitted data packets, and the number of data packets for each data transmission unit;
[0104] (5.3) Processed by the linear regression formula of the calibrated length obtained in step (4) to obtain the calibrated length sequence of the data transmission unit (see Figure 2 ).
[0105] In one embodiment of the present invention, in step (6), the specific process of training the CNN model based on the attention mechanism is as follows:
[0106] (6.1) Construct a CNN model based on the attention mechanism (such as Figure 3 ). Introduce the channel attention mechanism and the temporal attention mechanism on the basis of the traditional CNN model. The channel attention (SEBlock) mechanism (such as Figure 4 ) learns the weights of each channel through global average pooling, enhances the response of important features, and suppresses noise channels (packet loss, duplicate data). The temporal attention (TemporalAttention) mechanism (such as Figure 4 ) dynamically assigns weights in the time dimension, focuses on key frames, and ignores irrelevant frames (mistransmitted packets).
[0107] (6.2) Train the model. Train the CNN model based on the attention mechanism through supervised learning, learn the feature of the calibrated length sequence of the data transmission unit of the input video traffic, and realize the accurate classification of the QUIC encrypted video stream traffic.
[0108] Table 2 Convolutional layer hyperparameters
[0109] Hyperparameter Details Number of input channels 1 Convolution kernel size 3 Padding 1 Stride 1 Pooling layer Max pooling, pooling kernel size is 2, stride is 2 (the feature map size is halved after each pooling)
[0110] Table 2 Fully connected layer hyperparameters
[0111] Hyperparameter Details Input size of the fully connected layer 128 * (input_size / / 8) = 128 * 5 = 640 Number of neurons in the hidden layer First fully connected layer: 256; Second fully connected layer: 128 Activation function ReLU Dropout ratio 0.5
[0112] Table 4 Channel attention module (SEBlock) hyperparameters
[0113] Hyperparameter Details Compression ratio 16 (in the SEBlock, the number of channels is compressed to 1 / 16 of the original) Activation function ReLU (used in the fully connected layer of the SEBlock) Output activation function Sigmoid (used to generate channel attention weights)
[0114] Table 5 Temporal attention module (TemporalAttention) hyperparameters
[0115] Hyperparameter Details Number of channels of Query in_channels / / 8 (i.e., 128 / / 8 = 16) Number of channels of Key in_channels / / 8 (i.e., 128 / / 8 = 16) Number of channels of Value in_channels / / 8 (i.e., 128 / / 8 = 16) Attention scaling factor (gamma) Learnable parameter, initial value is 0
[0116] Table 6 Training hyperparameters
[0117]
[0118]
[0119] In one embodiment of the present invention, in step (7), the recognition process of the QUIC encrypted video traffic in the network is specifically as follows:
[0120] (7.1) Capture the QUIC encrypted video traffic in the network, and process it in the same way as step (5) to obtain the corrected length sequence.
[0121] (7.2) Input the processed corrected length sequence into the attention mechanism-based CNN model trained in step (6). According to the classification results output by the model, perform real-time recognition on the video title and record the result accuracy.
[0122] Table 7 Recognition Results
[0123] Environment Data used Accuracy Closed environment Closed world 99.25% Open environment Closed world + Open world 98.29% Complex environment 1 Closed world + 0.01% packet loss rate 96% Complex environment 2 Closed world + 0.05% packet loss rate 96% Complex environment 3 Closed world + 0.1% packet loss rate 96% Complex environment 4 Closed world + Mixed packet loss rate 96%
[0124] It should be noted that the above embodiments are not used to limit the protection scope of the present invention. Equivalent transformations or substitutions made on the basis of the above technical solutions all fall within the protection scope of the claims of the present invention.
Claims
1. A QUIC encrypted video traffic identification method based on the attention mechanism CNN deep learning model, characterized in that: The method comprises the following steps: Step (1) As a user, collect QUIC encrypted video traffic data and its corresponding original unencrypted traffic data under normal environment; Step (2) processes the video traffic data collected in step (1) by identifying the client request packet to realize the division of the data transmission unit; Step (3) using the obtained data transmission unit length sequence of the QUIC encrypted video traffic data and the corresponding data transmission unit length sequence of the original unencrypted traffic data to train a linear regression formula for the correction length; Step (4) as a third party, collect QUIC encrypted video traffic data in different network environments; Step (5) first divides the data transmission unit by identifying the client request packet, divides the data transmission unit of the encrypted traffic data, and then processes it through the linear regression formula of the correction length obtained in step (3) to obtain a data transmission unit correction length sequence; Step (6) sending the data transmission unit correction length sequence into the CNN model based on the attention mechanism for training to obtain a model that accurately identifies QUIC encrypted video traffic; Step (7) processes the QUIC encrypted video traffic data according to step (5), and uses the recognition model obtained in step (6) to identify the QUIC encrypted video traffic in different network environments.
2. According to claim 1, a QUIC encrypted video traffic identification method based on the attention mechanism CNN deep learning model is characterized in that: The step (1) specifically comprises the following sub-steps: (1.1) To build a collection environment for capturing QUIC encrypted video traffic data, one or more PCs and several network cables are required to play the required video on each PC; (1.2) Obtain the URLs of the videos to be identified on the corresponding platform and aggregate them to generate a video URL list file; (1.3) Deploy a script program on each PC to capture the transmitted encrypted video packets while automating the video playback; (1.4) The protocol handshake key is recorded when capturing video on each PC, and the captured encrypted video message is decrypted using the key to extract the actual data in the encrypted traffic for subsequent analysis and processing.
3. According to claim 1, a QUIC encrypted video traffic identification method based on the attention mechanism CNN deep learning model is characterized in that: The step (2) specifically comprises the following sub-steps: (2.1) By analyzing the principle of the video streaming protocol, when QUIC is applied to HTTP / 3, in order to achieve forward compatibility, some features of the HTTP protocol are still retained, including the basic request-response unit. Therefore, obvious segmentation can be seen in the traffic, that is, after the client sends a request message, it receives a response from the server, and based on this, the encrypted video stream is segmented; (2.2) The data packets exchanged between the two adjacent client request messages are grouped into a data transmission unit for subsequent analysis.
4. According to claim 1, a QUIC encrypted video traffic identification method based on the attention mechanism CNN deep learning model is characterized in that: The step (3) specifically comprises the following sub-steps: (3.1) Selecting protocol features related to the video stream as input data, including packet length and number of transmitted packets; (3.2) Using the actual data of the encrypted video stream collected (including the length and number of data packets), the parameters (U1, U2) in the formula are trained through linear regression methods such as the least squares method. The linear regression formula is as follows: cor_len=cip_len-pk_num×U1+U2 Let the actual video packet length of the i-th sample be cor_len i , the length of the encrypted data packet is cip_len i , the number of packets is pk_num i , calculate the predicted value based on the current parameters U1 and U2 Define the error function: Where n is the number of samples; (3.3) Gradient descent method is used to minimize the difference between the length of the encrypted data stream and the length of the actual data stream to obtain the final parameter value, Initialization parameter U1 (0) , U2 (0)和 Learning rate α = 0.01, The gradient calculation formula is: According to the update rule of gradient descent, update the parameters: Set a convergence threshold ε , calculate the change of error function after each iteration ΔJ=J(U1 (k) ,U2 (k) )-J(U1 (k +1) ,U2 (k+1) ), when ΔJ< ε Or when the maximum number of iterations is reached, when the maximum number of iterations is reached.
5. According to claim 1, a QUIC encrypted video traffic identification method based on the attention mechanism CNN deep learning model is characterized in that: The step (5) specifically comprises the following sub-steps: (5.1) dividing the data transmission unit by identifying the client request packet and dividing the data transmission unit of the encrypted traffic data; (5.2) Extract the packet length and number of packets of each data transmission unit; (5.3) The data transmission unit correction length sequence is obtained by processing the linear regression formula of the correction length obtained in step (4).
6. According to claim 1, a QUIC encrypted video traffic identification method based on the attention mechanism CNN deep learning model is characterized in that: The step (6) specifically comprises the following sub-steps: (6.1) Construct a CNN model based on the attention mechanism, establish a dual-attention enhanced convolutional neural network model, and form a CNN-STA hybrid architecture. Dynamically enhance the key channel feature expression of video clips through the channel attention module (SEBlock), combine the temporal attention module (TemporalAttention) to capture the important clip dependencies in the temporal dimension, and build a deep learning model that integrates the spatial-temporal dual attention mechanism to form an enhanced feature extraction architecture suitable for video traffic recognition. Design a hierarchical attention model architecture, which consists of a feature convolution unit, a dual attention unit, and a classification unit: The feature convolution unit adopts a three-level progressive structure. Each level includes Conv1D layer, BN layer, channel attention module, time attention module (only in the third level), ReLU activation function and MaxPool downsampling. The spatiotemporal features are extracted step by step through 3×3 convolution kernels. The dual attention unit is composed of channel attention and temporal attention in parallel. SEBlock uses global average pooling and fully connected layers to achieve channel dimension recalibration, and TemporalAttention uses the QKV attention mechanism to achieve time dimension feature weight allocation; The classification unit adopts a three-order fully connected layer structure, including 256-128 neuron layer dimensionality reduction, with ReLU activation function and Dropout regularization layer (drop rate 0.5), and finally outputs the category probability distribution through the linear projection layer; (6.2) Training model: Train the CNN model based on the attention mechanism through supervised learning: optimize the network parameters through end-to-end supervised learning, the channel attention module learns the feature channel importance weights, and the temporal attention module establishes long-range dependencies across frames. The dual attention mechanism synergistically improves the model's learning of the data transmission unit correction length sequence of the input video traffic, thereby achieving accurate classification of QUIC encrypted video stream traffic.
7. According to claim 1, a QUIC encrypted video traffic identification method based on the attention mechanism CNN deep learning model is characterized in that: The step (7) specifically comprises the following sub-steps: (7.1) Capture the QUIC encrypted video traffic in the network, divide it into data transmission units, and process it using the linear regression formula of the correction length to obtain the correction length sequence; (7.2) The processed corrected length sequence is input into the CNN model based on the attention mechanism trained in step (6), and the transmitted encrypted video traffic is identified in real time according to the classification result output by the model to determine whether it belongs to the video to be identified.
Citation Information
Patent Citations
Encrypted video identification method based on HTTP / 3 transmission characteristics
CN116668766A
Encrypted video identification method oriented to HTTP / 2 traffic multiplexing characteristics
CN116744052A
Application layer characterization of encrypted transport protocol
WO2024238001A1