A Fast Recognition Method for Resolution-Adaptive Encrypted Videos Based on Video Fingerprints

By acquiring plaintext fingerprints and correcting fingerprints at multiple resolutions of the video, combined with sliding windows and the Viterbi algorithm of the Hidden Markov model, the problem of difficult to identify encrypted videos under different network environments and resolutions in the prior art is solved, and a fast and accurate recognition effect is achieved.

CN116471460BActive Publication Date: 2025-06-13SOUTHEAST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310514450.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-08
Publication Date
2025-06-13
Estimated Expiration
2043-05-08

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify encrypted videos under different network environments and resolutions, and traditional methods require a large amount of sample data and time, making it difficult to adapt to scenes with resolution adaptability.

Method used

By obtaining video segment length sequences at multiple resolutions as plaintext fingerprints, and collecting video transmission data to obtain corrected fingerprints, the sliding window and hidden Markov model are used to identify them in combination with Viterbi algorithm.

Benefits of technology

It realizes the rapid and accurate identification of encrypted videos under different network environments and resolutions, reduces the time and cost of sample data acquisition, and adapts to resolution adaptive scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116471460B_ABST
    Figure CN116471460B_ABST
Patent Text Reader

Abstract

The present invention discloses a fast recognition method for resolution adaptive encrypted videos based on video fingerprints. First, the method obtains the video segment length sequences at multiple resolutions as plaintext fingerprints, and collects video transmission data. After correction processing, the corrected fingerprints of the encrypted videos are obtained. Based on the sliding window idea, the corrected fingerprint slices of the encrypted video to be recognized are compared with the plaintext fingerprint library to obtain possible matching results for each corrected fingerprint slice. Then, a hidden Markov model is constructed according to the corrected fingerprints and the plaintext fingerprint library corresponding to the matching results, and the Viterbi algorithm is used to solve the model to calculate the video plaintext fingerprint sequence with the highest probability, thereby realizing the recognition of encrypted videos. The present invention has universality and can quickly and accurately recognize encrypted videos, effectively solving the problem of encrypted video recognition caused by video resolution switching.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of network security, and in particular, relates to a fast recognition method for resolution adaptive encrypted videos based on video fingerprints. Background Art

[0002] With the continuous development of the Internet and the rapid popularization of various mobile devices, people's demand for video services has been continuously increasing. Therefore, various video platforms have emerged as the times require and become one of the important channels for people to pursue various needs such as entertainment, learning, and work. Moreover, the traffic proportion occupied by video services is also getting larger and larger and has become a major component of the total Internet traffic. According to the latest report released by the Ericsson Mobility website, about 70% of the global mobile network traffic in 2022 came from videos. With the continuous development of the Internet, it is expected that this proportion will reach an astonishing 80% by the end of 2028.

[0003] While videos bring great convenience to users, the number of harmful public nuisance videos mixed in them is also increasing continuously. These videos contain various harmful information, such as violence, hatred, discrimination, false propaganda, etc., which will have an adverse impact on the thoughts and mental health of viewers and even trigger violent behaviors. Moreover, these videos spread extremely fast in the network and society, causing serious negative impacts on the whole society. However, the number of videos in the network is extremely large, and the video sources are diverse and complex. Most videos are transmitted in an encrypted manner, making it difficult to easily know the content therein. For various reasons, it is difficult for regulatory agencies to identify such videos. Therefore, there is an urgent need for more effective technical means to help identify public nuisance videos in the network, reduce their influence range, and maintain a good environment for the Internet and society.

[0004] Among the methods for encrypted video recognition in the currently published literature, they are mainly divided into machine learning model prediction methods and video fingerprint recognition methods. Most of these methods identify videos under the condition of a fixed video resolution, without considering the impact of resolution adaptability on encrypted video recognition. There are mainly three problems with the existing recognition methods: (1) The existing methods extract data features from the data transmitted by encrypted videos, and then construct machine learning models or video fingerprints to identify encrypted videos, which overly rely on the network environment for video traffic collection. When the network environment changes, the features of the data transmitted by the same encrypted video will also change accordingly, making it difficult to effectively adapt to different network environments; (2) Traditional machine learning methods need to collect a large amount of video traffic sample data, extract features from it, and label the training data. Then, use the training data to train the video recognition model, and then use the trained model to identify encrypted videos. However, this process of collecting a large amount of sample data requires a lot of time and cost; (3) In order to ensure the smoothness and clarity of videos and avoid video stuttering and buffering problems caused by insufficient network bandwidth, current mainstream video platforms have all used HTTP adaptive streaming technology (HTTP adaptive streaming, HAS). This technology can automatically adjust the video resolution and bitrate according to the user's real-time network speed, ensuring that users can watch videos smoothly under various network conditions. However, the existing methods can only identify encrypted videos with a fixed resolution and are difficult to accurately and effectively identify resolution-adaptive encrypted videos. Summary of the Invention

[0005] Object of the Invention: Aiming at the above problems, the present invention discloses a fast recognition method for resolution-adaptive encrypted videos based on video fingerprints. This method first obtains the video segment length sequences at multiple resolutions as plaintext fingerprints, and collects video transmission data, and obtains the corrected fingerprints through correction and restoration; based on the sliding window idea, compares the corrected fingerprint shards of the encrypted video to be recognized with the plaintext fingerprint library to obtain possible matching results for each corrected fingerprint shard; then constructs a hidden Markov model according to the corrected fingerprints and the plaintext fingerprint library corresponding to the matching results, and then uses the Viterbi algorithm to solve the model to calculate the video plaintext fingerprint sequence with the highest probability, so as to realize the recognition of encrypted videos. The present invention has universality, can quickly and accurately identify encrypted videos, and effectively solves the problem of encrypted video recognition caused by video resolution switching.

[0006] Technical Solution: In order to achieve the object of the present invention, the technical solution adopted by the present invention is: A fast recognition method for resolution-adaptive encrypted videos based on video fingerprints, the method includes the following steps:

[0007] Step (1) Obtain the video segment length sequences of the encrypted video at multiple resolutions, and use them as plaintext fingerprints to construct a plaintext fingerprint library;

[0008] Step (2): Collect the transmission data during the playback of the encrypted video, eliminate the variations in the video segment length caused by the encapsulation of various protocols, and make it as close as possible to the original segment length to obtain the corrected fingerprint.

[0009] Step (3): In the full matching stage, using the idea of sliding matching, use the first few shards of the corrected fingerprint to calculate the videos that may match from the massive plaintext fingerprint database, and form a small plaintext fingerprint database with the corresponding plaintext fingerprints.

[0010] Step (4): In the fast matching stage, for the remaining shards in the corrected fingerprint, perform matching in the small plaintext fingerprint database obtained in step (3), and record the matching results for each corrected fingerprint shard.

[0011] Step (5): Use the information such as the plaintext fingerprint database corresponding to the matching results obtained in steps (3) and (4) and the corrected fingerprint of the encrypted video to be identified to construct a hidden Markov model.

[0012] Step (6): Use the Viterbi algorithm to solve the hidden Markov model constructed in step (5), and obtain the combination of plaintext fingerprints corresponding to the encrypted video data with the maximum probability as the identification result.

[0013] Furthermore, in the above step (1), obtain the video segment length sequences at multiple resolutions of the encrypted video, and use them as plaintext fingerprints to construct a plaintext fingerprint database. The method is as follows:

[0014] (1.1) Search for the selected keywords, and use a crawler program to obtain a list of video URLs.

[0015] (1.2) Construct the corresponding request headers according to the obtained URLs, access the interface of the video description file and download the video description file.

[0016] (1.3) Parse the video description file to obtain the sequences of video segment sizes at multiple video resolutions, that is, the plaintext fingerprints.

[0017] (1.4) Repeat steps (1.1)-(1.3) to obtain the plaintext fingerprints of multiple videos, and store the description information such as the video titles in the plaintext fingerprint database.

[0018] Furthermore, in the above step (2), collect the transmission data during the playback of the encrypted video, eliminate the variations in the video segment length caused by the encapsulation of various protocols, and make it as close as possible to the original segment length to obtain the corrected fingerprint. The method is as follows:

[0019] (2.1) Start the traffic collection tool Wireshark, open the browser and enter the video URL to play the video, and collect the encrypted transmission data during the video playback.

[0020] (2.2) Utilize the encrypted transmission data and the corresponding tags in the plaintext fingerprint database, and based on the principles of TLS and HTTP protocols for data encryption encapsulation and transmission, train the TLS correction model and the HTTP correction model to correct the interference caused by the TLS and HTTP protocols.

[0021] (2.3) According to the two correction models obtained in step (2.2), accurately restore the encrypted video transmission data and calculate the corrected fingerprint of the video.

[0022] Further, in step (3), in the full match stage, using the idea of sliding match, use the first few shards of the corrected fingerprint to calculate the potentially matching videos from the massive plaintext fingerprint database, and form a small plaintext fingerprint database with the corresponding plaintext fingerprints. The method is as follows:

[0023] (3.1) For the first few shards of the corrected fingerprint, set the corrected fingerprint window C. The corrected fingerprint window C slides on the video corrected fingerprint, and the window size is always 1.

[0024] (3.2) For each plaintext fingerprint in the plaintext fingerprint database, set the plaintext fingerprint window P. The plaintext fingerprint window P slides on the video plaintext fingerprint.

[0025] (3.3) Initially, set the window size of the plaintext fingerprint window P to 1 and the upper limit to n, where the value of n is related to the video content distribution mechanism of the video platform.

[0026] (3.4) If the difference in the sum of the fingerprint lengths in the plaintext fingerprint window P and the fingerprint length in the corrected fingerprint window C is within the set error range, it is considered that the plaintext fingerprint sequence in the plaintext fingerprint window P matches the corrected fingerprint in the corrected fingerprint window C successfully.

[0027] (3.5) Record the plaintext fingerprint sequence in the plaintext fingerprint window P at this time and its corresponding video ID.

[0028] (3.6) If the match fails, slide the plaintext fingerprint window P backward and then compare again. If the plaintext fingerprint window P slides to the end of the fingerprint, increase the window size by one.

[0029] (3.7) In the first few shards of the selected corrected fingerprint, repeat steps (3.4)-(3.6) until all shards are traversed.

[0030] (3.8) Form a small candidate plaintext fingerprint database with the plaintext fingerprints corresponding to all the obtained video IDs.

[0031] Further, in step (4), during the fast matching stage, for the remaining fragments in the corrected fingerprint, match them in the small plaintext fingerprint database obtained in step (3). The method for recording the matching results for each corrected fingerprint fragment is as follows: For the remaining fragments in the corrected fingerprint, sequentially perform the same matching process as in step (3), but only match in the small candidate plaintext fingerprint database obtained in step (3), thereby greatly reducing the scale of the matching and recording the successful matching results for each corrected fingerprint fragment.

[0032] Further, in step (5), use the plaintext fingerprint database corresponding to the matching results obtained in steps (3) and (4) and information such as the corrected fingerprint of the encrypted video to be recognized to construct a hidden Markov model. The method is as follows:

[0033] (5.1) Use the corrected fingerprint of the video to be recognized as the observation sequence, and each corrected fingerprint fragment as an output state;

[0034] (5.2) Use the set of plaintext fingerprint sequences corresponding to each corrected fingerprint fragment obtained in steps (3) and (4) as the hidden state sequence;

[0035] (5.3) Calculate the state transition matrix. The transition between plaintext fingerprint fragments when the video resolution switches is regarded as the transition between hidden states, and the transition probability between plaintext fingerprint fragments is regarded as the transition probability between hidden states. According to the sequential nature of video playback, specify the transition probability between different hidden states, and set the transition probability between plaintext fingerprint sequence fragments of different videos to a very small value. For the transition probability between plaintext fingerprint sequence fragments that conform to the transmission and download order of the same video, it can be calculated using the following formula:

[0036]

[0037] where S t is the hidden state at time t, n is the number of video segments included in S t , and N is the number of t-1 transferable states;

[0038] (5.4) Calculate the emission probability matrix. Since the main focus is on finding the optimal hidden state sequence and for the sake of simplifying the calculation, set the emission probability of different hidden states transitioning to the observation state at a certain moment to an equal probability;

[0039] (5.5) Calculate the initial state matrix. For the same reason as in (5.4), set it to a constant matrix;

[0040] (5.6) Construct a hidden Markov model based on the two state sets and three probability transition matrices obtained in (5.1)-(5.5).

[0041] Furthermore, in step (6), the Viterbi algorithm is used to solve the hidden Markov model constructed in step (5) to obtain the combination of plaintext fingerprints corresponding to the encrypted video data with the highest probability as the recognition result. The method is as follows:

[0042] (6.1) Multiply the probability of the initial state by the probability of observing the first data, and calculate the probability value δ corresponding to the most probable path in state i at t = 1 1 (i)

[0043]

[0044] where π i represents the probability of the initial state being i, represents the probability of generating the observed value o 1 when in state i;

[0045] (6.2) For t = 2, 3,..., T and each state j, calculate δ t (j)

[0046]

[0047] where a i,j represents the probability of transitioning from state i to state j, represents the probability of generating the observed value o t when in state j;

[0048] (6.3) After all moments are completed, the most likely state sequence needs to be selected, and the corresponding probability is the maximum δ t (j), that is, P * = max i (δ T (i));

[0049] (6.4) Starting from the last moment T, based on δ t (j) and the transition probability a i,j , gradually deduce the most likely state sequence:

[0050]

[0051] where represents the most likely hidden state at time T;

[0052] (6.5) According to the most likely state sequence I * , combine the corresponding plaintext fingerprints to obtain the combination of plaintext fingerprints corresponding to the encrypted video data with the highest probability as the recognition result.

[0053] Advantageous effects: Compared with the prior art, the technical solution of the present invention has the following advantageous technical effects.

[0054] (1) The present invention obtains the plaintext fingerprint of the video by parsing the video description file, collects the actual transmission data of the video, and uses the pre-trained correction and restoration model to obtain the corrected fingerprint. By comparing the corrected fingerprint with the plaintext fingerprint, encrypted videos are identified. In different network environments, the plaintext fingerprint of the video is fixed, while the corrected fingerprint can be obtained by the pre-trained correction and restoration model. Therefore, the method proposed by the present invention can stably identify encrypted videos in different network environments.

[0055] (2) Compared with traditional machine learning methods, the encrypted video recognition model proposed by the present invention does not need to repeatedly collect a large amount of encrypted transmission data of the video to be recognized for training the model. Only a part of the video traffic needs to be collected to train the correction and restoration model. Using the trained correction and restoration model and a plaintext fingerprint library containing the plaintext fingerprints of the encrypted videos to be recognized, the encrypted videos to be recognized can be accurately identified.

[0056] (3) In the recognition of encrypted videos, most existing studies are aimed at the recognition of encrypted videos at a fixed resolution, and it is difficult to apply this method to the scenario of resolution adaptation. In the method proposed by the present invention, videos with different resolutions correspond to different plaintext fingerprints, and the switching of resolutions corresponds to the transition of hidden states in the hidden Markov model. Then the Viterbi algorithm is used to solve the model to obtain the plaintext fingerprint sequence with the highest probability at different resolutions. The present invention can well adapt to the scenario of adaptive resolution playback and can stably and accurately identify encrypted videos using adaptive resolution playback.

[0057] (4) The present invention uses the method of sliding matching. In the full matching stage, the first few shards in the corrected fingerprint are first compared with a large plaintext fingerprint library to obtain a small candidate plaintext fingerprint library, thereby greatly reducing the number of plaintext fingerprints to be compared and improving the matching efficiency. Then the remaining shards in the corrected fingerprint are used to compare with the small candidate plaintext fingerprint library to obtain the final recognition result. Through this method, public nuisance videos can be quickly and accurately identified from a large amount of encrypted video traffic, providing a strong technical guarantee for preventing the rapid spread of public nuisance videos. Description of the Drawings

[0058] Figure 1 It is a schematic diagram of the system structure of a method for fast recognition of encrypted videos with resolution adaptation based on video fingerprints;

[0059] Figure 2 It is a flowchart of a method for fast recognition of encrypted videos with resolution adaptation based on video fingerprints;

[0060] Figure 3Schematic diagram of fingerprint sliding matching correction for encrypted videos;

[0061] Figure 4 Schematic diagram of encrypted video recognition based on the hidden Markov model. Specific implementation manners

[0062] The technical solutions of the present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0063] Specific embodiment: A method for fast recognition of resolution-adaptive encrypted videos based on video fingerprints provided by the present invention has an overall system architecture as Figure 1 shown, including the following steps:

[0064] (1) Obtain the video segment length sequences of the encrypted video at multiple resolutions, and use them as plaintext fingerprints to construct a plaintext fingerprint library;

[0065] (2) Collect the transmission data during the playback of the encrypted video, eliminate the changes in the video segment length caused by the encapsulation of various protocols, and make it as close as possible to the original segment length to obtain the corrected fingerprint;

[0066] (3) In the full matching stage, using the idea of sliding matching, use the first few shards of the corrected fingerprint to calculate the videos that may be matched from the massive plaintext fingerprint library, and form a small plaintext fingerprint library with the corresponding plaintext fingerprints;

[0067] (4) In the fast matching stage, for the remaining shards in the corrected fingerprint, perform matching in the small plaintext fingerprint library obtained in step (3), and record the results of each corrected fingerprint shard match;

[0068] (5) Use the information such as the plaintext fingerprint library corresponding to the matching results obtained in steps (3) and (4) and the corrected fingerprint of the encrypted video to be recognized to construct a hidden Markov model;

[0069] (6) Use the Viterbi algorithm to solve the hidden Markov model constructed in step (5), and obtain the combination of plaintext fingerprints corresponding to the encrypted video data with the maximum probability as the recognition result.

[0070] In an embodiment of the method of the present invention, in step (1), to obtain the video segment length sequences of the encrypted video at multiple resolutions and use them as plaintext fingerprints to construct a plaintext fingerprint library, the method is as follows:

[0071] (1.1) Search for the selected keywords, and use a crawler program to obtain a list of video URLs;

[0072] (1.2) Construct the corresponding request headers according to the obtained URLs, access the interface of the video description file, and download the video description file;

[0073] (1.3) Parse the video description file to obtain a sequence of video segment sizes at multiple video resolutions, i.e., the plaintext fingerprint.

[0074] (1.4) Repeat steps (1.1)-(1.3) to obtain the plaintext fingerprints of multiple videos, and store the description information such as the video title in the plaintext fingerprint database.

[0075] Different video platforms use different streaming media protocols. Mainstream video platforms usually use the HLS and DASH protocols to transmit video files. Among them, the video description file of the HLS protocol is an M3U8 file, and the video description file of the DASH protocol is an MPD file. In a few cases, it will be presented in JSON format. Different parsing methods are adopted for video description files in different formats to obtain the plaintext fingerprint. For example, the plaintext fingerprint of video ID 6rzlfG6Xkg0 on YouTube is shown in Table 1:

[0076] Table 1 Plaintext Fingerprint of Video 6rzlfG6Xkg0

[0077]

[0078] In an embodiment of the method of the present invention, in step (2), collect the transmission data during the playback of the encrypted video, eliminate the changes in the video segment length caused by the encapsulation of various protocols, and make it as close as possible to the original segment length to obtain the corrected fingerprint. The method is as follows:

[0079] (2.1) Start the traffic collection tool Wireshark, open the browser, enter the video URL to play the video, and collect the encrypted transmission data during the video playback.

[0080] (2.2) Use the encrypted transmission data and the corresponding tags in the plaintext fingerprint database, and according to the principle of data encryption encapsulation and transmission by the TLS and HTTP protocols, train the TLS correction model and the HTTP correction model to correct the interference caused by the TLS and HTTP protocols.

[0081] (2.3) According to the two correction models obtained in step (2.2), accurately restore the encrypted video transmission data and calculate the corrected fingerprint of the video.

[0082] When encapsulating video data, it will go through various protocol processes, such as segmentation, compression, and encryption, and add corresponding protocol headers and control information to ensure reliable data transmission. To obtain a corrected fingerprint that is relatively close to the original plaintext fingerprint, it is necessary to analyze these protocols and subtract the length of the added extra information from the actually transmitted payload obtained. For different HTTP protocol versions, it is necessary to analyze their protocol content to build different HTTP correction models. Currently, the widely used HTTP versions are HTTP / 1.1, HTTP / 2, and HTTP / 3. For example, compared with HTTP / 1.1, the HTTP / 2 protocol has a header compression mechanism, which will compress the HTTP header data and transmit it separately as a HEADERS frame. Therefore, for the HTTP / 2 protocol, this part of the length needs to be subtracted additionally.

[0083] In one embodiment of the method of the present invention, in step (3), in the full match stage, using the idea of sliding matching, the first few shards of the corrected fingerprint are used to calculate the videos that may match from the massive plaintext fingerprint library, and the corresponding plaintext fingerprints are formed into a small plaintext fingerprint library. The method is as follows:

[0084] (3.1) For the first few shards of the corrected fingerprint, set a corrected fingerprint window C. The corrected fingerprint window C slides on the video corrected fingerprint, and the window size is always 1.

[0085] (3.2) For each plaintext fingerprint in the plaintext fingerprint library, set a plaintext fingerprint window P. The plaintext fingerprint window P slides on the video plaintext fingerprint.

[0086] (3.3) Initially, set the window size of the plaintext fingerprint window P to 1 and the upper limit to n, where the value of n is related to the video content distribution mechanism of the video platform.

[0087] (3.4) If the difference between the sum of the fingerprint lengths in the plaintext fingerprint window P and the fingerprint lengths in the corrected fingerprint window C is within the set error range, it is considered that the plaintext fingerprint sequence in the plaintext fingerprint window P matches the corrected fingerprint in the corrected fingerprint window C successfully.

[0088] (3.5) Record the plaintext fingerprint sequence in the plaintext fingerprint window P at this time and its corresponding video ID.

[0089] (3.6) If the match fails, slide the plaintext fingerprint window P backward and then compare again. If the plaintext fingerprint window P slides to the end of the fingerprint, increase the window size by one.

[0090] (3.7) In the first few shards of the selected corrected fingerprint, repeat steps (3.4)-(3.6) until all shards are traversed.

[0091] (3.8) Combine the plaintext fingerprints corresponding to all the obtained video IDs to form a small candidate plaintext fingerprint library.

[0092] The corrected fingerprint shards obtained in step (2) and the plaintext fingerprint segments do not have a one-to-one correspondence. One corrected fingerprint shard may correspond to multiple plaintext fingerprint segments. And since the video may experience a resolution switching event during playback, at this time one corrected fingerprint shard may correspond to multiple plaintext fingerprint segments of different resolutions. Considering the above reasons, based on the sliding window idea, dynamically adjust the size of the window to achieve one-to-many matching. In the full matching stage, it is necessary to determine the maximum size n of the sliding window according to the content distribution mechanism of different video platforms. For example, after analyzing multiple Facebook videos collected, it is found that the size of n is related to the network quality. When the network quality is poor, n is usually 1 or 2, while when the network quality is good, n is usually 6 or 7. For example, the small candidate plaintext fingerprint library obtained by the corrected fingerprint with ID 6rzlfG6Xkg0 in the YouTube video during the full matching stage is shown in Table 2:

[0093] Table 2 Small candidate plaintext fingerprint library obtained from video 6rzlfG6Xkg0

[0094]

[0095] In an embodiment of the method of the present invention, in step (4), in the fast matching stage, for the remaining shards in the corrected fingerprint, match them in the small plaintext fingerprint library obtained in step (3), and record the results of each corrected fingerprint shard match. The method is as follows: For the remaining shards in the corrected fingerprint, sequentially perform the same matching process as in step (3), but only match in the small candidate plaintext fingerprint library obtained in step (3), thereby greatly reducing the scale of the matching, and record the successful matching results of each corrected fingerprint shard.

[0096] In an embodiment of the method of the present invention, in step (5), use the plaintext fingerprint library corresponding to the matching results obtained in steps (3) and (4) and information such as the corrected fingerprint of the encrypted video to be recognized to construct a hidden Markov model. The method is as follows:

[0097] (5.1) Use the corrected fingerprint of the video to be recognized as the observation sequence, and each corrected fingerprint shard as an output state;

[0098] (5.2) Use the set of plaintext fingerprint sequences corresponding to each corrected fingerprint shard obtained in steps (3) and (4) as the hidden state sequence;

[0099] (5.3) Calculate the state transition matrix. The switching between the plaintext fingerprint shards during video resolution switching is regarded as the transition between hidden states, and the switching probability between the plaintext fingerprint shards is regarded as the transition probability between hidden states. According to the sequentiality of video playback, the transition probabilities between different hidden states are specified. The transition probability between the plaintext fingerprint sequence shards of different videos is set to a minimum value. For the transition probability between the plaintext fingerprint sequence shards that conform to the transmission and download order of the same video, it can be calculated using the following formula:

[0100]

[0101] Among them, S t is the hidden state at time t, n is the number of video segments included in S t , and N is the number of states that S t-1 can be transferred to;

[0102] (5.4) Calculate the emission probability matrix. Since the main focus is on finding the optimal hidden state sequence and for the sake of simplifying the calculation, the emission probability of different hidden states at a certain moment being transferred to the observed state is set to an equal probability;

[0103] (5.5) Calculate the initial state matrix. For the same consideration as in (5.4), it is set to a constant matrix;

[0104] (5.6) Construct a hidden Markov model based on the two state sets and the three probability transition matrices obtained in (5.1)-(5.5).

[0105] In an embodiment of the method of the present invention, in step (6), the Viterbi algorithm is used to solve the hidden Markov model constructed in step (5), and the combination of plaintext fingerprints corresponding to the encrypted video data with the maximum probability is obtained as the recognition result. The method is as follows:

[0106] (6.1) Multiply the probability of the initial state by the probability of observing the first data, and calculate the probability value δ of the maximum probability path in state i at t = 1 1 (i)

[0107]

[0108] Among them, π i represents the probability of the initial state being i, represents the probability of generating the observed value o 1 when the state is i;

[0109] (6.2) For t = 2, 3,, T and each state j, calculate δ t (j)

[0110]

[0111] Among them, a i,j represents the probability of the state i transitioning to the state j, represents the probability of generating the observation value o when the state is j t ;

[0112] (6.3) After all moments end, the most likely state sequence needs to be selected, and the corresponding probability is the maximum δ t (j), that is, P * = max i (δ T (i));

[0113] (6.4) Starting from the last moment T, according to δ t (j) and the transition probability a i,j , gradually deduce the most likely state sequence:

[0114]

[0115] Among them, represents the most likely hidden state at moment T;

[0116] (6.5) According to the most likely state sequence I * , combine the corresponding plaintext fingerprints to obtain the combination of plaintext fingerprints corresponding to the encrypted video data with the maximum probability, as the recognition result.

[0117] In an embodiment of the present invention, when the video to be recognized is the YouTube video 6rzlfG6Xkg0, the final recognition result is shown in Table 3:

[0118] Table 3 Recognition result of video 6rzlfG6Xkg0

[0119]

[0120] The above embodiments are only the preferred embodiments of the present invention. It should be noted that: for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and equivalent replacements can be made. These technical solutions obtained by improving and equivalently replacing the claims of the present invention all fall within the protection scope of the present invention.

Claims

1. A fast recognition method for resolution - adaptive encrypted videos based on video fingerprints, characterized in that, the method comprises the following steps: Step (1): Obtain the video segment length sequences of the encrypted video at multiple resolutions, and use them as plaintext fingerprints to construct a plaintext fingerprint library; Step (2): Collect the transmission data during the playback of the encrypted video, eliminate the changes in the video segment length caused by the encapsulation of various protocols, and make it as close as possible to the original segment length to obtain a corrected fingerprint; Step (3): In the full - match stage, use a sliding window for matching. Use the first few shards of the corrected fingerprint to calculate the videos that may match from the massive plaintext fingerprint library, and form a small plaintext fingerprint library with the corresponding plaintext fingerprints; Step (4): In the fast - match stage, for the remaining shards in the corrected fingerprint, sequentially perform the same matching process as in Step (3), and perform the matching in the small plaintext fingerprint library obtained in Step (3), and record the results of each corrected fingerprint shard match; Step (5): Use the plaintext fingerprint library corresponding to the matching results obtained in Step (3) and Step (4) and the corrected fingerprint information of the encrypted video to be recognized to construct a hidden Markov model; Step (6): Use the Viterbi algorithm to solve the hidden Markov model constructed in Step (5), and obtain the combination of plaintext fingerprints corresponding to the encrypted video data with the maximum probability as the recognition result; The said Step (5) includes the following sub - steps: (5.1) Take the corrected fingerprint of the video to be recognized as the observation sequence, and each corrected fingerprint shard as an output state; (5.2) Take the set of plaintext fingerprint sequences corresponding to each corrected fingerprint shard obtained in Step (3) and Step (4) as the hidden state sequence; (5.3) Calculate the state transition matrix. The transition between plaintext fingerprint shards when the video resolution switches is regarded as the transition between hidden states, and the transition probability between plaintext fingerprint shards is regarded as the transition probability between hidden states. For the transition probability between plaintext fingerprint sequence shards that conform to the transmission and download order of the same video, use the following formula to calculate: Among them, S t is the hidden state at time t, n is the number of video segments included in S t and N is the number of transferable states in S t-1 ; (5.4) Calculate the emission probability matrix. Set the emission probability of different hidden states being transferred to the observation state at a certain moment as an equal probability; (5.5) Calculate the initial state matrix and set it as a constant matrix; (5.6) Construct a hidden Markov model according to the output state, hidden state sequence, state transition matrix, emission probability matrix, and initial state matrix obtained from (5.1)-(5.5).

2. The fast recognition method for resolution - adaptive encrypted videos based on video fingerprints according to Claim 1, characterized in that, the said Step (1) includes the following sub - steps: (1.1) Search for selected keywords, and use a crawler program to obtain a list of video URLs; (1.2) Construct the corresponding request headers according to the obtained URLs, access the interface of the video description file and download the video description file; (1.3) Parse the video description file to obtain the sequence of video segment sizes at multiple video resolutions, that is, the plaintext fingerprints; (1.4) Repeat steps (1.1)-(1.3) to obtain the plaintext fingerprints of multiple videos, and store the video title description information in the plaintext fingerprint database together.

3. A method for fast recognition of resolution adaptive encrypted videos based on video fingerprints according to claim 1, characterized in that, the step (2) includes the following sub-steps: (2.1) Start the traffic collection tool Wireshark, open the browser and enter the video URL to play the video, and collect the encrypted transmission data during video playback; (2.2) Use the encrypted transmission data and the corresponding tags in the plaintext fingerprint database, and according to the principle of data encryption packaging and transmission by the TLS and HTTP protocols, train the TLS correction model and the HTTP correction model to correct the interference caused by the TLS and HTTP protocols; (2.3) According to the two correction models obtained in step (2.2), accurately restore the encrypted video transmission data and calculate the corrected fingerprint of the video.

4. A method for fast recognition of resolution adaptive encrypted videos based on video fingerprints according to claim 1, characterized in that, the step (3) includes the following sub-steps: (3.1) For the first few shards of the corrected fingerprint, set a corrected fingerprint window C. The corrected fingerprint window C slides on the video corrected fingerprint, and the window size is always 1; (3.2) For each plaintext fingerprint in the plaintext fingerprint database, set a plaintext fingerprint window P. The plaintext fingerprint window P slides on the video plaintext fingerprint; (3.3) Initially, set the window size of the plaintext fingerprint window P to 1 and the upper limit to n, where the value of n is related to the video content distribution mechanism of the video platform; (3.4) If the difference in the sum of the fingerprint lengths in the plaintext fingerprint window P and the fingerprint length in the corrected fingerprint window C is within the set error range, it is considered that the plaintext fingerprint sequence in the plaintext fingerprint window P matches the corrected fingerprint in the corrected fingerprint window C successfully; (3.5) Record the plaintext fingerprint sequence in the plaintext fingerprint window P at this time and its corresponding video ID; (3.6) If the match fails, slide the plaintext fingerprint window P backward and then compare again. If the plaintext fingerprint window P slides to the end of the fingerprint, increase the window size by one; (3.7) In the first few shards of the selected corrected fingerprint, repeat steps (3.4)-(3.6) until all shards are traversed; (3.8) Combine the plaintext fingerprints corresponding to all the obtained video IDs into a small candidate plaintext fingerprint database.