Video transcoding classification detection method and system based on intra-frame and inter-frame vector information
By constructing video data sets and extracting intra-frame inter-vector information features, combining multi-forktree CTU division and inter-frame prediction direction analysis, the problem of poor detection performance in the existing technology under the new encoding standards is solved, and higher-precision video transcoding classification detection is achieved.
Patent Information
- Application Number
- CN202510017577.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-13
AI Technical Summary
The existing video transcoding detection methods perform poorly under the new encoding standards, especially under the HEVC encoding standards, which rely on PU block types, have a degradation in the VVC standard, and at the same time, the utilization of feature information in the B frame is insufficient.
By constructing the original video dataset and the transcoding video dataset, the vector information characteristics intra- and inter-frames are extracted, including the frequency information and motion vector information of the CU block, combined with the multi-forktree CTU division method and inter-frame prediction direction analysis, a richer feature set is generated, and the LIbSvm classifier of the RBF kernel function is used for classification detection.
This method can better capture the differentiated characteristics of the original video and the transcoding video, improve the accuracy of the video transcoding classification detection results, and perform more stably under the new encoding standards.
Smart Images

Figure CN119996704A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine vision image processing, and in particular to a video transcoding classification detection method and system based on intra-frame and inter-frame vector information. Background Art
[0002] With the development of the Internet and the popularity of smart devices, digital videos are closely connected with people's lives. Entertainment, learning, broadening horizons and sharing life through digital videos have become the norm. However, while the increasingly mature video editing technology brings convenience to life, it has a certain impact on the originality and integrity of digital videos. Video content can be easily tampered with through video editing software. Video tampering often uses operations such as scaling, cropping, frame deletion, frame insertion, etc. The process of tampering requires decoding the modified video, and then inevitably re-encoding operations. Tamperers may also re-encode the video using a coding standard different from the original video, tamper with the video or forge high-quality videos to obtain benefits, resulting in video transcoding. After the tamperer publishes the edited and modified video on the Internet, he spreads false information to guide bad speech and incite public opinion. In this regard, passive video forensics research is increasingly valued by scholars.
[0003] Related methods An effective exposure method for HEVC video transcoding based on coding unit (CU) and prediction unit (PU) partition types is proposed. It extracts the CU and PU partition types of I-type pictures and P-type pictures. Then, their average frequencies are calculated and concatenated as distinguishing features and further sent to support vector machine (SVM) for classification. Or a new method based on prediction unit (PU) statistics is proposed to detect HEVC transcoded videos in AVC format. According to the analysis of HEVC video footprint, 5D and 25D feature sets are extracted from I frames and P frames, respectively, and combined into the proposed 30-D feature set, which is finally input into the SVM classifier. Or it is of great significance to propose an algorithm to detect transcoded HEVC videos. First, a theoretical model of video transcoding is constructed, and a transcoding detection algorithm based on in-loop filtering and prediction unit (PU) partition (IFPP) is proposed. By analyzing the statistical characteristics of strong filtering and normal filtering modes in deblocking filtering and calculating the offset value in sample adaptive offset (SAO) filtering, the transcoding trajectory across coded frames can be captured. In addition, PU partition statistics are extracted to fully exploit the tracking in intra-coded frames. By fusing these sub-features, the proposed 17-dimensional IFPP feature is obtained and further fed into a support vector machine (SVM) classifier.
[0004] However, most of the related methods focus on AVC / HEVC transcoding detection, and the detection performance is poor under the new coding standard. First, the existing detection technology focuses on the CU block type division characteristics of intra-frame prediction. Although a similar coding structure is used under the new coding standard, the division method of the more detailed CTU block type is different from the past. In inter-frame prediction, in the HEVC coding standard, PU (Prediction Units) is used as an important feature to characterize transcoding traces, but under the new standard VVC, there is no PU type, so the traditional method that relies on PU block type shows a decrease in performance. Secondly, the existing detection methods often only consider the feature information of I frames and P frames in the video coding structure, and rarely use the feature information in B frames. Under the new coding standard, there are only I frames and B frames under the default coding structure (RA), which will affect the extraction of features, so the detection effect of the previous classification detection method will be reduced. Summary of the invention
[0005] In order to solve the above technical problems, the purpose of the present invention is to provide a video transcoding classification detection method and system based on intra-frame and inter-frame vector information, which can better capture the differentiated features of the original video and the transcoded video and improve the accuracy of the video transcoding classification detection results.
[0006] The first technical solution adopted by the present invention is: a video transcoding classification detection method based on intra-frame and inter-frame vector information, comprising the following steps:
[0007] Construct original video dataset and transcoded video dataset;
[0008] Performing intra-frame feature information extraction processing on the original video data set and the transcoded video data set to obtain intra-frame prediction feature information of the original video data set and intra-frame prediction feature information of the transcoded video data set;
[0009] Performing inter-frame feature information extraction processing on the original video data set and the transcoded video data set to obtain inter-frame prediction feature information of the original video data set and inter-frame prediction feature information of the transcoded video data set;
[0010] Based on the classifier, the intra-frame prediction feature information of the original video data set and the intra-frame prediction feature information of the transcoded video data set, as well as the inter-frame prediction feature information of the original video data set and the inter-frame prediction feature information of the transcoded video data set are classified and detected to obtain the video transcoding classification detection result.
[0011] Furthermore, the step of constructing the original video dataset and the transcoded video dataset specifically includes:
[0012] Get uncompressed video dataset;
[0013] The uncompressed video data set is compressed by a HEVC video encoder and a VVC video encoder to obtain a first compressed video data set and a second compressed video data set;
[0014] The first compressed video data set and the second compressed video data set are segmented and integrated to construct an original video data set;
[0015] The first compressed video data set is first decoded and then compressed by a VVC video encoder to obtain a transcoded video data set.
[0016] Further, the step of performing intra-frame feature information extraction processing on the original video dataset and the transcoded video dataset to obtain intra-frame prediction feature information of the original video dataset and intra-frame prediction feature information of the transcoded video dataset specifically includes:
[0017] By using a multi-tree CTU partitioning method, different CTU partitioning type blocks are extracted and marked for the original video data set and the transcoded video data set, so as to obtain several types of CU blocks of the original video data set and several types of CU blocks of the transcoded video data set;
[0018] Determine frequency information of several types of CU blocks of the original video data set and frequency information of several types of CU blocks of the transcoded video data set;
[0019] Acquire motion vector information of the original video data set and the transcoded video data set and map them to several types of CU blocks of the original video data set and several types of CU blocks of the transcoded video data set to obtain a mapping result;
[0020] Based on the mapping results, the number of intra-frame key frames in the video stream is set, and the frequencies of different CU block types in the key frames of the video stream are calculated in combination with the frequency information to obtain the intra-frame prediction feature information of the original video dataset and the intra-frame prediction feature information of the transcoded video dataset.
[0021] Furthermore, the calculation expression of the frequency information of the CU block is specifically as follows:
[0022]
[0023] In the above formula, Indicates the frequency information of different CU blocks, N CU Indicates the number of different CU block types, S CU Represents the total number of different CU block types in a single frame, and i represents several CU block division types.
[0024] Further, the calculation expression for calculating the frequency of different CU block types in the key frame of the video stream is specifically as follows:
[0025]
[0026] In the above formula, Indicates the frequency of different CU block types, n indicates the number of key frames in the video stream, Represents the frequency information of different CU blocks, and k represents several separate key frames.
[0027] Further, the step of performing inter-frame feature information extraction processing on the original video dataset and the transcoded video dataset to obtain inter-frame prediction feature information of the original video dataset and inter-frame prediction feature information of the transcoded video dataset specifically includes:
[0028] Setting a prediction direction of inter-frame prediction, the prediction direction includes unidirectional prediction and bidirectional prediction, and the unidirectional prediction includes forward prediction and backward prediction;
[0029] Obtain motion vector information of the original video data set and the transcoded video data set and map them to different CU block partition type frequencies on the CTU;
[0030] According to the frequency of different CU block partition types mapped to the CTU according to the motion vector information, the frequency of unidirectional prediction and the frequency of bidirectional prediction are determined in combination with the prediction direction of inter-frame prediction;
[0031] The frequency of unidirectional prediction and the frequency of bidirectional prediction are combined to obtain inter-frame prediction feature information of the original video data set and inter-frame prediction feature information of the transcoded video data set.
[0032] Further, the calculation expression of the frequency of different CU block partition types mapped to the motion vector information on the CTU is specifically as follows:
[0033]
[0034] In the above formula, Indicates the frequency of different CU block partition types mapped to the CTU. Represents the frequency information of different CU blocks, n represents the number of key frames in the video stream, and j represents several CU block division types.
[0035] Furthermore, the calculation expressions of the unidirectional prediction frequency and the bidirectional prediction frequency are specifically as follows:
[0036]
[0037] In the above formula, represents the frequency of one-way prediction, represents the frequency of bidirectional prediction, represents forward prediction, represents backward prediction, represents bidirectional prediction, and j represents several CU block partition types.
[0038] Further, the step of performing classification detection on the intra-frame prediction feature information of the original video data set and the intra-frame prediction feature information of the transcoded video data set, as well as the inter-frame prediction feature information of the original video data set and the inter-frame prediction feature information of the transcoded video data set based on the classifier to obtain the video transcoding classification detection result specifically includes:
[0039] The inter-frame prediction feature information of the original video data set is combined with the inter-frame prediction feature information of the transcoded video data set, and the intra-frame prediction feature information of the original video data set is combined with the intra-frame prediction feature information of the transcoded video data set to obtain the prediction features of the combined original video data set and the prediction features of the combined transcoded video data set;
[0040] The prediction features of the fused original video data set and the prediction features of the fused transcoded video data set are standardized to obtain the prediction features of the standardized original video data set and the prediction features of the standardized transcoded video data set;
[0041] The predicted features of the standardized original video dataset and the predicted features of the standardized transcoded video dataset are classified and detected by the LIbSvm classifier based on the RBF kernel function to obtain the video transcoding classification detection results.
[0042] The second technical solution adopted by the present invention is: a video transcoding classification detection system based on intra-frame and inter-frame vector information, comprising:
[0043] The first module is used to construct the original video dataset and the transcoded video dataset;
[0044] The second module is used to extract intra-frame feature information of the original video data set and the transcoded video data set to obtain intra-frame prediction feature information of the original video data set and intra-frame prediction feature information of the transcoded video data set;
[0045] The third module is used to extract inter-frame feature information of the original video data set and the transcoded video data set to obtain inter-frame prediction feature information of the original video data set and inter-frame prediction feature information of the transcoded video data set;
[0046] The fourth module is used to classify and detect the intra-frame prediction feature information of the original video data set and the intra-frame prediction feature information of the transcoded video data set, as well as the inter-frame prediction feature information of the original video data set and the inter-frame prediction feature information of the transcoded video data set based on the classifier to obtain the video transcoding classification detection result.
[0047] The beneficial effects of the method and system of the present invention are as follows: the present invention constructs an original video data set and a transcoded video data set, and then extracts and processes intra-frame feature information of the original video data set and the transcoded video data set to obtain intra-frame prediction feature information of the original video data set and intra-frame prediction feature information of the transcoded video data set, and extracts and processes inter-frame feature information of the original video data set and the transcoded video data set to obtain inter-frame prediction feature information of the original video data set and inter-frame prediction feature information of the transcoded video data set. Starting from the difference in intra-frame and inter-frame frequency characteristics between the transcoded video and the original video, the number of different division types of CU blocks in the I frame and the consistent representation of the prediction direction of the inter-frame motion vector in the B frame are analyzed, so that the differentiated features of the original video and the transcoded video can be better captured. Finally, based on the classifier, the intra-frame prediction feature information of the original video data set and the intra-frame prediction feature information of the transcoded video data set, as well as the inter-frame prediction feature information of the original video data set and the inter-frame prediction feature information of the transcoded video data set are classified and detected, so that the difference between the original video and the transcoded video can be better learned, and the accuracy of the video transcoding classification detection result can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 It is a flow chart of the steps of the video transcoding classification detection method based on intra-frame and inter-frame vector information of the present invention;
[0049] Figure 2 It is a structural block diagram of a video transcoding classification detection system based on intra-frame and inter-frame vector information of the present invention;
[0050] Figure 3 It is a schematic diagram of a video transcoding process provided by a specific embodiment of the present invention;
[0051] Figure 4 is a schematic diagram of CU division types provided by a specific embodiment of the present invention;
[0052] Figure 5 is a schematic diagram of motion vector mapping provided by a specific embodiment of the present invention;
[0053] Figure 6 It is a schematic diagram of a framework for video transcoding classification detection provided by a specific embodiment of the present invention. DETAILED DESCRIPTION
[0054] The present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. The step numbers in the following embodiments are only provided for the convenience of explanation and description, and the order between the steps is not limited in any way. The execution order of each step in the embodiment can be adaptively adjusted according to the understanding of those skilled in the art.
[0055] Reference Figure 1 and Figure 6The present invention provides a video transcoding classification detection method based on intra-frame and inter-frame vector information, the method comprising the following steps:
[0056] S100, constructing an original video dataset and a transcoded video dataset;
[0057] Specifically, an uncompressed video data set is obtained; the uncompressed video data set is compressed by a HEVC video encoder and a VVC video encoder respectively to obtain a first compressed video data set and a second compressed video data set; the first compressed video data set and the second compressed video data set are segmented and integrated to construct an original video data set; the first compressed video data set is first decoded and then compressed by a VVC video encoder to obtain a transcoded video data set.
[0058] In this embodiment, if Figure 3 As shown in the figure, 36 uncompressed videos are obtained from open source websites to construct a dataset. The uncompressed videos are recompressed using the H.265 (HEVC) / H.266 (VVC) video encoder, and the bit rates are set to {3, 4, 5, 6, 7} Mb / s respectively. Each video is further split into non-overlapping video segments containing only 100 frames, which are used as the original video dataset.
[0059] The compressed H.265 is re-decoded and encoded into H.266 video, which is used as the negative sample of the dataset (i.e., transcoded video) according to different parameter settings.
[0060] S200, performing intra-frame feature information extraction processing on the original video data set and the transcoded video data set to obtain intra-frame prediction feature information of the original video data set and intra-frame prediction feature information of the transcoded video data set;
[0061] Specifically, different CTU partition type blocks are extracted and marked for the original video dataset and the transcoded video dataset respectively through a multi-tree CTU partitioning method to obtain several types of CU blocks of the original video dataset and several types of CU blocks of the transcoded video dataset; the frequency information of the several types of CU blocks of the original video dataset and the frequency information of the several types of CU blocks of the transcoded video dataset are determined; the motion vector information of the original video dataset and the transcoded video dataset is obtained and mapped to the several types of CU blocks of the original video dataset and the several types of CU blocks of the transcoded video dataset to obtain a mapping result; based on the mapping result, the number of intra-frame key frames of the video stream is set, and the frequency of different CU block types in the key frames of the video stream is calculated in combination with the frequency information to obtain the intra-frame prediction feature information of the original video dataset and the intra-frame prediction feature information of the transcoded video dataset.
[0062] In this embodiment, the division methods of CTU blocks are analyzed and compared. The new coding standard adopts the division methods such as quadtree, ternary tree, binary tree, etc. Figure 4 As shown in the figure, VTM (the official reference codec for the H.266 coding standard) is used to extract H.266 decoding information, extract and mark different CTU division type blocks, divide them into 17 different types of CU blocks, and further count the frequency information of each different type of block CU.
[0063] Furthermore, the motion vector mapping proposed in the embodiment of the present invention is specifically as follows: Figure 5 As shown, the motion vector information is extracted from the VTM, the extracted feature information is mapped to the CU block, and the direction consistency of the motion vector is represented by counting the number of CU blocks.
[0064] Among them, the frequency of each CU block type of a single frame extracted is calculated as follows:
[0065]
[0066] Where N CU is the number of different CU block types, S CU is the sum of the number of different CU block types in a single frame, i is one of the 17 CU block division types, and the specific division blocks are shown in Table 1. Assuming the number of key frames in a video stream is n, the frequency of different CU block types in the key frames of the video stream is calculated as:
[0067]
[0068] Get the 17-D detection features of the intra-frame prediction frame
[0069] Table 1 Key frame CU block partition type data table
[0070] CUtype 4x4 4x8 4x16 4x32 8x8 Index 0 1 2 3 4 CUtype 8x4 8x16 8x32 16x16 16x4 Index 5 6 7 8 9 CUtype 16x8 16x32 32x32 32x4 32x8 Index 10 11 12 13 14 CUtype 32x16 64x64 \ \ \ Index 15 16 \ \ \
[0071] S300, performing inter-frame feature information extraction processing on the original video data set and the transcoded video data set to obtain inter-frame prediction feature information of the original video data set and inter-frame prediction feature information of the transcoded video data set;
[0072] Specifically, a prediction direction of inter-frame prediction is set, the prediction direction includes unidirectional prediction and bidirectional prediction, and the unidirectional prediction includes forward prediction and backward prediction; motion vector information of the original video data set and the transcoded video data set is obtained and mapped to different CU block partition type frequencies on the CTU; according to the motion vector information mapped to different CU block partition type frequencies on the CTU, combined with the prediction direction of inter-frame prediction, the frequency of unidirectional prediction and the frequency of bidirectional prediction are determined; the frequency of unidirectional prediction and the frequency of bidirectional prediction are combined to obtain the inter-frame prediction feature information of the original video data set and the inter-frame prediction feature information of the transcoded video data set.
[0073] In this embodiment, P is defined i (i=1, 2, 3) are the three prediction directions (forward prediction, backward prediction, and bidirectional prediction) in inter-frame prediction. The frequency of different CU block partition types mapped to CTU by a single direction motion vector is calculated. The expression is:
[0074]
[0075] Where j is one of 17 CU block partition types, and the specific partition blocks are shown in Table 2. In order to better utilize the feature information, the prediction direction is divided into unidirectional prediction and bidirectional prediction, and the frequencies of the two prediction methods are calculated respectively. The expressions are:
[0076]
[0077] Get the 34-D detection features of the inter-frame
[0078] Table 2 Inter-frame prediction frame CU block partition type data table
[0079] CUtype 8x8 16x16 32x32 64x64 128x128 Index 0 1 2 3 4 CUtype 8x16 16x8 16x32 32x16 64x32 Index 5 6 7 8 9 CUtype 32x64 64x128 128x64 8x32 32x8 Index 10 11 12 13 14 CUtype 16x64 64x16 \ \ \ Index 15 16 \ \ \
[0080] S400, based on the classifier, classify and detect the intra-frame prediction feature information of the original video data set and the intra-frame prediction feature information of the transcoded video data set, as well as the inter-frame prediction feature information of the original video data set and the inter-frame prediction feature information of the transcoded video data set to obtain a video transcoding classification detection result.
[0081] Specifically, the inter-frame prediction feature information of the original video dataset is fused with the inter-frame prediction feature information of the transcoded video dataset, as well as the intra-frame prediction feature information of the original video dataset and the intra-frame prediction feature information of the transcoded video dataset are fused to obtain the prediction features of the fused original video dataset and the prediction features of the fused transcoded video dataset; the prediction features of the fused original video dataset and the prediction features of the fused transcoded video dataset are standardized to obtain the standardized prediction features of the original video dataset and the prediction features of the standardized transcoded video dataset; the prediction features of the standardized original video dataset and the prediction features of the standardized transcoded video dataset are classified and detected by the LIbSvm classifier based on the RBF kernel function to obtain the video transcoding classification detection result.
[0082] In this embodiment, the detection features of the key frame and the non-key frame are fused to obtain the final 51-D detection feature, which is expressed as:
[0083] F inter ={F intra ,F inter}
[0084] The obtained features are standardized and the expression is:
[0085]
[0086] In the formula, x is a data in the original data, max(x) represents the maximum value in the original data, and min(x) represents the maximum value in the original data.
[0087] The standardized feature vector is put into the classifier for training. In the embodiment of the present invention, LIbSvm of RBF kernel function is used as the classifier, and the values of gamma, C and cache_size are set to 0.5, 1 and 100 respectively, and 5-fold cross validation is used to obtain the optimal training parameters, in which a single compressed VVC video clip is marked as a negative sample, and the transcoded clip is marked as a positive sample.
[0088] After extracting the features of the test set, put them into the trained model, and divide the sample data set in the test set into negative samples and positive samples. Negative samples represent normal videos, and positive samples represent tampered videos.
[0089] Among them, the detection evaluation index is:
[0090]
[0091] Where TP and TN are the number of true positive and true negative samples respectively. P is the number of positive samples, and N is the number of negative samples.
[0092] The average result of 20 detections of the test set is taken as the final result of transcoding detection, and the effectiveness and robustness of the algorithm are verified under different sample sets and parameter settings. The experimental results are shown in Tables 1 to 4.
[0093] Among them, Table 3 is a performance comparison of the proposed algorithm and sub-features as well as existing transcoding detection methods. It can be seen that the proposed algorithm can have higher detection accuracy after fusion, and the detection performance is also better than the previously proposed algorithms.
[0094] Table 3 Transcoding detection performance data table of the method of the embodiment of the present invention compared with the existing method (%)
[0095]
[0096]
[0097] Table 4 shows the performance of the proposed algorithm and sub-features on different resolution datasets. It can be seen that the proposed algorithm has better performance in different datasets after fusion, and it can better utilize features at higher resolutions.
[0098] Table 4 Transcoding detection accuracy data table under different data sets (%)
[0099]
[0100]
[0101] Table 5 shows the sub-features F of the proposed algorithm. inter Performance on the 1080p dataset under different GOP structures. It can be seen that the proposed sub-features show good robustness under different GOP structures.
[0102] Table 5 Detection results under different GOP structures
[0103]
[0104]
[0105] Table 6 shows the detection results of the proposed algorithm for different encoder preset parameter data sets. The three parameters Fast, Medium, and Slow represent low-quality, medium-quality, and high-quality video settings, respectively. This further verifies the detection performance of the proposed algorithm.
[0106] Table 6 Detection results under different encoding preset parameters (%)
[0107]
[0108]
[0109] Reference Figure 2 , a video transcoding classification detection system based on intra-frame and inter-frame vector information, comprising:
[0110] The first module 201 is used to construct an original video dataset and a transcoded video dataset;
[0111] The second module 202 is used to extract intra-frame feature information from the original video data set and the transcoded video data set to obtain intra-frame prediction feature information of the original video data set and intra-frame prediction feature information of the transcoded video data set;
[0112] The third module 203 is used to extract inter-frame feature information from the original video data set and the transcoded video data set to obtain inter-frame prediction feature information of the original video data set and inter-frame prediction feature information of the transcoded video data set;
[0113] The fourth module 204 is used to classify and detect the intra-frame prediction feature information of the original video data set and the intra-frame prediction feature information of the transcoded video data set, as well as the inter-frame prediction feature information of the original video data set and the inter-frame prediction feature information of the transcoded video data set based on the classifier to obtain the video transcoding classification detection result.
[0114] The contents of the above method embodiments are all applicable to the present system embodiments. The functions specifically implemented by the present system embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0115] The above is a specific description of the preferred implementation of the present invention, but the invention is not limited to the embodiments. Those skilled in the art may make various equivalent modifications or substitutions without violating the spirit of the present invention. These equivalent modifications or substitutions are all included in the scope defined by the claims of this application.
Claims
1. A video transcoding classification detection method based on intra-frame and inter-frame vector information, characterized in that: The following steps are involved: Construct original video dataset and transcoded video dataset; Performing intra-frame feature information extraction processing on the original video data set and the transcoded video data set to obtain intra-frame prediction feature information of the original video data set and intra-frame prediction feature information of the transcoded video data set; Performing inter-frame feature information extraction processing on the original video data set and the transcoded video data set to obtain inter-frame prediction feature information of the original video data set and inter-frame prediction feature information of the transcoded video data set; Based on the classifier, the intra-frame prediction feature information of the original video data set and the intra-frame prediction feature information of the transcoded video data set, as well as the inter-frame prediction feature information of the original video data set and the inter-frame prediction feature information of the transcoded video data set are classified and detected to obtain the video transcoding classification detection result.
2. The video transcoding classification detection method based on intra-frame and inter-frame vector information according to claim 1 is characterized in that: The step of constructing the original video dataset and the transcoded video dataset specifically includes: Get uncompressed video dataset; The uncompressed video data set is compressed by a HEVC video encoder and a VVC video encoder to obtain a first compressed video data set and a second compressed video data set; The first compressed video data set and the second compressed video data set are segmented and integrated to construct an original video data set; The first compressed video data set is first decoded and then compressed by a VVC video encoder to obtain a transcoded video data set.
3. The video transcoding classification detection method based on intra-frame and inter-frame vector information according to claim 2 is characterized in that: The step of extracting intra-frame feature information from the original video dataset and the transcoded video dataset to obtain intra-frame prediction feature information of the original video dataset and intra-frame prediction feature information of the transcoded video dataset specifically includes: By using a multi-tree CTU partitioning method, different CTU partitioning type blocks are extracted and marked for the original video data set and the transcoded video data set, so as to obtain several types of CU blocks of the original video data set and several types of CU blocks of the transcoded video data set; Determine frequency information of several types of CU blocks of the original video data set and frequency information of several types of CU blocks of the transcoded video data set; Acquire motion vector information of the original video data set and the transcoded video data set and map them to several types of CU blocks of the original video data set and several types of CU blocks of the transcoded video data set to obtain a mapping result; Based on the mapping results, the number of intra-frame key frames in the video stream is set, and the frequencies of different CU block types in the key frames of the video stream are calculated in combination with the frequency information to obtain the intra-frame prediction feature information of the original video dataset and the intra-frame prediction feature information of the transcoded video dataset.
4. The video transcoding classification detection method based on intra-frame and inter-frame vector information according to claim 3 is characterized in that: The calculation expression of the frequency information of the CU block is specifically as follows: In the above formula, Sρ i Indicates the frequency information of different CU blocks, N CU Indicates the number of different CU block types, S CU Represents the total number of different CU block types in a single frame, and i represents several CU block division types.
5. The video transcoding classification detection method based on intra-frame and inter-frame vector information according to claim 4 is characterized in that: The calculation expression for calculating the frequency of different CU block types in the key frame of the video stream is specifically as follows: In the above formula, represents the frequency of different CU block types, n represents the number of key frames in the video stream, Sρ i Represents the frequency information of different CU blocks, and k represents several separate key frames.
6. The video transcoding classification detection method based on intra-frame and inter-frame vector information according to claim 5 is characterized in that: The step of extracting inter-frame feature information from the original video dataset and the transcoded video dataset to obtain inter-frame prediction feature information of the original video dataset and inter-frame prediction feature information of the transcoded video dataset specifically includes: Setting a prediction direction of inter-frame prediction, the prediction direction includes unidirectional prediction and bidirectional prediction, and the unidirectional prediction includes forward prediction and backward prediction; Obtain motion vector information of the original video data set and the transcoded video data set and map them to different CU block partition type frequencies on the CTU; According to the frequency of different CU block partition types mapped to the CTU according to the motion vector information, the frequency of unidirectional prediction and the frequency of bidirectional prediction are determined in combination with the prediction direction of inter-frame prediction; The frequency of unidirectional prediction and the frequency of bidirectional prediction are combined to obtain inter-frame prediction feature information of the original video data set and inter-frame prediction feature information of the transcoded video data set.
7. The video transcoding classification detection method based on intra-frame and inter-frame vector information according to claim 6 is characterized in that: The calculation expression of the frequency of different CU block partition types mapped to the motion vector information on the CTU is specifically as follows: In the above formula, Indicates the frequency of different CU block partition types mapped to CTU by motion vector information, Sρ i It represents the frequency information of different CU blocks, n represents the number of key frames in the video stream, and j represents several CU block division types.
8. The video transcoding classification detection method based on intra-frame and inter-frame vector information according to claim 7 is characterized in that: The calculation expressions of the unidirectional prediction frequency and the bidirectional prediction frequency are specifically as follows: In the above formula, represents the frequency of one-way prediction, represents the frequency of bidirectional prediction, represents forward prediction, represents backward prediction, represents bidirectional prediction, and j represents several CU block partition types.
9. The video transcoding classification detection method based on intra-frame and inter-frame vector information according to claim 8, characterized in that: The step of classifying and detecting the intra-frame prediction feature information of the original video data set and the intra-frame prediction feature information of the transcoded video data set, as well as the inter-frame prediction feature information of the original video data set and the inter-frame prediction feature information of the transcoded video data set based on the classifier to obtain the video transcoding classification detection result specifically includes: The inter-frame prediction feature information of the original video data set is combined with the inter-frame prediction feature information of the transcoded video data set, and the intra-frame prediction feature information of the original video data set is combined with the intra-frame prediction feature information of the transcoded video data set to obtain the prediction features of the combined original video data set and the prediction features of the combined transcoded video data set; The prediction features of the fused original video data set and the prediction features of the fused transcoded video data set are standardized to obtain the prediction features of the standardized original video data set and the prediction features of the standardized transcoded video data set; The predicted features of the standardized original video dataset and the predicted features of the standardized transcoded video dataset are classified and detected by the LIbSvm classifier based on the RBF kernel function to obtain the video transcoding classification detection results.
10. A video transcoding classification detection system based on intra-frame and inter-frame vector information, characterized in that: Includes the following modules: The first module is used to construct the original video dataset and the transcoded video dataset; The second module is used to extract intra-frame feature information of the original video data set and the transcoded video data set to obtain intra-frame prediction feature information of the original video data set and intra-frame prediction feature information of the transcoded video data set; The third module is used to extract inter-frame feature information of the original video data set and the transcoded video data set to obtain inter-frame prediction feature information of the original video data set and inter-frame prediction feature information of the transcoded video data set; The fourth module is used to classify and detect the intra-frame prediction feature information of the original video data set and the intra-frame prediction feature information of the transcoded video data set, as well as the inter-frame prediction feature information of the original video data set and the inter-frame prediction feature information of the transcoded video data set based on the classifier to obtain the video transcoding classification detection result.