Depth video frame interpolation identification method and system based on double-layer routing attention and space-time inconsistency learning
By building a dual-stream network for deep video frame interpolation detection, combined with feature interactive learning of RGB information and noise information, the problems of low detection efficiency and insufficient generalization ability in the existing methods are solved, and more efficient deep video frame interpolation detection is achieved.
Patent Information
- Application Number
- CN202510332702.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-01
AI Technical Summary
When facing deep video frame interpolation, the existing deep video frame interpolation detection methods have problems such as low detection efficiency, great influence on multiple compression artifacts, low performance on unlearned encoding factors, performance deterioration in compression environments, and overfitting training data. The fusion of multimodal and multi-scale features is not fully considered, so the deep video frame interpolation operation cannot be effectively recognized.
The deep video frame interpolation detection method based on double-layer routing attention and space-time inconsistency learning is adopted. By building a dual-stream network, combining RGB information and filtered noise information, the feature interactive learning of time-level streams and frame-level streams is used to improve detection performance and pay attention to generalization performance.
It improves the accuracy and generalization ability of deep video frame interpolation detection, can effectively identify deep video frame interpolation operations, reduce visual artifacts, and enhances the ability to fusion of multimodal and multi-scale features.
Smart Images

Figure QLYQS_1 
Figure QLYQS_3 
Figure QLYQS_4
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video forensics, in particular to a deep video frame interpolation detection method based on bi-level routing attention and spatio-temporal inconsistency learning. Specifically, by combining the bi-level routing attention mechanism (BRA) with spatio-temporal inconsistency learning, the present invention proposes a binary classification detection method that only requires the original video data for the interpolated frames of a deep video, thus having important application value in scenarios such as video forensics and video tampering recognition. Background Art
[0002] Frame replication or frame averaging is the simplest video frame interpolation scheme, which inserts adjacent frames or average frames in the middle of two frames. This produces good visual effects in static scenes, but causes jitter and blurring in dynamic videos. Traditional video frame interpolation methods focus on the calculation of motion vectors, followed by motion correction and refinement of transition details. Although good results are produced, these methods can still produce visual artifacts such as blurring and ghosting in areas with complex motion and texture. These artifacts are caused by inaccurate motion estimation, illumination changes, and occlusion of objects. These are also the motivation for detecting conventional video frame interpolation. With the development of CNN, deep video frame interpolation has become prominent in this field by effectively alleviating the previously mentioned visual artifacts by generating more realistic high-frame-rate videos. They adopt deep generative models and obtain the transformation from a pair of adjacent frames to the intermediate frame through iterative refinement of model parameters. Their key lies in designing an innovative network architecture, plus a carefully designed loss function. The architecture usually includes two stages: motion estimation or feature matching and subsequent frame synthesis. The former aims to predict optical flow, convolutional kernels, bidirectional encoding, deformable convolutions, or motion and appearance extraction, while the latter creates an intermediate frame with more details based on the prediction results. In particular, the loss function helps the interpolated frame to match the real frame through backpropagation. In addition, the Transformer architecture and Diffusion models exhibit excellent generative capabilities and have gradually become popular architectures in this field. Since these interpolation models can generate more realistic interpolated frames, they pose a severe challenge to the detection of deep video frame interpolation.
[0003] Currently, the research on video frame interpolation detection is mainly divided into two categories. The first category is based on manual feature extraction, including legacy trace mining, forensic feature design, and SVM classification. For example, the original frame rate is deduced by predicting the periodic change of the error signal, the frame rate is estimated by the texture change curve, the artifact index map is designed by calculating the difference of interpolated pixels, and the frame rate is estimated by using noise change and motion effects. These methods rely on the visual artifacts generated by video frame interpolation operations, such as motion blur and boundary distortion. The second category uses deep learning to autonomously extract features and realizes video frame interpolation detection through preprocessing layers and spatio-temporal learning. For example, steganalysis CNN and hybrid CNN are used to detect interpolated frames, the deep video frame interpolation forensics method based on two-stream multi-scale spatio-temporal representation, and the deep video frame interpolation detection method with abnormal region awareness. However, deep video frame interpolation improves the fidelity of video generation and reduces visual traces, resulting in a decrease in the efficiency of existing detection methods. Existing methods also have limitations, such as being affected by multiple compression artifacts, having low performance for unlearned coding factors, deteriorating performance in a compressed environment, and overfitting to training data. In addition, existing research has not fully considered the fusion of multi-modal and multi-scale features and cannot predict the deep video frame interpolation models used. Therefore, there is an urgent need to design a specialized technology with excellent comprehensive performance to identify deep video frame interpolation operations. Summary of the Invention
[0004] The present invention aims to solve the problems existing in the prior art. To this end, the present invention proposes a deep video frame interpolation detection method and system based on double-layer routing attention and spatio-temporal inconsistency learning, constructs a two-stream network through RGB information and filtered noise information, and performs interactive learning on the features between the time-level stream and the frame-level stream, improving the detection performance while paying more attention to the generalization performance.
[0005] In a first aspect, an embodiment of the present invention provides a deep video frame interpolation detection method based on double-layer routing attention and spatio-temporal inconsistency learning. The deep video frame interpolation detection method based on double-layer routing attention and spatio-temporal inconsistency learning includes:
[0006] Obtain a sequence of 5 consecutive RGB video frames, and process the 5 consecutive RGB frames with a high-pass filter (HPF) to obtain a filtered frame sequence;
[0007] Construct a deep video frame interpolation detection network, perform patch embedding on the filtered frame sequence and then send it into the first BIFB module for frame-level feature extraction. Patch embedding can convert an image into sequence data, thus facilitating the Transformer model to process image data. Then input the RGB frame sequence into the time difference module (TDM) for time-level feature extraction. Among them, the obtained frame-level features and time-level features are interactively learned by a preset attention-based interactive feature fusion module (AIF2M) to obtain the fusion features of the first stage;
[0008] The first-stage fused features are first directly fed into the conv3_x residual block of ResNet18, and then after patch merging operation, they are fed into the second BIFB module, so as to perform temporal-level and frame-level feature extraction respectively. The patch merging operation can reduce the spatial resolution of the input feature image while increasing the number of channels. Then, the obtained frame-level features and temporal-level features are interactively learned using a preset attention-based interactive feature fusion module (AIF2M) to obtain the second-stage fused features;
[0009] The second-stage fused features are first directly fed into the conv4_x residual block of ResNet18, and then after patch merging operation, they are fed into the third BIFB module, so as to perform temporal-level and frame-level feature extraction respectively. Then, the obtained frame-level features and temporal-level features are interactively learned using a preset attention-based interactive feature fusion module (AIF2M) to obtain the third-stage fused features;
[0010] The third-stage fused features are input into the whole-part feature fusion module (WPF2M) for processing to obtain the final spatio-temporal features, and these features are input into a preset classifier to finally determine whether the video frame is an original frame or an interpolated frame.
[0011] According to some embodiments of the present invention, the method for extracting filtered frames from RGB video frames includes:
[0012] The RGB frame is grayscale-converted using the default weighted average method in the cvtColor function to obtain a grayscale image;
[0013] The obtained grayscale image is passed through a first predefined filter to remove low-frequency information to obtain new_img1;
[0014] The obtained new_img1 is normalized by dividing by 12 to obtain new_img2;
[0015] The obtained new_img2 is processed through a second predefined filter to obtain the filtered frame.
[0016] According to some embodiments of the present invention, the patch embedding operation includes two convolutional layers, which have a kernel size of 3×3, a stride size of 2, a padding value of 1, two batch normalization (BN) layers, and a Gaussian error linear unit (GELU) activation function; through the patch embedding operation, a video frame with a size of 224×224 is segmented into 56×56 blocks, and then each block is flattened into a one-dimensional vector and fed into the first BIFB block for feature extraction.
[0017] According to some embodiments of the present invention, the composition of the frame-level stream includes 3 BIFB modules, and the composition of the BIFB module is as follows:
[0018] A 3×3 depth convolution layer, a Layer Norm layer, a BRA module, a Layer Norm layer, and a multi-layer perceptron (MLP) with two layers and an expansion ratio of 3;
[0019] s in the 3 BIFB modules are all set to 8; the top-k of the first BIFB module is 1, the top-k of the second BIFB module is 4, and the top-k of the third BIFB module is 16.
[0020] According to some embodiments of the present invention, the working process of the bilayer routing attention mechanism (BRA) is as follows:
[0021] Region division and input prediction. First, the filtered frame I i filter ∈R H×W×C is divided into s regions, and each region has feature vectors. However, these regions do not overlap. Then I i filter becomes Next, the query, key, and value tensors Q, K, Q = X r W q , K = X r W k , V = X r W v (1) where W q , W k , W v ∈R C×C represent the corresponding projection weights of the query, key, and value.
[0022] Region-to-region routing index matrix. Then a directed graph is created to find the relationship between each given region and other regions. Specifically, by taking the average of Q and K for each region respectively, the region-level query Q r and the key K r ∈R S2×C are obtained. Secondly, through the matrix product of Q r and the transpose of K r , the adjacency matrix of the inter-region affinity graph Z r = Q r (K r ) T (2) Adjacency matrix Z r The entries of measure the semantic relatedness between two regions. Finally, the inter-region affinity graph for each region is pruned to contain only the top-k connections, and then the row-wise top-k operator is used to obtain the path index matrix I r = topkIndex(Z r ) (3) where the i-th row of I r contains the k indicators of the regions most relevant to region i.
[0023] Fine-grained token-to-token attention. For each query token in region i, the attention is directed over all key-value pairs concentrated by . However, effectively performing this step is challenging because these regions are likely to be distributed across the entire feature space. Therefore, the key tensor and value tensor are gathered as: K g = gather(K, I r ), V g = gather(V, I r ) (4) where K g , are the gathered key and value tensors. Next, the gathered key-value pairs are attended to by using the attention operation: O = Attention(Q, K g , V g ) + LCE(V) (5) where LCE(·) is the local context enhancement term, parameterized by a depth convolution with a kernel size of 5.
[0024] According to some embodiments of the present invention, each of the patch merging modules consists of a convolutional layer and a batch normalization (BN) layer. The convolutional layer has a kernel size of 3×3, a stride of 2, and a padding value of 1. The patch merging operation in S3 divides the feature map of size 56×56 into blocks of size 28×28, and the patch merging operation in S4 divides the feature map of size 28×28 into blocks of size 14×14. The patch merging operation gradually reduces the spatial resolution of the feature map in the deeper layers of the model while increasing its number of channels, helping the model capture more global features.
[0025] According to some embodiments of the present invention, the process of the time difference module (TDM) for extracting time-level features is as follows:
[0026] TDM takes five consecutive frames ({I t-2 , It-1 , I t , I t+1 , I t+2 The video group of {}) is processed into two branches;
[0027] In the first branch, it extracts features from the middle frame I t Extract features, and then obtain the feature map Y through 4 convolutional blocks and two max pooling layers t , where each convolutional block contains a convolutional layer with a kernel size of 3×3, a Batch Norm layer, and a ReLU activation function;
[0028] In the second branch, first perform frame difference operations on 5 consecutive frames, and then add the results of these frame differences. Next, use an average pooling layer to downsample the added information to minimize redundancy, and then use 1 convolutional block and a max pooling layer to derive the feature Y S , where the convolutional block contains a convolutional layer with a kernel size of 7×7, a Batch Norm layer, and a ReLU activation function.
[0029] In addition, 3 ConvGRU units with a kernel size of 3×3 are used to aggregate temporal information at different scales. Y S Capture temporal features through the first ConvGRU unit, then upsample, and add the resulting feature map Y Sc With the feature map Y t Add them together, and the resulting result is processed through 4 convolutional blocks to obtain Y t1 , where each convolutional block contains a convolutional layer with a kernel size of 3×3, a Batch Norm layer, and a ReLU activation function. At the same time, the feature map Y S Pass through 4 convolutional blocks, then through the second ConvGRU unit, and finally obtain the feature map Y through upsampling t2 , where each convolutional block contains a convolutional layer with a kernel size of 3×3, a BatchNorm layer, and a ReLU activation function. Finally, Y t1 And Y t2 Are added element-wise, and then input into the third ConvGRU unit to obtain the final feature Y out .
[0030] According to some embodiments of the present invention, the attention-based interactive feature fusion module (AIF2M) is set as follows:
[0031] The input features of the frame-level stream and the temporal-level stream both pass through an average pooling layer and a convolutional layer with a kernel size of 3, and then after a series of operations, pass through a depth convolutional layer with a kernel size of 3×3, a Batch Norm layer, a ReLU activation function, a convolutional layer with a kernel size of 1×1, and a Batch Norm layer.
[0032] According to some embodiments of the present invention, the overall-part feature fusion module (WPF2M) is arranged as follows:
[0033] The input features of both the frame-level stream and the temporal-level stream pass through a convolutional layer with a kernel size of 1×1, and then the input stream features of the frame-level pass through a convolutional layer with a kernel size of 1×1, a ReLU activation function, a Batch Norm layer, a convolutional layer with a kernel size of 1×1, and a Batch Norm layer. The input features of the temporal-level stream pass through an average pooling layer, a convolutional layer with a kernel size of 3, a ReLU activation function, a Batch Norm layer, a convolutional layer with a kernel size of 3, and a Batch Norm layer.
[0034] In a second aspect, an embodiment of the present invention provides a deep video frame interpolation detection system based on double-layer routing attention and spatio-temporal inconsistency learning. The deep video frame interpolation detection system based on double-layer routing attention and spatio-temporal inconsistency learning includes:
[0035] An image acquisition module: configured to acquire an RGB video frame sequence, and process five consecutive RGB frames with a high-pass filter (HPF) to obtain a filtered frame sequence;
[0036] A frame-level feature extraction stream: The first BIFB module extracts features from the filtered frame sequence after patch embedding operation, the second BIFB module extracts features from the first-stage fusion features after patch merging operation, and the third BIFB module extracts features from the second-stage fusion features after patch merging operation;
[0037] A temporal-level feature extraction stream: The temporal difference module extracts features from the RGB video frame sequence, the conv3_x residual block of ResNet18 extracts features from the first-stage fusion features, and the conv4_x residual block of ResNet18 extracts features from the second-stage fusion features;
[0038] An intermediate layer feature fusion module: configured to separately learn the frame-level features and temporal-level features in S2, S3, and S4, so as to obtain the first, second, and third-stage fusion features;
[0039] An overall-local feature fusion module: configured to process the third-stage fusion features to obtain the final spatio-temporal fusion features;
[0040] A video frame authenticity judgment module: configured to input the spatio-temporal fusion features into a preset classifier to obtain the authenticity situation of the RGB video frame output by the classifier.
[0041] Other features and advantages of the present invention will be set forth in the following description, and in part will be obvious from the description, or may be learned by practice of the present invention. Description of the Drawings
[0042] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the description of the embodiments in conjunction with the following drawings, where:
[0043] Figure 1 is a flowchart of a deep video frame interpolation detection method based on double-layer routing attention and spatio-temporal inconsistency learning provided by an embodiment of the present invention, and a structural diagram of a time difference module (TDM) and a BIFB module;
[0044] Figure 2 is a flowchart of extracting filtered frames from an RGB video frame sequence provided by an embodiment of the present invention;
[0045] Figure 3 is a flowchart of the operation of a double-layer routing attention mechanism (BRA) provided by an embodiment of the present invention;
[0046] Figure 4 is a schematic diagram of the feature processing process of a ConvGRU unit provided by an embodiment of the present invention;
[0047] Figure 5 is a structural diagram of an attention-based interactive feature fusion module (AIF2M) and a whole-part feature fusion module (WPF2M) provided by an embodiment of the present invention. Detailed Embodiments
[0048] The embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary only for explaining the present invention and should not be construed as limiting the present invention.
[0049] In the description of the present invention, if the first, second, etc. are described only for the purpose of distinguishing technical features, they should not be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence of the indicated technical features.
[0050] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by terms such as up, down, etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as limiting the present invention.
[0051] In the description of the present invention, it should be noted that, unless otherwise clearly defined, terms such as "set", "installed", "connected", etc. should be understood in a broad sense, and those skilled in the art can reasonably determine the specific meanings of the above terms in the present invention in combination with the specific content of the technical solution.
[0052] Next, the technical solution of the present invention will be clearly and completely described in conjunction with the accompanying drawings. Obviously, the following described embodiments are some embodiments of the present invention, not all embodiments.
[0053] Refer to Figure 1 , in some embodiments of the present invention, a deep video frame interpolation detection method based on double-layer routing attention and spatio-temporal inconsistency learning is provided, including:
[0054] S1: Obtain a continuous 5-frame RGB video frame sequence, and use a high-pass filter (HPF) to process the continuous 5-frame RGB frames to obtain a filtered frame sequence;
[0055] S2: Construct a deep video frame interpolation detection network, perform patch embedding on the filtered frame sequence and then send it into the first BIFB module for frame-level feature extraction. Patch embedding can convert an image into sequence data, which is convenient for the Transformer model to process image data. Then input the RGB frame sequence into the time difference module (TDM) for time-level feature extraction. Among them, the obtained frame-level feature and time-level feature are interactively learned using a preset attention-based interactive feature fusion module (AIF2M) to obtain the fusion feature of the first stage;
[0056] S3: First directly send the fusion feature obtained in S2 into the conv3_x residual block of ResNet18, and then perform patch merging operation and send it into the second BIFB module to perform time-level and frame-level feature extraction respectively. The patch merging operation can reduce the spatial resolution of the input feature image and increase the number of channels at the same time. Then the obtained frame-level feature and time-level feature are interactively learned using a preset attention-based interactive feature fusion module (AIF2M) to obtain the fusion feature of the second stage;
[0057] S4: First directly send the fusion feature obtained in S3 into the conv4_x residual block of ResNet18, and then perform patch merging operation and send it into the third BIFB module to perform time-level and frame-level feature extraction respectively. Then the obtained frame-level feature and time-level feature are interactively learned using a preset attention-based interactive feature fusion module (AIF2M) to obtain the fusion feature of the third stage;
[0058] S5: Input the third-stage fusion features obtained in S4 into the Whole-Part Feature Fusion Module (WPF2M) for processing to obtain the final spatio-temporal features, and input these features into a preset classifier to finally determine whether the video frame is an original frame or an interpolated frame.
[0059] First, through step S1, a high-pass filter (HPF) is used to process 5 consecutive RGB frames to obtain a filtered frame sequence. RGB information and the filtered noise information are used, and the relationship between the two types of information is considered to promote the learning of the constructed two-stream network. Then, through the deep video frame interpolation detection network constructed in step S2, the filtered frame sequence is patch-embedded and then sent to the first BIFB module for frame-level feature extraction. Patch embedding can convert an image into sequence data, thus facilitating the Transformer model to process image data. Then, the RGB frame sequence is input into the Temporal Difference Module (TDM) for temporal-level feature extraction. Among them, the obtained frame-level features and temporal-level features are interactively learned using a preset Attention-based Interactive Feature Fusion Module (AIF2M) to obtain the first-stage fusion features. Then, through step S3, the fusion features obtained in S2 are first directly sent into the conv3_x residual block of ResNet18, and then after patch merging operation, they are sent to the second BIFB module to respectively perform temporal-level and frame-level feature extraction. The patch merging operation can reduce the spatial resolution of the input feature image while increasing the number of channels. Then, the obtained frame-level features and temporal-level features are interactively learned using a preset Attention-based Interactive Feature Fusion Module (AIF2M) to obtain the second-stage fusion features. Then, through step S4, the fusion features obtained in S3 are first directly sent into the conv4_x residual block of ResNet18, and then after patch merging operation, they are sent to the third BIFB module to respectively perform temporal-level and frame-level feature extraction. Then, the obtained frame-level features and temporal-level features are interactively learned using a preset Attention-based Interactive Feature Fusion Module (AIF2M) to obtain the third-stage fusion features. Finally, through step S5, the third-stage fusion features obtained in S4 are input into the Whole-Part Feature Fusion Module (WPF2M) for processing to obtain the final spatio-temporal features, and these features are input into a preset classifier to finally determine whether the video frame is an original frame or an interpolated frame.
[0060] Refer to Figure 2 , in some embodiments of the present invention, extracting filtered frames from 5 consecutive RGB video frames includes:
[0061] Convert the RGB frame to grayscale using the default weighted average method in the cvtColor function to obtain a grayscale image;
[0062] The obtained grayscale image is passed through a first predefined filter to remove low-frequency information, resulting in new_img1;
[0063] The obtained new_img1 is normalized by dividing it by 12 to obtain new_img2;
[0064] The obtained new_img2 is processed through a second predefined filter to obtain the filtered frame.
[0065] Preferably, the size of the input RGB frame is 3×H×W, where 3 represents the three RGB channels, and H and W correspond to the height and width respectively.
[0066] It should be noted that by using a high-pass filter (HPF), the high-frequency information (such as edges, details, noise) of the RGB frame can be retained and the low-frequency information (such as smooth background) can be removed. In video frame interpolation detection, it can help extract the abnormal traces introduced by the interpolation operation, thereby improving the detection accuracy.
[0067] In some embodiments of the present invention, the patch embedding operation includes two convolutional layers with a kernel size of 3×3, a stride size of 2, a padding value of 1, two batch normalization (BN) layers, and a Gaussian error linear unit (GELU) activation function.
[0068] It should be noted that through the patch embedding operation, the input video frame with a size of 224×224 is divided into 56×56 blocks, and then each block is flattened into a one-dimensional vector and sent into the first BIFB block for feature extraction.
[0069] Refer to Figure 1 , in some embodiments of the present invention, the composition of the BIFB module is as follows:
[0070] A 3×3 depth convolutional layer, a Layer Norm layer, a BRA module, a Layer Norm layer, and a two-layer multi-layer perceptron (MLP) with an expansion ratio of 3;
[0071] It should be noted that s in the 3 BIFB modules is set to 8; the top-k of the first BIFB module is 1, the top-k of the second BIFB module is 4, and the top-k of the third BIFB module is 16.
[0072] Refer to Figure 3 , in some embodiments of the present invention, the working process of the bilayer routing attention mechanism (BRA) is as follows:
[0073] Region division and input prediction. First, the filtered frame I i filter ∈R H×W×C is divided into s regions, and each region has A feature vector. However, these regions do not overlap. Then I i filter becomes Next, the query, key, and value tensors Q, K, Q = X r W q , K = X r W k , V = X r W v (1) where W q , W k , W v ∈ R C×C represent the corresponding projection weights for the query, key, and value.
[0074] The region-to-region routing index matrix. Then a directed graph is created to find the relationships between each given region and other regions. Specifically, by averaging the Q and K for each region respectively, the region-level query Q r and key Secondly, through the matrix product of Q r and the transpose of K r , the adjacency matrix of the inter-region affinity graph is obtained Z r = Q r (K r ) T (2) The entries of the adjacency matrix Z r measure the semantic relatedness of two regions. Finally, the inter-region affinity graph for each region is pruned to contain only the top-k connections, and then the row-wise top-k operator is used to obtain the path index matrix I r = topkIndex(Z r ) (3) where the i-th row of I r contains the k indicators of the regions most relevant to region i.
[0075] Fine-grained token-to-token attention. For each query token in region i, the attention is directed at all key-value pairs concentrated by . However, effectively performing this step is challenging because these regions are likely to be distributed across the entire feature space. Therefore, the key tensor and value tensor are collected as: K g=gather(K, I r ), V g =gather(V, I r ) (4) where K g , are the key and value tensors for gathering. Next, the gathered key-value pairs are attended to by using an attention operation: O = Attention(Q, K g , V g ) + LCE(V) (5) where LCE(·) is the local context enhancement term, parameterized by a depth convolution with a kernel size of 5.
[0076] According to some embodiments of the present invention, each of the patch merging modules consists of a convolutional layer and a batch normalization (BN) layer. The convolutional layer has a kernel size of 3×3, a stride of 2, and a padding value of 1.
[0077] It should be noted that the patch merging operation in S3 divides the feature map of size 56×56 into 28×28 blocks, and the patch merging operation in S4 divides the feature map of size 28×28 into 14×14 blocks. The patch merging operation gradually reduces the spatial resolution of the feature map while increasing its number of channels in the deeper layers of the model, which helps the model capture more global features.
[0076] Referring to Figure 1 , in some embodiments of the present invention, the process of the temporal difference module (TDM) for extracting temporal-level features is as follows:
[0077] TDM processes a group of five consecutive frames ({I t-2 , I t-1 , I t , I t+1 , I t+2}) of a video into two branches.
[0027] In the first branch, it extracts features from the middle frame I t , and then passes through 4 convolutional blocks and two max pooling layers to obtain the feature map Y t , where each convolutional block contains a convolutional layer with a kernel size of 3×3, a Batch Norm layer, and a ReLU activation function;
[0078] In the second branch, first, frame difference operations are performed on the five consecutive frames, and then the frame difference results are added together. Next, average pooling is used to downsample the added information to minimize redundancy, and then 1 convolutional block and a max pooling layer are used to derive the feature Y S, where the convolutional block includes a convolutional layer with a kernel size of 7×7, a Batch Norm layer, and a ReLU activation function.
[0079] It should be noted that in the structure of the entire Temporal Difference Module (TDM), 3 ConvGRU units with a kernel size of 3×3 are used to aggregate temporal information at different scales. S Capture temporal features through the first ConvGRU unit, then perform upsampling, and add the resulting feature map Sc to the feature map t The result is processed through 4 convolutional blocks to obtain t1 , where each convolutional block includes a convolutional layer with a kernel size of 3×3, a Batch Norm layer, and a ReLU activation function. At the same time, the feature map S passes through 4 convolutional blocks, then through the second ConvGRU unit, and finally through upsampling to obtain the feature map t2 , where each convolutional block includes a convolutional layer with a kernel size of 3×3, a Batch Norm layer, and a ReLU activation function. Finally, t1 and t2 are added element-wise, and then input into the third ConvGRU unit to obtain the final feature out .
[0080] Referring to Figure 4 , in some embodiments of the present invention, the feature processing process of the ConvGRU unit is as follows:
[0081] The input feature is first divided into two independent parts along the channel axis. One part is directly input into the ConvGRU unit to capture temporal motion and appearance information, and the other part retains the initial spatial features. Finally, these two sets of features are concatenated to obtain the final output.
[0082] Referring to Figure 5 , in some embodiments of the present invention, the setting of the Attention-based Interactive Feature Fusion Module (AIF2M) is as follows:
[0083] The input features of the frame-level flow and the temporal-level flow both pass through an average pooling layer and a convolutional layer with a kernel size of 3, and then after a series of operations, pass through a depth convolutional layer with a kernel size of 3×3, a Batch Norm layer, a ReLU activation function, a convolutional layer with a kernel size of 1×1, and a Batch Norm layer.
[0084] It should be noted that the frame-level flow is used to capture the inconsistencies in the spatial structure, while the temporal-level flow can be used to highlight the inconsistencies in the temporal motion. To fully supplement these two types of inconsistencies, it is beneficial to represent these features in the middle layer of the network to prevent excessive network layers from affecting the extraction of inconsistent information.
[0085] Referring to Figure 5 , in some embodiments of the present invention, the whole-part feature fusion module (WPF2M) is set as follows:
[0086] The input features of both the frame-level stream and the temporal-level stream pass through a convolutional layer with a kernel size of 1×1, and then the input stream features of the frame-level pass through a convolutional layer with a kernel size of 1×1, a ReLU activation function, a Batch Norm layer, a convolutional layer with a kernel size of 1×1, and a Batch Norm layer. The input features of the temporal-level stream pass through an average pooling layer, a convolutional layer with a kernel size of 3, a ReLU activation function, a Batch Norm layer, a convolutional layer with a kernel size of 3, and a Batch Norm layer.
[0087] It should be noted that since the final decision result of the deep video frame interpolation detection depends on the final recognition features of the two streams, a specific module must be used to effectively integrate the features of the two streams. Traditional feature fusion modules usually use summation or concatenation, but these modules cannot effectively utilize the features of the two streams. Since the whole-part pattern can fully represent the important information in the feature mapping process, the whole-part feature fusion module (WPF2M) is designed to select discriminative features from the whole and part perspectives for the final frame-level decision.
[0088] Referring to Figure 1 , for the convenience of those skilled in the art to understand, a specific embodiment of the present invention provides a deep video frame interpolation detection method based on double-layer routing attention and spatio-temporal inconsistency learning, including:
[0089] The first step is to implement based on the Pytorch deep learning framework and use the UCF101 dataset. UCF101 contains 101 action categories. To maximize the diversity of the test dataset, we select 100 from 101 video categories and randomly extract 4 videos from each category. For these 400 videos, the training and test subsets are set in a ratio of 4:1, and the resolution remains the original 240P. DAVIS covers four evenly distributed categories (humans, animals, vehicles, objects) and several actions, with rich content, and each video is 2 - 4 seconds long. We select 80 videos from DAVIS as a supplementary test dataset for cross-dataset experiments and adjust their resolution to 240P due to resource limitations. The original frame rate of both is 15fps. After performing the frame sequence extraction operation on the video, the input is then adjusted to a size of 224×224 and input into the constructed network.
[0090] Step 2: Use a high-pass filter (HPF) in the frame-level stream to process the input RGB frame sequence to retain the high-frequency information (such as edges, details, noise) of the RGB frames. The filtered frame sequence obtained after processing is used as the input of the frame-level stream. The RGB frame sequence is used as the input of the temporal-level stream. The frame-level stream and the temporal-level stream extract features from the two inputs respectively, and then the entire network fuses and learns the features of the input, and inputs the finally obtained result into the classifier for discrimination.
[0091] During the training process of the network, cross-entropy is used to calculate the loss and backpropagation is performed to update the weights of the network. When the set conditions are met, the training stops, and a trained network is obtained. Input the image to be detected into the trained network, and the forged interpolated frames can be discriminated.
[0092] Refer to Figure 1 , an embodiment of the present invention further provides a deep video frame interpolation detection system based on double-layer routing attention and spatio-temporal inconsistency learning, including an image acquisition module, a frame-level feature extraction stream, a temporal-level feature extraction stream, an intermediate layer feature fusion module, a global-local feature fusion module, and a video frame authenticity judgment module.
[0093] It should be noted that since the deep video frame interpolation detection system based on double-layer routing attention and spatio-temporal inconsistency learning in this embodiment has the same inventive concept as the above-mentioned deep video frame interpolation detection method based on double-layer routing attention and spatio-temporal inconsistency learning. Therefore, the corresponding content in the method embodiment is also applicable to the system embodiment of the present invention, and will not be elaborated here.
[0094] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "illustrative embodiments", "examples", "specific examples", or "some examples" etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0095] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and purposes of the present invention. The scope of the present invention is defined by the claims and their equivalents.
Claims
1. A deep video frame interpolation detection method based on two-layer routing attention and spatiotemporal inconsistency learning, characterized in that: The method comprises the following steps: S1: Obtain a sequence of 5 consecutive RGB video frames, and use a high-pass filter (HPF) to process the 5 consecutive RGB frames to obtain a filtered frame sequence; S2: Build a deep video frame interpolation detection network, embed the filtered frame sequence into patches and then send it to the first BIFB module for frame-level feature extraction. Patch embedding can convert images into sequence data, which makes it easier for the Transformer model to process image data. Then the RGB frame sequence is input into the temporal difference module (TDM) for temporal feature extraction. Among them, the obtained frame-level features and temporal-level features are interactively learned using the preset attention-based interactive feature fusion module (AIF2M) to obtain the fusion features of the first stage; S3: The fused features obtained in S2 are first directly sent to the conv3_x residual block of ResNet18, and then sent to the second BIFB module after the patch merging operation, so as to extract features at the time level and frame level respectively. The patch merging operation can reduce the spatial resolution of the input feature image and increase the number of channels. Then the obtained frame-level features and time-level features are interactively learned using the preset attention-based interactive feature fusion module (AIF2M) to obtain the fused features of the second stage; S4: The fused features obtained in S3 are first directly sent to the conv4_x residual block of ResNet18, and then sent to the third BIFB module after the patch merging operation, so as to extract features at the time level and frame level respectively. Then the obtained frame-level features and time-level features are interactively learned using the preset attention-based interactive feature fusion module (AIF2M) to obtain the fused features of the third stage; S5: The third-stage fusion features obtained in S4 are input into the whole-part feature fusion module (WPF2M) for processing to obtain the final spatiotemporal features, and the features are input into a preset classifier to finally determine whether the video frame is an original frame or an interpolated frame.
2. The deep video frame interpolation detection method based on dual-layer routing attention and spatiotemporal inconsistency learning according to claim 1, characterized in that: The method for extracting a filter frame from an RGB video frame in S1 includes: The RGB frame is converted to grayscale using the default weighted average method in the cvtColor function to obtain a grayscale image; The obtained grayscale image is passed through the first predefined filter to remove low-frequency information to obtain new_img1; Divide the obtained new_img1 by 12 and normalize it to get new_img2; The obtained new_img2 is processed by a second predefined filter to obtain the filtered frame.
3. The deep video frame interpolation detection method based on dual-layer routing attention and spatiotemporal inconsistency learning according to claim 1, characterized in that: The patch embedding operation in S2 consists of two convolutional layers with a kernel size of 3×3, a stride size of 2, a padding value of 1, two batch normalization (BN) layers, and a Gaussian error linear unit (GELU) activation function; the input video frame of size 224×224 is split into 56×56 blocks through the patch embedding operation, and then each block is flattened into a one-dimensional vector and fed into the first BIFB block for feature extraction. The patch merging operation in S3 and S4 consists of a convolutional layer and a batch normalization (BN) layer, and the kernel size of the convolutional layer is 3×3, the stride size is 2, and the padding value is 1. The patch merging operation in S3 splits the feature map of size 56×56 into 28×28 blocks, and the patch merging operation in S4 splits the feature map of size 28×28 into 14×14 blocks. The patch merging operation gradually reduces the spatial resolution of the feature map in the deep layer of the model while increasing its number of channels, which helps the model capture more global features.
4. The deep video frame interpolation detection method based on dual-layer routing attention and spatiotemporal inconsistency learning according to claim 1, characterized in that: The frame-level stream consists of three BIFB modules, which are: 3×3 deep convolutional layer, Layer Norm layer, BRA module, Layer Norm layer and a two-layer multilayer perceptron (MLP) with a dilation ratio of 3; s in the three BIFB modules is set to 8; the top-k of the first BIFB module is 1, the top-k of the second BIFB module is 4, and the top-k of the third BIFB module is 16.
5. The deep video frame interpolation detection method based on dual-layer routing attention and spatiotemporal inconsistency learning according to claim 3, characterized in that: The workflow based on the two-layer routing attention mechanism (BRA) is as follows: Region division and input prediction. First, the filtered frame Divided into s regions, each region has eigenvectors. But these regions do not overlap. Then becomes Next, we use linear projection to derive the query, key, and value tensors Q, K. Q=X r W q ,K=X r W k ,V=X r W v (1) Where W q , W k , W v ∈R C×C Represents the corresponding projection weights for query, key, and value. A region-to-region routing index matrix. Then a directed graph is created to find the relationship between each given region and other regions. Specifically, the region-level query Q is obtained by averaging Q and K for each region separately. r and key Secondly, through Q r The matrix product and K r The adjacency matrix of the affinity graph between regions is obtained by transposing Z r =Q r (K r ) T (2) Adjacency Matrix Z r The entries of measure the semantic relevance of two regions. Finally, the inter-region affinity graph of each region is pruned to contain only top-k connections, and then the path index matrix is obtained using the row-by-row top-k operator I r =topkIndex(Z r ) (3) Among them I r The i-th row of contains the k indicators of the regions that are most relevant to region i. Fine-grained token-to-token attention. For each query token in region i, attention is directed to the All key-value pairs in the set. However, performing this step efficiently is challenging because these regions are likely to be distributed across the entire feature space. Therefore, the key tensor and value tensor are collected as: K g =gather(K,I r ),V g =gather(V,I r ) (4) Where K g , are the gathered key and value tensors. Next, we focus on the gathered key-value pairs by using the attention operation: O=Attention(Q,K g ,V g )+LCE(V) (5) where LCE(·) is the local context enhancement term parameterized by a depthwise convolution with a kernel size of 5.
6. The deep video frame interpolation detection method based on dual-layer routing attention and spatiotemporal inconsistency learning according to claim 1, characterized in that: The process of extracting time-level features by the temporal difference module (TDM) in S2 is as follows: TDM converts five consecutive frames ({I t-2 , I t-1 , I t , I t+1 , I t+2 })'s video group is processed into two branches; In the first branch, it starts from the intermediate frame I t Extract features, then pass through 4 convolution blocks and two maximum pooling layers to get the feature map Y t , where each convolution block contains a convolution layer with a kernel size of 3×3, a Batch Norm layer, and a ReLU activation function; In the second branch, the frame difference operation is first performed on five consecutive frames, and then the frame difference results are added. Next, the added information is downsampled using an average pooling layer to minimize redundancy, and then 1 convolution block and a maximum pooling layer are used to derive the feature Y S , where the convolution block contains a convolution layer with a kernel size of 7×7, a Batch Norm layer, and a ReLU activation function. In addition, three ConvGRU units with a kernel size of 3×3 are used to aggregate temporal information at different scales. S The temporal features are captured by the first ConvGRU unit, then upsampled, and the resulting feature map Y Sc With feature map Y t The result is processed by 4 convolution blocks to get Y t1 , where each convolution block contains a convolution layer with a kernel size of 3×3, a Batch Norm layer, and a ReLU activation function. At the same time, the feature map Y S Through 4 convolution blocks, then through the second ConvGRU unit, and finally through upsampling to obtain the feature map Y t2 , where each convolution block contains a convolution layer with a kernel size of 3×3, a Batch Norm layer, and a ReLU activation function. Finally, Y t1 and Y t2 It is added element by element and then input into the third ConvGRU unit to get the final feature Y out .
7. The deep video frame interpolation detection method based on dual-layer routing attention and spatiotemporal inconsistency learning according to claim 1, characterized in that: The settings of the attention-based interactive feature fusion module (AIF2M) in S2, S3, and S4 are as follows: The input features of the frame-level stream and the time-level stream are passed through the average pooling layer and the convolution layer with a kernel size of 3, and then after a series of operations, they pass through the deep convolution layer with a kernel size of 3×3, the Batch Norm layer, the ReLU activation function, the convolution layer with a kernel size of 1×1, and the Batch Norm layer.
8. The deep video frame interpolation detection method based on dual-layer routing attention and spatiotemporal inconsistency learning according to claim 1, characterized in that: The settings of the whole-part feature fusion module (WPF2M) in S5 are as follows: The input features of the frame-level stream and the time-level stream are all passed through a convolutional layer with a kernel size of 1×1. Then the input stream features of the frame level pass through a convolutional layer with a kernel size of 1×1, a ReLU activation function, a Batch Norm layer, a convolutional layer with a kernel size of 1×1, and a Batch Norm layer. The input features of the time-level stream pass through an average pooling layer, a convolutional layer with a kernel size of 3, a ReLU activation function, a Batch Norm layer, a convolutional layer with a kernel size of 3, and a Batch Norm layer.
9. A detection system for implementing the method according to any one of claims 1 to 8, characterized in that: The system includes: Image acquisition module: obtains RGB video frame sequence, and uses high-pass filter (HPF) to process 5 consecutive RGB frames to obtain filtered frame sequence; Frame-level feature extraction flow: The first BIFB module extracts features from the filtered frame sequence after the patch embedding operation, the second BIFB module extracts features from the first-stage fused features after the patch merging operation, and the third BIFB module extracts features from the second-stage fused features after the patch merging operation; Time-level feature extraction flow: The temporal difference module extracts features from the RGB frame sequence, the conv3_x residual block of ResNet18 extracts features from the first-stage fusion features, and the conv4_x residual block of ResNet18 extracts features from the second-stage fusion features; Intermediate layer feature fusion module: used to learn the frame-level features and time-level features in S2, S3 and S4 respectively, so as to obtain the fusion features of the first, second and third stages; Global-local feature fusion module: used to process the fusion features of the third stage to obtain the final spatiotemporal fusion features; Video frame authenticity judgment module: used to input the spatiotemporal fusion features into a preset classifier to obtain the authenticity of the RGB video frame output by the classifier.