Video restoration tampering detection method based on noise residual error and average optical flow
By extracting the noise residual and average optical flow mode of the video, using SegFormer to generate multi-scale features and perform feature interaction, the problem of tamper repair in the prior art is solved, and the accurate positioning and detection of the tampering area is achieved.
Patent Information
- Application Number
- CN202510291159.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art fails to effectively utilize the inherent characteristics of the video when detecting deep video repair tampering, and cannot fundamentally change the noise and optical flow patterns of the tampering video, making tampering traces difficult to detect.
By extracting the noise residual and average optical flow mode of the video, multi-scale features are generated using SegFormer as the encoder of the backbone network, and inter-feature interaction and multi-scale features are fused through the decoder to locate the tamper area.
Passive evidence collection for deep video repair tampering can be realized, the tampering area can be accurately identified and positioned, and the accuracy of tampering detection is improved.
Smart Images

Figure SMS_1 
Figure SMS_2 
Figure SMS_3
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of multimedia content security, and particularly relates to a video repair tampering detection method based on noise residuals and average optical flow Background Art
[0002] With the popularization of portable video acquisition devices such as smart phones and tablet computers, as well as the rapid development of Internet technology, the shooting, editing and dissemination of videos have become increasingly convenient. At the same time, video editing software represented by Adobe Premiere, Monkey, etc. has become increasingly powerful, making the acquisition and tampering of multimedia information simple and easy. In this context, it is of great theoretical and practical significance to distinguish the authenticity of videos or detect whether videos have been tampered with. Therefore, digital video forensics technology urgently needs to be developed
[0003] In recent years, some research results have been accumulated in this field. Wei et al. used the intra-frame and inter-frame residuals of videos to enhance the traces of repaired frames. The extraction of inter-frame residuals was obtained based on optical flow-based frame alignment. In order to learn discriminative features from frame residuals, a two-stream network was designed as an encoder, and a bidirectional convolutional LSTM was embedded into the decoder network to generate per-pixel prediction results of the repaired area for each frame. Yao et al. extracted motion residuals, used 3D convolution and a spatially rich model for filtering enhancement, and at the same time extracted the filtered residuals based on the Laplacian of Gaussian filter and the Laplacian filter, and proposed a two-stream video detection network, including a convNeXt two-stream encoder and a multi-scale feature cross-fusion decoder
[0004] Although certain development has been achieved in passive forensics of deep video repair, the existing methods for detection have not yet considered analyzing tampered videos after deep video repair from the perspective of the characteristics of the video itself. Although deep repair technology can improve or replace some content in the video to a certain extent, making the tampering traces visually imperceptible, it cannot fundamentally change the inherent characteristics and laws of the video. After the original video is deeply repaired, for the repaired frames, although the video content is tampered with, the noise and optical flow patterns that have existed since the video was created cannot be modified. Therefore, it is necessary to consider passive forensics of deep video repair from this aspect
[0005] Therefore, the present invention proposes a video repair tampering detection method based on noise residuals and average optical flow Summary of the Invention
[0006] The purpose of the present invention is to propose a hybrid network method based on noise residuals and average optical flow patterns, and to realize passive forensics of deep video repair tampering by analyzing the inherent attributes of tampered videos processed by deep video repair methods
[0007] To achieve the above object, the present invention provides the following technical solutions:
[0008] S1: Extract the noise, forward optical flow pattern, and backward optical flow pattern of the video, obtain the noise residual based on the noise, and take the average of the forward and backward optical flow patterns to obtain the average optical flow pattern;
[0009] S2: Input the noise residual and the average optical flow pattern into an encoder with SegFormer as the backbone network to generate multi-scale features;
[0010] S3: Use the decoder to perform intra-feature interaction on the generated multi-scale features of the noise residual and the multi-scale features of the average optical flow pattern;
[0011] S4: Pass the interacted features through a multi-scale feature fusion module to perform inter-feature interaction;
[0012] S5: Combine all the features to obtain the tampering localization area of the final depth-repaired video.
[0013] In summary, the present invention proposes a video repair tampering detection method based on noise residual and average optical flow. This method first analyzes the intrinsic attributes of the video, thereby extracting the noise, forward optical flow pattern, and backward optical flow pattern of the video, then obtaining the noise residual based on the noise, and taking the average of the forward and backward optical flow patterns to obtain the input of the hybrid network. At the same time, this method also proposes a hybrid network composed of an encoder and a decoder. The decoder will obtain two types of multi-scale features generated by the noise residual and the average optical flow pattern from the encoder. First, perform intra-feature interaction on them, and then perform inter-feature interaction to obtain the final output, and obtain the tampering localization area of the final depth-repaired video. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 are examples of video repair frames, masks, and extracted features of the present invention;
[0015] Figure 2 is the self-attention calculation process of the present invention;
[0016] Figure 3 is the network framework diagram of the present invention;
[0017] Figure 4 is the multi-scale feature fusion module of the present invention; DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] In order to make the object and advantages of the present invention clearer, the present invention will be specifically described below in conjunction with embodiments. It should be understood that the following text is only used to describe one or several specific implementation manners of the present invention, and does not strictly limit the specific protection scope claimed by the present invention.
[0019] Example:
[0020] S1: Extract the noise, forward optical flow pattern, and backward optical flow pattern of the video. Obtain the noise residual based on the noise, and take the average of the forward and backward optical flow patterns to obtain the average optical flow pattern;
[0021] S2: Input the noise residual and the average optical flow pattern into the encoder with SegFormer as the backbone network to generate multi-scale features;
[0022] S3: Use the decoder to perform feature internal interaction on the generated multi-scale features of the noise residual and the multi-scale features of the average optical flow pattern;
[0023] S4: Pass the interacted features through the multi-scale feature fusion module to perform feature inter interaction;
[0024] S5: Combine all the features to obtain the tampering localization area of the final depth restored video.
[0025] The process of obtaining the noise residual and the average optical flow pattern in S1 can be expressed by the following mathematical formula:
[0026] N_R t = F(f t+1 ) - F(f t ) (1)
[0027]
[0028] In the tampered video generated by the depth video restoration technology, the noise of the video is extracted using a high-pass filter to obtain the spatial domain features. Given a video V = {…, f_t, …}, where f_t represents the t-th frame of the video. After converting it to a grayscale image, a high-pass filter is implemented by using a two-dimensional convolution with a step size of 1. The input is convolved with a set of filter kernels, and then the convolution results are concatenated together as the input to the subsequent network to obtain the noise of the video, highlighting the restored area. Then, the noise residual is obtained using the noise of consecutive frames, leaving only the tampered area. Here, F represents the high-pass filter, and N_R t represents the noise residual.
[0029] At the same time, calculate the light component in the direction of the brightness gradient, extract the forward optical flow pattern and the backward optical flow pattern, and then take the average of them to better aggregate the motion information of the video. Here, avg_flow_pattern represents the average of the extracted forward optical flow pattern and backward optical flow pattern, and I represents the function for extracting the optical flow pattern of two consecutive frames. Removing regions of the video or restoring missing regions of the video inevitably destroys the latent features of the video and damages the integrity of the pre-established noise pattern and motion changes. The video restoration frames, masks, and extracted features are as Figure 1 shown.
[0030] The process of generating multi-scale features in S2 can be described as follows:
[0031] The encoder module is based on the backbone network of SegFormer, as Figure 3 shown, to generate multi-scale features of a given input, which provide high-resolution coarse features and low-resolution fine-grained features. Specifically, for an input of size H×W×3, hierarchical feature maps [C_1, C_2, C_3, C_4] are obtained through patch merging, with sizes that are 1 / 4, 1 / 8, 1 / 16, and 1 / 32 times that of the original input respectively. This encoder consists of 4 stacked SegFormer blocks, and each SegFormer block includes three parts: efficient self-attention, hybrid feed-forward network, and overlapping patch embedding. Among them:
[0032] Efficient self-attention uses an operation with spatial reduction based on self-attention. A model using self-attention considers the interdependencies between all elements of the entire input sequence during the calculation process, and the computational complexity increases quadratically with the growth of the sequence length, which may lead to excessive consumption of computational resources. To overcome this problem, efficient self-attention with a spatial reduction operation is introduced, which reduces the computational amount while maintaining the capture of key information and improving the model efficiency. The process of self-attention is as Figure 2 shown and is expressed by Equation 3 as:
[0033]
[0034] The hybrid feed-forward network simulates the ability of position perception by introducing convolutional layers with specific parameter settings. Specifically, a convolutional layer with a kernel size of 3×3, a stride of 2, and a zero-padding of 1 is used to replace the position encoding, which cleverly realizes the position sensitivity of the input data, retains to a certain extent the advantages of spatial consistency of convolutional neural networks in processing images, and combines the advantages of the Transformer architecture in global information processing and parallel computing. It is expressed by Equation 4 as:
[0035] x out = MLP(GELU(Conv(MLP(x in )))) + x in (4)
[0036] where x_in is the feature of the self-attention module, MLP represents the multi-layer perceptron, Conv represents the convolutional layer, and GELU represents the activation function.
[0037] Overlapping block embedding uses overlapping blocks for merging to model local continuous information. By dividing the input data into multiple overlapping blocks and performing feature extraction and embedding operations on these blocks separately, the sensitivity and understanding ability of the model to local information are enhanced. Compared with the traditional non-overlapping block processing method, overlapping block embedding allows a certain overlapping area between adjacent blocks. In this way, when splicing or fusing the embedding results of different blocks, more context detail information can be retained and fused, which is expressed by Equation 5 as follows:
[0038] x out =LN(Flatten(Conv(x in ))) (5)
[0039] Where x_in is the feature output by the hybrid feed-forward network, Conv represents the convolutional layer, the convolutional kernel is 7×7, the stride is 4, the zero-padding is 3, Flatten is to flatten the last two dimensions of the feature, and LN represents normalization.
[0040] The process of intra-feature interaction in S3 can be described as follows:
[0041] Intra-feature interaction first fuses each type of feature from top to bottom, and always fuses the higher-level feature with the lower-level feature. This is because high-level features usually contain more abstract and global semantic information, while low-level features focus on local details and edge information. Finally, the two types of features are concatenated on the channel, enabling the model to utilize both high-level global semantic features and low-level local detail features simultaneously. This process can be specifically described as follows:
[0042] f i =Conv([up ×2 (n i+1 )+n i ,up ×2 (m i+1 )+m i ),i∈{0,1,2}(6a)
[0043] f i =Conv([n i ,m i ),i=4(6b)
[0044] Where, UP ×2 represents 2x upsampling, [·] represents channel concatenation, and Conv represents the convolutional layer.
[0045] The process of inter-feature interaction in S4 can be described as follows:
[0046] The interaction between features refers to the Multi-scale Feature Fusion Module (MFFM), which uses addition, multiplication, and channel concatenation operations to fuse features between adjacent levels, achieving cross-level complementarity and collaboration between features. The multiplication operation is to eliminate redundant information between two features, and the channel concatenation operation is to integrate complementary information. This module adopts a top-down fusion method. First, the input features are upsampled by a factor of 2 to have the same spatial resolution as the features of the adjacent level, and then multiplication, addition, and channel concatenation operations are performed in sequence. After the obtained fused features are adjusted in dimension and resolution through 2 convolutional layers, they are added to the input features, and the output is obtained through an activation layer. This process is as Figure 4 shown, and can be specifically described as follows:
[0047]
[0048] f a = f m + f i ′ +1 + f i (7c)
[0049] f i ′ +1 = ReLU(f i+1 + BN(Conv(ReLU(BN(Conv(f cc )))))), i ∈ {0, 1, 2}(7d)
[0050]
[0051] Finally, the multi-scale features are upsampled to the same size and then merged to obtain the final output, and the tampering localization area of the final depth-repaired video is obtained.
[0052] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention. The structures, devices, and operation methods not specifically described and explained in the present invention are implemented according to the conventional means in the art without special description and limitation.
Claims
1. A video restoration and tampering detection method based on noise residual and average optical flow, characterized in that: The method comprises the following steps: S1: extract the noise and forward and backward optical flow patterns of the video, obtain the noise residual based on the noise, and take the average of the forward and backward optical flow patterns to obtain the average optical flow pattern; S2: The noise residual and the average optical flow pattern are input into the encoder with SegFormer as the backbone network to generate multi-scale features; S3: Use the decoder to perform intra-feature interaction between the generated noise residual multi-scale features and the average optical flow pattern multi-scale features; S4: The interacted features are passed through a multi-scale feature fusion module to perform feature interaction; S5: Merge all features to obtain the final tampering positioning area of the deep repair video.
2. According to claim 1, a video restoration and tampering detection method based on noise residual and average optical flow is characterized in that: The method for extracting the noise residual of the video in S1 includes: First, given a video, convert the video frame into a grayscale image; Using a high-pass filter, convolve the input with a set of filter kernels; Connect the convolution results together as the input of the subsequent network to obtain the noise of the video; Calculate the noise difference of consecutive frames to get the noise residual, which can be expressed as follows: N_R t =F(f t+1 )-F(f t ) (1)。 3. The video restoration and tampering detection method based on noise residual and average optical flow according to claim 1 is characterized in that: The method for extracting the average optical flow pattern of the video in S1 includes: Calculate the brightness component in the direction of the brightness gradient; According to the calculated brightness component, the forward optical flow pattern and the backward optical flow pattern are extracted; The average of the forward and backward optical flow patterns can be expressed as follows:
4. According to claim 2, a video restoration and tampering detection method based on noise residual and average optical flow is characterized in that: The encoder in S2 comprises: 4 stacked SegFormer blocks; Each SegFormer block consists of three parts: effective self-attention, hybrid feedforward network, and overlapping block embedding, which can be expressed as follows: x out =MLP(GELU(Conv(MLP(x in ))))+x in (3b) x out =LN(Flatten(Conv(x in )))(3c)。 5. The video restoration and tampering detection method based on noise residual and average optical flow according to claim 3 is characterized in that: The intra-feature interaction in S3 can be expressed as follows: By fusing each type of features from top to bottom; Concatenate high-level features with low-level features on the channel: f i =Conv([up ×2 (n i+1 )+n i ,up ×2 (m i+1 )+m i ]),i∈{0,1,2}(4a) f i =Conv([n i ,m i ]),i=4(4b)。 6. The video restoration and tampering detection method based on noise residual and average optical flow according to claim 4 is characterized in that: The interaction between features in S4 can be expressed as: The multi-scale feature fusion module is used to upsample the previous layer features by a factor of 2: Use the multiplication operation to eliminate redundant information between two features: Use addition operation to reduce feature information loss: f a =f m +f i ′ +1 +f i (7a) f i ′ +1 =ReLU(f i+1 +BN(Conv(ReLU(BN(Conv(f cc )))))),i∈{0,1,2}(7b) Take advantage of integrating complementary information through concatenation operations: The obtained fusion features are adjusted in dimension and resolution through two convolutional layers, added to the input features, and the output is obtained through the activation layer.
7. The video restoration and tampering detection method based on noise residual and average optical flow according to claim 5 is characterized in that: The feature merging in S5 is performed by upsampling the multi-scale features to the same size and then merging them.