A passive evidence collection method for video tampering and its network

By using a passive video tampering evidence network with a Transformer architecture and a feature enhancement module, the problem of difficult detection of tampering traces in video matting under deep learning methods is solved, achieving efficient identification of subtle tampering traces and complete processing of long videos.

CN117237853BActive Publication Date: 2025-10-31HUNAN UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311402080.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-27
Publication Date
2025-10-31
Estimated Expiration
2043-10-27

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively detect and locate tampering traces in video matting techniques based on deep learning methods, resulting in insufficient accuracy and robustness of video tampering detection methods.

Method used

A passive video tampering evidence collection network based on the Transformer architecture is adopted. Through a 2D spatial feature extraction model and a temporal attention encoder, combined with a self-attention mechanism and a feature enhancement module, it captures the temporal relationship and subtle tampering traces between video frames. The loss function is used to optimize the model parameters and generate more effective masks.

Benefits of technology

It improves the accuracy and robustness of passive evidence collection from video tampering, effectively captures tampering traces from image matting techniques using deep learning methods, and is suitable for the complete processing of long videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117237853B_ABST
    Figure CN117237853B_ABST
Patent Text Reader

Abstract

This application provides a passive evidence collection method for video tampering. The input sequence is fed into a time-attention-based encoder that iterates 12 times to complete the encoding process. Edge extraction is performed on the results of the 3rd, 6th, 9th, and 12th iterations, and the operation edge extracted from the 12th iteration is used as the forgery feature F. for The operation edges extracted from the results of the 3rd, 6th, 9th, and 12th iterations are fused together and used as the edge feature F. edge ; to forge feature F for After enhancing the edge location in the spatial domain, it is compared with the edge feature F. edge Feature fusion is performed to obtain enhanced feature F enh For the enhancement feature F enh Decoding to obtain the final mask of the operation area can effectively capture videos tampered with using deep learning-based matting techniques, improving the accuracy of passive video tampering evidence collection. This application also provides a passive video tampering evidence collection network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of image processing technology, specifically relating to a passive evidence collection method and network for video tampering. Background Technology

[0002] Passive video tampering evidence collection is a technique used to detect whether specific targets or regions have been replaced or modified in videos using video matting techniques, thereby verifying the authenticity and integrity of the video. It is widely used to extract key information from video data such as surveillance videos and camera recordings, to identify criminal suspects, vehicles, and objects, or to analyze accident scenes and behavioral patterns. With the development of deep learning technology, video matting techniques have undergone revolutionary improvements. Deep learning methods utilize deep neural networks to learn features and representations from data, enabling more effective capture of target object features, improving the accuracy and robustness of matting. This allows video matting technology to better handle blurred boundaries, occlusions, and moving objects in complex scenes, making tampering traces less noticeable to the naked eye. As a result, existing forensic methods cannot capture these traces and are unable to effectively detect and locate tampering.

[0003] Therefore, it is necessary to provide a passive evidence collection method and network for video tampering to solve the problems mentioned in the background art. Summary of the Invention

[0004] This application provides a passive evidence collection method and network for video tampering. Based on deep learning methods, it can detect more subtle tampering traces and effectively capture videos tampered with using deep learning-based image matting techniques, thereby improving the accuracy of passive evidence collection from video tampering.

[0005] To solve the above-mentioned technical problems, the technical solution of this application is as follows:

[0006] A passive evidence collection method for video tampering is provided, including the following steps:

[0007] S1: Provide continuous RGB video frames containing tampered content, extract features using a 2D spatial feature extraction model, and use the extracted features as the input sequence;

[0008] S2: Construct a passive video tampering evidence collection network. The input sequence is fed into the network for training. The training process involves feeding the input sequence into a time-attention-based encoder for 12 iterations to complete the encoding process. Edge extraction is performed on the results of the 3rd, 6th, 9th, and 12th iterations. The operation edge extracted from the 12th iteration is used as the forgery feature F. for The operation edges extracted from the results of the 3rd, 6th, 9th, and 12th iterations are fused together and used as the edge feature F. edge ; to forge feature Ffor After enhancing the edge location in the spatial domain, it is compared with the edge feature F. edge Feature fusion is performed to obtain enhanced feature F enh For the enhancement feature F enh Decode the final mask of the operation area, construct a loss function between the final mask and the real mask, train the network with the goal of minimizing the loss function, output the optimal model parameters, and complete the training of the passive evidence collection network for video tampering.

[0009] S3: For passive evidence collection of tampering of any consecutive RGB video frames, the trained passive evidence collection network for video tampering is used to make predictions and output the prediction mask of the consecutive RGB video frames to be detected.

[0010] Preferably, in step S2, the encoder uses a Transformer architecture.

[0011] Preferably, the encoder is composed of multiple encoder layers with the same structure stacked together. Each encoder layer includes a self-attention layer and a feedforward neural network sublayer connected in sequence. The self-attention layer interacts with each element in the input sequence with other elements in the input sequence, calculates the relative weights with other elements, and generates a weighted representation of each element, capturing the dependencies and contextual information between elements. The feedforward neural network sublayer performs a nonlinear transformation on the representation generated by the self-attention layer to enhance the expressive power of the representation. By stacking multiple encoder layers, the input sequence is encoded into a richer and more abstract representation.

[0012] Preferably, the edge extraction process specifically includes the following steps:

[0013] The output features after the input features are convolved by a 1×1 convolution are divided into two branches. One branch is convolved again by two concatenated first convolutional blocks. The two first convolutional blocks have the same structure, each including a 5×5 convolutional layer, a batch normalization layer and a Leaky ReLU activation function layer connected in sequence. The other branch is left unprocessed. The output features of the two branches are then fused.

[0014] The fused features are convolved through a second convolutional block to output the extracted operation edges. The second convolutional block includes a 5×5 convolutional layer, a batch normalization layer, and a ReLU activation function layer connected in sequence.

[0015] Preferably, in step S2, the forged feature F is... for After enhancing the edge location in the spatial domain, it is compared with the edge feature F. edge Feature fusion is performed to obtain enhanced feature F enh Specifically, the process includes the following steps:

[0016] In the forgery feature F forand edge features F edge Element-wise multiplication between them yields the convolution tensor S∈R. C ×(H / 16)×(W / 16) Then, the dimension of the convolution tensor S is reshaped to [(H / 16)×(W / 16)]×C, resulting in three convolution tensors S1, S2 and S3.

[0017] In convolution tensor S3 and Perform matrix multiplication between them, and then generate matrix A∈R through a softmax operation. (HW / 16)×(HW / 16) , represented as: Where T represents the matrix transpose; Represents matrix multiplication; softmax(·) represents the softmax operation;

[0018] Perform a matrix multiplication between matrix A and the convolutional tensor S1, and reshape the result of the matrix multiplication to match the forged feature F. for Same dimension, and with fake feature F for Perform matrix summation, and then sum the results to the edge features F. edge After cascading, the enhanced feature Fenh∈R was obtained. 2C×(H / 16)×(W / 16) , represented as: In the formula, reshape indicates reshaping using the reshape function; This represents matrix summation.

[0019] Preferably, the decoding process enhances feature F enh Perform bilinear upsampling and 1×1 convolution operations to transform the fake features F for and edge features F edge Restore to the original resolution.

[0020] Preferably, the loss function during training is expressed as:

[0021] L total =L for +L edge ;

[0022] In the formula, L total Indicates total loss; L for Indicates forged loss; L edge L represents the edge loss. for and L edge All are defined by the binary cross-entropy function.

[0023] This application also provides a passive video tampering evidence collection network for performing the above-described passive video tampering evidence collection method, including:

[0024] Encoder: Used to iterate the input sequence 12 times to complete the encoding process, wherein the input sequence is the feature extracted from continuous RGB video frames by a 2D spatial feature extraction model;

[0025] Edge extraction module: Used to extract edges from the results of the 3rd, 6th, 9th, and 12th iterations, and the operation edge extracted from the 12th iteration is used as the forgery feature F. for The operation edges extracted from the results of the 3rd, 6th, 9th, and 12th iterations are fused together and used as the edge feature F. edge ;

[0026] Feature enhancement module: used to enhance the appearance of forged features F for After enhancing the edge location in the spatial domain, it is compared with the edge feature F. edge Feature fusion is performed to obtain enhanced feature F enh ;

[0027] Decoder: Used for enhancing feature F enh Decode the final mask of the operating region.

[0028] The beneficial effects of this application are as follows:

[0029] (1) The Transformer architecture is used. By inputting the features of each frame into the Transformer architecture, the self-attention mechanism of the Transformer can be used to capture the temporal relationship and contextual information between different frames, and richer and more abstract features can be obtained.

[0030] (2) In order to capture rich tampering traces, shallow and deep layers are used to capture forged features and edge features. A self-attention-based feature enhancement module is used to enhance the artifacts in the edge region and generate more effective enhanced features. This can improve the effectiveness and accuracy of identifying subtle tampering traces, enabling it to effectively capture videos tampered with by matting techniques based on deep learning methods.

[0031] (3) Allows full video processing during testing, making it more suitable for processing long videos. Attached Figure Description

[0032] Figure 1 This is a framework diagram illustrating the passive evidence collection network for video tampering provided in this application;

[0033] Figure 2 This represents the transformer architecture diagram;

[0034] Figure 3 This diagram illustrates the architecture of the edge extraction module.

[0035] Figure 4An architecture diagram showing the feature enhancement module;

[0036] Figure 5 This chart shows a comparison of prediction results from various models on the VideoMatte240K dataset. Detailed Implementation

[0037] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0038] Please see Figure 1-4 This application provides a passive evidence collection method for video tampering, comprising the following steps:

[0039] S1: Provides continuous RGB video frames containing tampered content, extracts features using a 2D spatial feature extraction model, and uses the extracted features as the input sequence.

[0040] The 2D spatial feature extraction model adopts a conventional structure in this field, and its extraction process also adopts conventional techniques in this field, which will not be elaborated in this embodiment.

[0041] S2: Construct a passive video tampering evidence collection network. The input sequence is fed into the network for training. The training process involves feeding the input sequence into a time-attention-based encoder for 12 iterations to complete the encoding process. Edge extraction is performed on the results of the 3rd, 6th, 9th, and 12th iterations. The operation edge extracted from the 12th iteration is used as the forgery feature F. for The operation edges extracted from the results of the 3rd, 6th, 9th, and 12th iterations are fused together and used as the edge feature F. edge ; to forge feature F for After enhancing the edge location in the spatial domain, it is compared with the edge feature F. edge Feature fusion is performed to obtain enhanced feature F enh For the enhancement feature F enh Decode the final mask of the operation area, construct a loss function between the final mask and the real mask, train the model with the goal of minimizing the loss function, output the optimal model parameters, and complete the training of the passive evidence collection network for video tampering.

[0042] The encoder uses the Transformer architecture. Video data is typically a sequence of frames, each with spatial and channel dimensions. To design and optimize for these specific dimensions and better capture the relationships within the video data, the Transformer structure needs to be extended or modified to accommodate the temporal and spatial dimensions of the video data. Such extensions and modifications can be performed using conventional techniques in the field.

[0043] The encoder consists of multiple stacked encoder layers with the same structure. Each encoder layer includes a self-attention layer and a feedforward neural network sublayer connected in sequence. The self-attention layer interacts with each element in the input sequence with other elements in the input sequence, calculates the relative weights with other elements, and generates a weighted representation of each element, capturing the dependencies and contextual information between elements. The feedforward neural network sublayer performs a nonlinear transformation on the representation generated by the self-attention layer to enhance the expressive power of the representation. Through the stacking of multiple encoder layers, the input sequence is encoded into a richer and more abstract representation.

[0044] In existing technologies, detection typically involves analyzing the temporal information and contextual consistency of video frames to detect the presence of matting operations and whether the matting behavior is consistent with the surrounding environment. However, in complex scenes, temporal information and contextual consistency can be affected by various factors, such as lighting changes, dynamic backgrounds, and occlusions, which may make it difficult for detection methods to accurately analyze temporal relationships and contextual information. Therefore, this application uses a Transformer architecture. By inputting the features of each frame into the Transformer architecture, the self-attention mechanism of the Transformer can be used to capture the temporal relationships and contextual information between different frames, improving the accuracy of obtaining temporal relationships and contextual information.

[0045] For video tampering detection tasks, the traces of tampering are very subtle. Besides the traces in the manipulated area, the subtle differences between the manipulated edges and the surrounding non-manipulated areas are also crucial. Therefore, this application obtains the forgery feature F through edge extraction. for and edge features F edge Furthermore, a self-attention-based feature enhancement module is used to enhance artifacts in edge regions, thereby generating more effective enhanced features F. enh .

[0046] The edge extraction process specifically includes the following steps:

[0047] The output features after the input features are convolved by a 1×1 convolution are divided into two branches. One branch is convolved again by two concatenated first convolutional blocks. The two first convolutional blocks have the same structure, each including a 5×5 convolutional layer, a batch normalization layer and a Leaky ReLU activation function layer connected in sequence. The other branch is left unprocessed. The output features of the two branches are then fused.

[0048] The fused features are convolved through a second convolutional block to output the extracted operation edges. The second convolutional block includes a 5×5 convolutional layer, a batch normalization layer, and a ReLU activation function layer connected in sequence.

[0049] In addition to the area within the tampered region, the operational edges also contain abundant forgery traces. Considering the importance of the operational edges, this invention enhances the forgery features F in the spatial domain before connection. for The operation edge position is made easier to capture, thereby improving the effectiveness and accuracy of tamper detection. Specifically, the forgery feature F for After enhancing the edge location in the spatial domain, it is compared with the edge feature F. edge Feature fusion specifically includes the following process:

[0050] In the forgery feature F for and edge features F edge Element-wise multiplication between them yields the convolution tensor S∈R. C ×(H / 16)×(W / 16) Then, the dimension of the convolution tensor S is reshaped to [(H / 16)×(W / 16)]×C, resulting in three convolution tensors S1, S2 and S3.

[0051] In convolution tensor S3 and Perform matrix multiplication between them, and then generate matrix A∈R through a softmax operation. (HW / 16)×(HW / 16) , represented as: Where T represents the matrix transpose; Represents matrix multiplication; softmax(·) represents the softmax operation;

[0052] Perform a matrix multiplication between matrix A and the convolutional tensor S1, reshape the result of the matrix multiplication to the same dimension as the fake feature Ffor, and perform a matrix summation with the fake feature Ffor. Then, sum the result with the edge feature F. edge After cascading, the enhanced feature Fenh∈R was obtained. 2C×(H / 16)×(W / 16) , represented as: In the formula, reshape indicates reshaping using the reshape function; This represents matrix summation.

[0053] The decoding process enhances feature Fenh Perform bilinear upsampling and 1×1 convolution operations to transform the fake features F for and edge features F edge Restore to the original resolution.

[0054] The loss function during training is expressed as:

[0055] L total =L for +L edge ;

[0056] In the formula, L total Indicates total loss; L for Indicates forged loss; L edge L represents the edge loss. for and L edge All are defined by the binary cross-entropy function.

[0057] S3: For passive evidence collection of tampering of any consecutive RGB video frames, the trained passive evidence collection network for video tampering is used to make predictions and output the prediction mask of the consecutive RGB video frames to be detected.

[0058] This application also provides a passive video tampering evidence collection network for performing the above-described passive video tampering evidence collection method, including:

[0059] Encoder: Used to iterate the input sequence 12 times to complete the encoding process, wherein the input sequence is the feature extracted from continuous RGB video frames by a 2D spatial feature extraction model;

[0060] Edge extraction module: Used to extract edges from the results of the 3rd, 6th, and 9th iterations, and merge the extracted features to form the fake feature F. for Edge extraction is performed on the results of the 12th iteration to form edge features F. edge ;

[0061] Feature enhancement module: used to enhance the appearance of forged features F for After enhancing the edge location in the spatial domain, it is compared with the edge feature F. edge Feature fusion is performed to obtain enhanced feature F enh ,

[0062] Decoder: Used for enhancing feature F enh Decode the final mask of the operating region.

[0063] The original video data of the input sequence used in this application for training is long video, which realizes the entire video analysis and allows for complete video processing during testing, making it more suitable for processing long videos.

[0064] Example 1

[0065] Please see Figure 5 We performed character matting and background replacement on the existing large-scale video dataset VideoMatte240K to generate a modified video set. Fifty 4K videos from VideoMatte240K were selected; this dataset is rich in various motion types and moving objects, with each video lasting 15 to 30 seconds and including manually annotated object segmentation masks. Complete, unmodified videos were then fed into the modified dataset as the original videos for training and testing, providing a more accurate prediction mask as a reference.

[0066] Forty videos from the tampered video set were used for training, and ten videos were used for testing. The video rows were horizontally and vertically flipped to enhance dataset diversity. This invention uses F1 and IoU evaluation metrics to comprehensively evaluate the performance of video tamper localization. F1 and IoU are defined as follows:

[0067]

[0068]

[0069] In the formula, TP represents the number of samples that the model correctly predicts as positive; FN represents the number of samples that the model incorrectly predicts as negative; FP represents the number of samples that the model incorrectly predicts as positive; and TN represents the number of samples that the model correctly predicts as negative.

[0070] To evaluate the performance of the passive forensic method and system for detecting frequent tampering in this application, the following comparative experiment was conducted: The comparative models were RRUNet and MantraNet; the test datasets selected for the comparative experiment were CASIAV1.0, CASIAV2.0, and VideoMatte240K. In CASIA, the image resolution was 384×256; CASIAv2.0 contained 5123 images with resolutions ranging from 240×600 to 800×600; CASIAv1.0 and CASIAv2.0 included splicing and copying moving forgeries. For all datasets, 80% of the images were randomly selected as the training set, 10% as the validation set, and 10% as the test set. The performance metrics of each model obtained from the tests are shown in Table 1.

[0071] Table 1 Performance Indicators of Each Model

[0072]

[0073] As shown in Table 1, in the tests on the three test datasets CASIA V1.0, CASIA V2.0 and VideoMatte240K, this application outperforms the Mantra-Net model and the RRUNet model in both F1 and IoU evaluation metrics, indicating that this application has good performance in locating tampered regions.

[0074] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A method for passively obtaining evidence of video tampering, characterized in that, Includes the following steps: S1: Provide continuous RGB video frames containing tampered content, extract features using a 2D spatial feature extraction model, and use the extracted features as the input sequence; S2: Construct a passive video tampering evidence collection network. The input sequence is fed into the network for training. The training process involves feeding the input sequence into a time-attention-based encoder for 12 iterations to complete the encoding process. Edge extraction is performed on the results of the 3rd, 6th, 9th, and 12th iterations. The operation edge extracted from the 12th iteration is used as the forgery feature F. for The operation edges extracted from the results of the 3rd, 6th, 9th, and 12th iterations are fused together and used as the edge feature F. edge ; to forge feature F for After enhancing the edge location in the spatial domain, it is compared with the edge feature F. edge Feature fusion is performed to obtain enhanced feature F enh For the enhancement feature F enh Decode the final mask of the operation area, construct a loss function between the final mask and the real mask, train the network with the goal of minimizing the loss function, output the optimal model parameters, and complete the training of the passive evidence collection network for video tampering. S3: For passive evidence collection of tampering of any consecutive RGB video frames, the trained passive evidence collection network for video tampering is used to make predictions and output the prediction mask of the consecutive RGB video frames to be detected. In step S2, the encoder uses the Transformer architecture; The encoder consists of multiple stacked encoder layers with the same structure. Each encoder layer includes a self-attention layer and a feedforward neural network sublayer connected in sequence. The self-attention layer interacts with each element in the input sequence with other elements in the input sequence, calculates the relative weights with other elements, and generates a weighted representation of each element, capturing the dependencies and contextual information between elements. The feedforward neural network sublayer performs a nonlinear transformation on the representation generated by the self-attention layer to enhance the expressive power of the representation. Through the stacking of multiple encoder layers, the input sequence is encoded into a richer and more abstract representation. The edge extraction process specifically includes the following steps: The output features after the input features are convolved by a 1×1 convolution are divided into two branches. One branch is convolved again by two concatenated first convolutional blocks. The two first convolutional blocks have the same structure, each including a 5×5 convolutional layer, a batch normalization layer and a Leaky ReLU activation function layer connected in sequence. The other branch is left unprocessed. The output features of the two branches are then fused. The fused features are convolved through a second convolutional block to output the extracted operation edges. The second convolutional block includes a 5×5 convolutional layer, a batch normalization layer, and a ReLU activation function layer connected in sequence.

2. The passive evidence collection method for video tampering according to claim 1, characterized in that, In step S2, the forged feature F is... for After enhancing the edge location in the spatial domain, it is compared with the edge feature F. edge Feature fusion is performed to obtain enhanced feature F enh Specifically, the process includes the following steps: In the forgery feature F for and edge features F edge Element-wise multiplication is performed between them to obtain the convolution tensor. Then convolution tensors Dimensional reshaping This yields three convolutional tensors. , and ; In convolution tensor and Perform matrix multiplication between them, and then generate a matrix using the softmax operation. , is represented as: Where T denotes matrix transpose; Represents matrix multiplication; This indicates the softmax operation; In matrix A and convolution tensor Perform matrix multiplication between them, and reshape the result of the matrix multiplication to match the forged feature F. for The same dimensions are used, and the matrix is ​​summed with the forged feature Ffor. The summation result is then compared with the edge feature F. edge Enhanced features were obtained after cascading. , is represented as: In the formula, reshape indicates reshaping using the reshape function; This represents matrix summation.

3. The passive evidence collection method for video tampering according to claim 2, characterized in that, The decoding process enhances feature F enh Perform bilinear upsampling and 1×1 convolution operations to transform the fake features F for and edge features F edge Restore to the original resolution.

4. The passive evidence collection method for video tampering according to claim 1, characterized in that, The loss function during training is expressed as: ; In the formula, L total Indicates total loss; L for Indicates forged loss; L edge L represents the edge loss. for and L edge All are defined by the binary cross-entropy function.

5. A passive video tampering evidence collection network, used to execute the passive video tampering evidence collection method according to any one of claims 1-4, characterized in that, include: Encoder: Used to iterate the input sequence 12 times to complete the encoding process, wherein the input sequence is the feature extracted from continuous RGB video frames by a 2D spatial feature extraction model; Edge extraction module: Used to extract edges from the results of the 3rd, 6th, 9th, and 12th iterations, and the operation edge extracted from the 12th iteration is used as the forgery feature F. for The operation edges extracted from the results of the 3rd, 6th, 9th, and 12th iterations are fused together and used as the edge feature F. edge ; Feature enhancement module: used to enhance the appearance of forged features F for After enhancing the edge location in the spatial domain, it is compared with the edge feature F. edge Feature fusion is performed to obtain enhanced feature F enh ; Decoder: Used for enhancing feature F enh Decode the final mask of the operating region.

Citation Information

Patent Citations

  • Depth video object restoration tampering detection method

    CN111814543A

  • Depth video frame insertion detection method and device and computer readable storage medium

    CN115909160A