Video Tampering Detection Method and System Based on Multi-Scale Features and Hybrid 3D Network
By using multi-scale features and hybrid 3D network methods in video tamper detection, combined with EMA feature fusion module and hybrid three-dimensional convolution structure, the problem of difficult to identify the frames and complexity generated by video frame interpolation technology in the prior art is solved, and video tamper detection with efficient recognition and generalization capabilities is achieved.
Patent Information
- Application Number
- CN202411576849.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-06
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2044-11-06
AI Technical Summary
The prior art has an adversarial attack that is difficult to understand when detecting frames generated by video frame interpolation technology, and multi-scale feature extraction and hybrid three-dimensional networks are highly complex when processing high-resolution images or videos, which can easily lead to overfitting and gradient vanishing problems.
The video tamper detection method based on multi-scale features and hybrid 3D network is adopted to extract shallow features, temporal features and spatial features through multi-scale internal cascade networks, and feature fusion is used using EMA feature fusion module, combining 3D convolution and parallel grouping convolution structures to improve the efficiency and accuracy of time feature extraction.
Effectively identifying whether video frames have been illegally inserted and tampered, improve the model's ability to recognize objects of different sizes, reduce network parameters and calculation load, enhance the model's performance and generalization capabilities, and resist adversarial attacks.
Smart Images

Figure CN119445345B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a video tampering detection method and system based on multi-scale features and a hybrid 3D network. Background Art
[0002] The statements in this part only provide background technical information related to the present invention and do not necessarily constitute prior art.
[0003] Video frame interpolation technology is a video tampering method for generating intermediate frames in a video sequence, aiming to increase the frame rate of the video, smooth the motion, or enhance the visual effect. Currently, deep video frame interpolation technology has applications in multiple fields, including post-production of movies, video games, virtual reality, and augmented reality, etc., making it more efficient in generating smooth and realistic intermediate frames.
[0004] However, while it has significant advantages in improving video quality and enhancing the visual experience, it may also be misused, bringing some potential risks and hazards. Therefore, it is particularly crucial to develop frame insertion detection technology, which can not only expose the illegal tampering of video content but also provide key technical support for legal procedures and content verification.
[0005] Meanwhile, due to the increasing demand for efficient frame insertion detection technology in society, the frames generated by modern video frame interpolation technology are becoming increasingly difficult to be detected by traditional detection methods. Currently, the methods for detecting the frames generated by video frame interpolation technology mainly include technologies such as multi-scale feature extraction, spatio-temporal analysis, and hybrid three-dimensional networks.
[0006] Multi-scale feature extraction requires determining appropriate scale parameters, which usually involves the selection and adjustment of hyperparameters. When dealing with high-resolution images or videos, efficient algorithms and powerful hardware support are required, which may increase the complexity of the model and lead to an increased risk of overfitting.
[0007] Spatio-temporal data is usually high-dimensional, involving time series and multiple spatial dimensions. Processing and analyzing such data require efficient algorithms to manage the data complexity. In videos or other dynamic scenarios, maintaining temporal consistency is a challenge, especially in cases where the target or scene changes rapidly.
[0008] Regarding the significant issue of the volume of 3D convolutional neural networks (3D CNNs) in practical applications, someone invented a hybrid convolutional network method called MC3. This method combines 3D convolution and 2D convolution for spatio-temporal feature learning. Although the hybrid 3D network combines the advantages of 2D convolution and 3D convolution, its training may be more unstable, especially when using different types of convolutional layers, which may lead to problems such as vanishing or exploding gradients. And motion modeling is required in the early layers of the network structure. In contrast, in higher-level semantic abstractions, motion modeling may be less critical and may even be ignored.
[0009] Video frame insertion technology, that is, using complex technologies such as generative adversarial networks to generate inserted frames that are extremely similar to the original video frames, thus evading detection. This adversarial attack makes it difficult for existing detection methods to effectively identify. Summary of the Invention
[0010] To overcome the deficiencies of the above-mentioned prior art, the present invention provides a video tampering detection method and system based on multi-scale features and a hybrid 3D network. Aiming at the improvement of the hybrid 3D network architecture based on multi-scale feature extraction and the EMA feature fusion module, it can effectively identify whether the video frames have undergone illegal insertion tampering operations.
[0011] To achieve the above object, one or more embodiments of the present invention provide the following technical solutions:
[0012] The first aspect of the present invention provides a video tampering detection method based on multi-scale features and a hybrid 3D network.
[0013] A video tampering detection method based on multi-scale features and a hybrid 3D network includes:
[0014] Obtain the video frames to be detected and perform preprocessing;
[0015] Perform sampling processing on the preprocessed video frames to obtain feature maps of different scales;
[0016] Input the feature maps of different scales into a multi-scale internal cascaded network, and respectively use the shallow feature extraction module, temporal feature extraction module, and spatial feature extraction module in the multi-scale internal cascaded network to extract shallow features, temporal features, and spatial features;
[0017] Fuse the shallow features, temporal features, and spatial features using the EMA fusion module of multi-scale features to obtain the fused features;
[0018] Use the fused features to perform video frame detection to obtain the video tampering detection result;
[0019] Among them, a hybrid three-dimensional convolution structure is introduced in the time feature extraction module, and the time features between multiple adjacent video frames are obtained by paralleling two-dimensional convolution and three-dimensional convolution.
[0020] The second aspect of the present invention provides a video forgery detection system based on multi-scale features and a hybrid 3D network.
[0021] A video forgery detection system based on multi-scale features and a hybrid 3D network includes:
[0022] A data acquisition module, configured to acquire video frames to be detected and perform preprocessing;
[0023] A feature map acquisition module, configured to perform sampling processing on the preprocessed video frames to obtain feature maps of different scales;
[0024] A feature extraction module, configured to input feature maps of different scales into a multi-scale internal cascade network, and respectively use a shallow feature extraction module, a time feature extraction module, and a spatial feature extraction module in the multi-scale internal cascade network to extract shallow features, time features, and spatial features;
[0025] A feature fusion module, configured to fuse the shallow features, time features, and spatial features using an EMA fusion module for multi-scale features to obtain fused features;
[0026] A result detection module, configured to perform video frame detection using the fused features to obtain a video forgery detection result;
[0027] Among them, a hybrid three-dimensional convolution structure is introduced in the time feature extraction module, and the time features between multiple adjacent video frames are obtained by paralleling two-dimensional convolution and three-dimensional convolution.
[0028] The third aspect of the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps in a method as described in the second aspect of the present invention are implemented.
[0029] The fourth aspect of the present invention provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, the steps in a method as described in the second aspect of the present invention are implemented.
[0030] The fifth aspect of the present invention provides a computer program product containing instructions. When it runs on a computer, it causes the computer program to be executed by a processor to implement the steps in a method as described in the second aspect of the present invention.
[0031] The above one or more technical solutions have the following beneficial effects:
[0032] The present invention extracts features of different scales from a video by using a multi-scale feature extraction framework to capture spatio-temporal features in the video. By extracting features at different scales, the model can capture information at different levels in the image, thereby improving the model's recognition ability for objects of different sizes.
[0033] Through the shallow feature extraction module, not only can the most relevant features in multiple scales be effectively extracted from the original dataset for training, but it can also help reduce network parameters to construct a lightweight network.
[0034] Through the temporal feature extraction module, i.e., using 2D convolution in some parts of the network to reduce the computational burden of processing the additional time dimension while still retaining the ability to capture spatial features. It not only maintains computational efficiency but also improves the performance and generalization ability of the model.
[0035] By constructing a spatial feature extraction module, through two depthwise separable convolutions at three scales, the network architecture can be optimized to improve model efficiency. It can reduce the computational load and the number of parameters of the model and learn the temporal features between adjacent frames. It can also capture multi-scale and multi-directional features while maintaining fewer parameters.
[0036] The present invention introduces a hybrid network architecture in the temporal feature extraction module, combining 3D convolution and parallel group convolution to more effectively extract temporal features from video frames. It helps to obtain the temporal features between multiple adjacent frames and can also reduce the computational burden of processing the additional time dimension. 3D convolution can capture spatio-temporal features at a lower level, while 2D convolution can integrate these features at a higher level. Group convolution allows the model to learn different features within different groups, enhancing the model's representation ability for complex spatio-temporal data. By combining with 3D convolution, the network can better capture the dynamic and static features in the video sequence. Due to the addition of the time dimension, this module can more effectively extract temporal features from the video frame sequence.
[0037] The present invention realizes the effective fusion of features of different scales through an improved efficient multi-scale attention module (EMA) technology and conducts final classification to identify forged frames. It enhances the model's understanding and expression ability for complex spatio-temporal features, which is crucial for collecting and integrating cross-space information from different spatial dimensions.
[0038] Advantages of additional aspects of the present invention will be given in part in the following description, will become apparent in part from the following description, or will be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] The accompanying drawings forming a part of this invention are used to provide a further understanding of the invention. The schematic embodiments of the invention and their descriptions are used to explain the invention and do not unduly limit the invention.
[0040] Figure 1 This is the overall framework of the video tampering detection method based on multi-scale features and hybrid 3D network in the first embodiment of the present invention;
[0041] Figure 2 This is the schematic diagram of the original data set in the first embodiment of the present invention;
[0042] Figure 3 This is the schematic diagram of the EMA fusion module structure in the first embodiment of the present invention. Detailed implementation manners
[0043] It should be noted that the following detailed descriptions are all exemplary and are intended to provide further explanations of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0044] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present invention.
[0045] Without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0046] Embodiment 1
[0047] This embodiment discloses a video tampering detection method based on multi-scale features and hybrid 3D network. In order to improve the recognition ability of illegally inserted video frames, a more robust and generalized feature extraction method is developed, and a model that can resist adversarial attacks is designed, which can identify more effectively, such as Figure 1 shown, the specific steps include:
[0048] Step 1: Obtain the video frames to be detected and perform preprocessing;
[0049] Step 2: Perform sampling processing on the preprocessed video frames to obtain feature maps of different scales;
[0050] Step 3: Input the feature maps of different scales into a multi-scale internal cascade network, and respectively use the shallow feature extraction module, temporal feature extraction module and spatial feature extraction module in the multi-scale internal cascade network to extract shallow features, temporal features and spatial features;
[0051] Step 4: Perform feature fusion on the shallow features, temporal features and spatial features by using the EMA fusion module of multi-scale features to obtain the fused features;
[0052] Step 5: Input the fused features into the classification network to determine whether it is a frame-inserted tampered image.
[0053] To describe this embodiment more clearly, a video tampering detection method based on multi-scale features and a hybrid 3D network can be specifically described as follows:
[0054] Step 1: Obtain the video frames to be detected and perform preprocessing.
[0055] In this embodiment, a total of 26,633 videos with an original frame rate of 25 frames per second (fps) from UCF101 and Vimeo90K are selected, and the training and test data are constructed in a ratio of 7:1. A total of 186,431 frames (26,633 videos × 7 frames / video = 186,431 frames) constitute the original dataset.
[0056] Use two representative video frame insertion technology (DVFI) methods, AdaCoF and IFRNet. Take seven consecutive frames as a group, as Figure 2 shown. Therefore, a total of 35,660 groups of frames are used in the training set and 3,760 groups of frames are tested in the test set.
[0057] This rich and diverse dataset provides strong support for our research and enables the comprehensive performance evaluation of the DVFI forensics task.
[0058] In this embodiment, the preprocessing is to randomly crop seven consecutive video frames to obtain 256 × 256 blocks, that is, to obtain the cropped images.
[0059] Step 2: Perform sampling processing on the preprocessed video frames to obtain feature maps of different scales.
[0060] In this embodiment, specifically: perform bilinear upsampling and downsampling by a factor of two on the cropped images in Step 1 to obtain twice-scale images and half-scale images respectively.
[0061] Specifically, it can be expressed in mathematical form as:
[0062] ;
[0063] ;
[0064] where F up represents the feature after bilinear upsampling, F down represents the feature after downsampling by a factor of two, and F si represents the feature map extracted at a specific scale si.
[0065] Step 3: Input the feature maps of different scales into the multi-scale internal cascaded network, and use the shallow feature extraction module, temporal feature extraction module, and spatial feature extraction module in the multi-scale internal cascaded network to extract shallow features, temporal features, and spatial features respectively.
[0066] In this embodiment, the specific steps are as follows:
[0067] Step 3-1: Use the shallow feature extraction module in the multi-scale internal cascaded network to extract shallow features.
[0068] In this embodiment, the shallow feature extraction module includes three groups of convolutions. These features are characterized through training, which can help reduce network parameters and build a lightweight network.
[0069] Among them, the convolution kernel, stride, and padding of each group of convolution blocks are set to 3×3, 1, and 1 respectively. Each group of convolution blocks includes three operations: Batch Normalization, Relu, and group conv, that is, batch normalization, activation function, and grouped convolution; the BN layer is used to reduce overfitting, and ReLU is used as the activation function, and pooling operations are not used.
[0070] Input the cropped image, double-scale image, and half-scale image into the three groups of convolutions respectively for convolution, and obtain the feature maps after the first convolution, the feature maps after the second convolution, and the feature maps after the third convolution.
[0071] Then, perform skip connection on the feature map after the first convolution and the feature map after the three convolutions to obtain more stable shallow features.
[0072] The specific calculation process of this module can be expressed in mathematical form as:
[0073] ;
[0074] In the formula, W0 represents the weight of the 3×3 convolutional layer, represents element-wise addition, and Upsample(·) and Downsample(·) represent upsampling and downsampling respectively.
[0075] By constructing the shallow feature extraction module, it helps to reduce network parameters to build a lightweight network, and can further suppress the content and abnormal noise of video frames; in addition, this module does not use pooling operations, effectively avoiding the weakening of interpolation trajectories; also, by extracting features at different scales, the model can capture different levels of information in the image, thereby improving the model's recognition ability for objects of different sizes.
[0076] Step 3-2: Use the temporal feature extraction module in the multi-scale internal cascaded network to extract temporal features.
[0077] In this embodiment, the temporal feature extraction module introduces a hybrid three-dimensional convolutional network structure, which integrates two grouped convolutions (2D) and 3D convolution in parallel, and obtains the temporal features between multiple adjacent video frames by paralleling the two-dimensional convolution and the three-dimensional convolution.
[0078] The specific process of extracting temporal features is as follows:
[0079] 1. Upsample the half-scale features obtained in step 2 to the same size as the original-scale feature map and then concatenate them; upsample the original-scale features obtained in step 2 to the same size as the double-scale feature map and then concatenate them. Pass the two concatenated groups of feature maps through a point convolution block and a parallel convolution block of 2D and 3D respectively; directly pass the half-scale feature map through the convolution block composed of a point convolution and a parallel convolution of 2D and 3D.
[0080] By using a 1x1 convolution kernel to perform convolution on each position of the feature map, the number of channels of the feature map can be effectively changed.
[0081] 2. Input the features extracted by the point convolution into two grouped convolution modules to obtain the feature tensors of the corresponding scales.
[0082] Grouped convolution allows the model to learn different features within different groups, enhancing the model's representation ability for complex spatio-temporal data. Compared with only using 3D convolution, this method reduces the computational burden of processing the additional time dimension by using 2D convolution in some parts of the network, while still retaining the ability to capture spatial features.
[0083] 3. Then input the features extracted by the point convolution into the 3D convolution block. The 3D convolution can capture spatio-temporal features at a lower level, concatenate the captured feature maps along the time and space dimensions, and integrate these features at a higher level.
[0084] The convolution kernel size of this 3D convolution block is 3×3×3, which has dimensions in both the spatial and time dimensions. Adding an additional depth dimension represents the depth in the time dimension, which can capture the relationship between adjacent video frames and obtain a feature tensor containing temporal information; after the 2D convolution operation, a feature tensor containing spatial information is obtained. The BN layer is used to reduce overfitting, and ReLU is used as the activation function without using pooling operations. The two feature tensors are concatenated along the channel dimension to obtain a feature map with temporal and spatial information.
[0085] The 3D convolution can capture spatio-temporal features at a lower level, while the 2D convolution can integrate these features at a higher level. Grouped convolution allows the model to learn different features within different groups, enhancing the model's representation ability for complex spatio-temporal data.
[0086] The time feature extraction module can be expressed in mathematical form as follows:
[0087] ;
[0088] Wherein, W0 and W1 represent the weights of the 3×3 convolutional layer, and W2 represents the weights of the 3×3×3 convolutional layer.
[0089] By combining 2D and 3D convolutions, constructing a hybrid 3D convolutional network can better capture the dynamic and static features in the video sequence. Due to the addition of the time dimension, this module can more effectively extract time features from the video frame sequence. This method not only maintains computational efficiency but also improves the performance and generalization ability of the model.
[0090] Step 3-2: Use the spatial feature extraction module in the multi-scale internal cascade network to extract spatial features.
[0091] In this embodiment, the enhanced features obtained after the previous two modules are used to estimate the pixel-level parameters of each target pixel on input frames of different scales through three groups of sub-networks.
[0092] The spatial feature extraction module includes three convolutional operations for extracting spatial features between adjacent frames at three scales. That is, for each scale, each convolutional operation includes a point convolution and two depthwise separable convolutions. The point convolution combines the feature maps along the channel direction, mixes the features at each position to aggregate the time features. The depthwise separable convolution is a combination of a depth convolution and a point convolution, which is used to reduce the computational load and the number of parameters of the model and learn the time features between adjacent frames. This enables our model to capture multi-scale and multi-directional features while maintaining fewer parameters.
[0093] The specific steps are as follows:
[0094] 1. Upsample the half-scale features obtained in step 2 to have the same size as the original-scale feature map and then concatenate them; upsample the original-scale features obtained in step 2 to have the same size as the double-scale feature map and then concatenate them. Pass the two sets of concatenated feature maps through a point convolution block and two depthwise separable convolution blocks respectively; directly pass the half-scale feature map through the convolution block composed of one point convolution and two depthwise separable convolutions.
[0095] 2. The double-scale features, half-scale features, and original-scale features obtained from the previous module are first passed through a pointwise convolution respectively, with a stride of 2, a padding of 1, and a kernel size of 3×3, to obtain the multi-scale features after convolution. We use it to combine the feature maps in the channel direction, perform pointwise mixing on the features at each position, and aggregate the temporal features.
[0096] It can mix features in the depth dimension while maintaining the resolution in the spatial dimension.
[0097] Among them, the channel direction refers to the depth of the feature map, that is, the number of channels of the feature map. It contains the feature representations of the input data at different levels. Each channel can be regarded as a filter for capturing different features in the input data.
[0098] 3. The multi-scale features after convolution are passed through two depthwise separable convolution blocks. Depthwise separable convolution is a combination of depthwise convolution and pointwise convolution. Among them, the pointwise convolution, stride, and padding are 3×3, 1, and 0 respectively.
[0099] For depthwise convolution, a convolution kernel is independently applied to each input channel, which keeps the number of channels of the input feature map unchanged, and the spatial features of each channel are independently extracted and encoded. Pointwise convolution can be regarded as a linear transformation in the channel dimension. It combines the features of each channel from depthwise convolution through weighted summation to create new feature channels. A BN layer is used between the two convolutions to reduce overfitting, and ReLU is used as the activation function. After passing through two depthwise separable convolution modules, a feature map with spatial information is obtained.
[0100] This module can be represented in mathematical form as:
[0101] ;
[0102] In the formula, W0 is the weight of the 3×3 depthwise separable convolution layer, and W1 is the weight of the 3×3 convolution layer.
[0103] Step 4: The shallow features, temporal features, and spatial features are fused using the EMA fusion module of multi-scale features to obtain the fused features.
[0104] In this embodiment, the fusion module adopts an improved efficient multi-scale attention module (EMA), adds the EMA module to the end of the overall framework, and does not change the size of the feature vector.
[0105] 1. For any given input feature map X∈R C×H×W , EMA divides X into G sub-features in the channel dimension direction to learn different semantics; the grouping can be expressed as X = [X0, X1,...X G−1, X i ∈ R C / / G×H×W 。
[0106] Two tensors are introduced, one of which is the output of the 1×1 branch and the other is the output of the 5×5 branch. The attention weight descriptor of the grouped feature map is extracted using three parallel paths; to better collect multi-scale spatial information, we change the original 3×3 branch to a 5×5 branch, increasing the capacity of the model.
[0107] Among them, C×H×W is the size of the feature map tensor input to the EMA after a series of previous convolutions, which represents a feature map with C channels, H height, and W width. X represents the feature map itself, and R represents the set of real numbers, meaning the data is real-valued.
[0108] EMA groups the number of channels of the input feature map to achieve the aggregation of information in different spatial dimension directions, thereby extracting richer feature representations. The channel dimension C is divided into G groups, and each group contains C / G channels. Here, G is a hyperparameter representing dividing the number of channels into G equal groups. Therefore, the shape of the grouped feature map becomes (C / G)×H×W.
[0109] The above C / G channels are divided into two parts, R1 and R3, where R1 and R3 are two factors of C / G. We need to reorganize the channel dimension C / G into two new dimensions, R1 and R3, so that this reorganized feature map can be used in subsequent operations to capture spatial features at different scales and enhance the model's perception ability of key information through the attention mechanism.
[0110] 2. Use two-dimensional global average pooling to encode the global spatial information in the output of the 1×1 branch. Before the channel feature joint activation mechanism, directly convert the output of the smallest branch into the shape of the corresponding dimension; that is, R1 1×C / / G × R3 C / / G×HW 。
[0111] The formula for the 2D global pooling operation is:
[0112] ;
[0113] In the formula, Z c represents the output related to the Cth channel. Its purpose is to encode the global information and model the long-term dependencies.
[0114] Step 5: Input the fused features into the classification network to determine whether it is a frame-inserted tampered image.
[0115] In this embodiment, it is specifically as follows: the extracted multi-scale features are input into the classification network, and a multi-layer perceptron composed of multiple fully connected layers is used for binary classification. Finally, it includes an output unit, and the sigmoid activation function is used to output the probabilities belonging to the two categories.
[0116] Embodiment II
[0117] The purpose of this embodiment is to provide a video tampering detection system based on multi-scale features and a hybrid 3D network, including:
[0118] A data acquisition module, configured to acquire video frames to be detected and perform preprocessing;
[0119] A feature map acquisition module, configured to perform sampling processing on the preprocessed video frames to obtain feature maps of different scales;
[0120] A feature extraction module, configured to input feature maps of different scales into a multi-scale internal cascaded network, and respectively use a shallow feature extraction module, a temporal feature extraction module, and a spatial feature extraction module in the multi-scale internal cascaded network to extract shallow features, temporal features, and spatial features;
[0121] A feature fusion module, configured to fuse the shallow features, temporal features, and spatial features by using an EMA fusion module for multi-scale features to obtain the fused features;
[0122] A result detection module, configured to use the fused features to perform video frame detection and obtain video tampering detection results;
[0123] Among them, a hybrid three-dimensional convolution structure is introduced into the temporal feature extraction module, and temporal features between multiple adjacent video frames are obtained by paralleling two-dimensional convolution and three-dimensional convolution.
[0124] Embodiment III
[0125] The purpose of this embodiment is to provide a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the above method are implemented.
[0126] Embodiment IV
[0127] The purpose of this embodiment is to provide a computer-readable storage medium.
[0128] A computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the above method are executed.
[0129] Embodiment V
[0130] The purpose of this embodiment is to provide a computer program product containing instructions, which, when running on a computer, enables the computer to execute the methods and functions involved in any one of the above embodiments.
[0131] Each step involved in the device of the above embodiments corresponds to the first method embodiment. For specific implementation manners, reference may be made to the relevant description part of the first embodiment. The term "computer-readable storage medium" should be understood to include a single medium or multiple media containing one or more instruction sets; it should also be understood to include any medium that can store, encode, or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.
[0132] Those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computer device. Optionally, they can be implemented by program codes executable by a computing device. Thus, they can be stored in a storage device for execution by the computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.
[0133] Although the specific implementation manners of the present invention have been described above in conjunction with the accompanying drawings, this is not a limitation on the protection scope of the present invention. Those skilled in the art should understand that based on the technical solutions of the present invention, various modifications or deformations that can be made without creative efforts by those skilled in the art are still within the protection scope of the present invention.
Claims
1. A video tampering detection method based on multi-scale features and hybrid 3D networks, characterized in that: include: Obtain the video frame to be detected and perform preprocessing; The preprocessed video frames are sampled to obtain feature maps of different scales; Input feature maps of different scales into the multi-scale internal cascade network, and respectively use the shallow feature extraction module, the temporal feature extraction module and the spatial feature extraction module in the multi-scale internal cascade network to extract shallow features, temporal features and spatial features; The shallow features, temporal features and spatial features are fused using an EMA fusion module of multi-scale features to obtain fused features; The EMA fusion module using multi-scale features performs feature fusion, and the specific steps are as follows: For any given input feature map, the EMA fusion module divides X into G sub-features in the channel dimension direction; Two tensors are introduced, one of which is the output of the 1×1 branch and the other is the output of the 5×5 branch. Three parallel paths are used to extract the attention weight descriptor of the grouped feature map. The global spatial information in the 1×1 branch output is encoded using 2D global average pooling, and the output of the smallest branch is directly converted to the shape of the corresponding dimension before the channel feature joint activation mechanism; Use the fused features to perform video frame detection and obtain video tampering detection results; Among them, a hybrid three-dimensional convolution structure is introduced into the temporal feature extraction module, and the temporal features between multiple adjacent video frames are obtained by parallel two-dimensional convolution and three-dimensional convolution.
2. The video tampering detection method based on multi-scale features and hybrid 3D network as claimed in claim 1, characterized in that: The preprocessing is to randomly crop the video frame to obtain a cropped image; the preprocessed video frame is sampled to obtain feature maps of different scales, specifically, the cropped image is upsampled by two times and downsampled by one half to obtain a twice-scale image and a half-scale image, respectively.
3. The video tampering detection method based on multi-scale features and hybrid 3D network as claimed in claim 1, characterized in that: The shallow feature extraction module includes three groups of convolutions. The shallow feature extraction module in the multi-scale internal cascade network is used to input feature maps of different scales into the three groups of convolutions for convolution, and obtain the feature map after the first convolution, the feature map after the second convolution, and the feature map after the third convolution. Then, the feature map after the first convolution is jump-connected with the feature map after the third convolution to obtain shallow features; Among them, the convolution kernel, stride and padding of each group of convolution blocks are set to 3×3, 1 and 1 respectively. Each group of convolution blocks includes three operations: batch normalization, activation function and grouped convolution; BN layer is used to reduce overfitting, and ReLU is used as the activation function without pooling operation.
4. The video tampering detection method based on multi-scale features and hybrid 3D network as claimed in claim 1, characterized in that: The temporal feature extraction module is a hybrid three-dimensional convolutional network structure; the network integrates two group convolutions and 3D convolutions in parallel, and obtains the temporal features between multiple adjacent video frames by parallelizing two-dimensional convolutions and three-dimensional convolutions; The temporal feature extraction module is used to extract the temporal feature, specifically: the spliced feature map is passed through a point convolution block and a 2D and 3D parallel convolution block respectively; the features extracted by the point convolution are obtained; the features extracted by the point convolution are input into two group convolution modules and a 3D convolution module respectively, and two feature tensors of corresponding scales are obtained; the two feature tensors of corresponding scales are spliced along the channel dimension to obtain a feature map with temporal and spatial information.
5. The video tampering detection method based on multi-scale features and hybrid 3D network as claimed in claim 1, characterized in that: The spatial feature extraction module includes three convolution operations for extracting spatial features between adjacent frames at three scales; the three convolution operations include a point convolution block and two depth-separable convolution blocks.
6. A video tampering detection system based on multi-scale features and hybrid 3D networks, characterized in that: include: The data acquisition module is configured to acquire the video frame to be detected and perform preprocessing; A feature map acquisition module is configured to sample the preprocessed video frames to obtain feature maps of different scales; The feature extraction module is configured to input feature maps of different scales into the multi-scale internal cascade network, and respectively extract shallow features, temporal features, and spatial features using the shallow feature extraction module, the temporal feature extraction module, and the spatial feature extraction module in the multi-scale internal cascade network; The feature fusion module is configured to perform feature fusion on the shallow features, temporal features and spatial features using an EMA fusion module of multi-scale features to obtain fused features; the EMA fusion module using multi-scale features performs feature fusion, and the specific steps are: for any given input feature map, the EMA fusion module divides X into G sub-features in the channel dimension direction; introduces two tensors, one of which is the output of the 1×1 branch and the other is the output of the 5×5 branch, and uses three parallel paths to extract the attention weight descriptor of the grouped feature map; uses two-dimensional global average pooling to encode the global spatial information in the 1×1 branch output, and directly converts the output of the minimum branch into the shape of the corresponding dimension before the channel feature joint activation mechanism; The result detection module is configured to use the fused features to perform video frame detection and obtain video tampering detection results; Among them, a hybrid three-dimensional convolution structure is introduced in the time feature extraction module, and the time features between multiple adjacent video frames are obtained by parallel two-dimensional convolution and three-dimensional convolution.
7. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the method described in any one of claims 1 to 5 are implemented.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method described in any one of claims 1 to 5 are performed.
9. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Face-changing video tampering detection method and system based on multi-domain feature fusion
CN112734696A
Video action recognition method based on spatial-temporal enhanced network
WO2023065759A1