ViT and spatial feature fused depth video forgery detection method
Through the deep video forgery detection method that integrates VisionTransformer (ViT) and spatial characteristics, the problem of the existing technology being difficult to accurately detect high-quality face forgery videos is solved, and higher detection accuracy and ability to respond to a variety of forgery technical means is achieved.
Patent Information
- Application Number
- CN202510198525.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-23
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-02-23
AI Technical Summary
The existing deep video forgery detection technology is difficult to accurately detect high-quality face forgery videos in complex scenarios, and it faces the challenge of a variety of forgery technical means.
A deep video forgery detection method that integrates VisionTransformer (ViT) and spatial features is adopted. By constructing a spatial feature extraction network and a detection network that integrates ViT and spatial features, combining the multi-head self-attention mechanism and channel attention mechanism, the detection accuracy is improved.
Without increasing the calculation cost, the accuracy of deep video forgery detection is significantly improved, and it can effectively deal with a variety of forgery techniques, which enhances the generalization ability and robustness of the detection model.
Smart Images

Figure CN120126053A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine vision and forgery detection, and particularly to a deep video forgery detection method that integrates ViT (Vision Transformer) and spatial features. Background Art
[0002] In recent years, artificial intelligence technology has developed rapidly and has been applied more and more widely. In video generation, using generative adversarial network (GAN) technology, realistic images and videos can be forged and generated. However, deep forgery technology is a double-edged sword. While bringing convenience to aspects such as film and television video generation, industrial content production, old photo restoration, and videoization, it also poses severe challenges to the supervision and management of the rationality and authenticity of video content, as well as the identification of forgeries. With the development of the global Internet, the rapid spread and influence of false videos have become increasingly prominent. Especially in high-spread environments such as social media and news platforms, deep forgery videos are extremely likely to have a significant negative impact on public perception and social opinion.
[0003] Among various deep video forgery technologies, the most widely applied, most influential, and most advanced one is deep video face forgery technology, which is divided into three branches: face replacement, attribute editing, and face generation. With the high integration and modularization of technology, the technical threshold of face replacement has gradually decreased, the acquisition method has become simpler, and the affected fields have become more extensive. Currently, the research and application of deep video forgery detection have become an important topic that the society and academia urgently need to solve.
[0004] Although current face forgery detection technologies based on artificial intelligence have made certain progress. However, the generation technology of deep video forgery is constantly updated and iterated, making the detection technology still face many challenges. On the one hand, the quality of the generated face forgery videos is getting higher and higher, and it is difficult to distinguish the forgery details through traditional visual detection means; on the other hand, the videos generated by different forgery means show complex and diverse changes in spatio-temporal features, limiting the generalization ability and robustness of the detection model. Therefore, how to accurately detect face forgery videos in complex scenarios and environments has become an urgent technical problem to be solved. Considering the above problems, a deep video forgery detection method that integrates ViT and spatial features is specifically proposed, which will not only greatly improve the accuracy of deep forgery detection but also be able to cope with various forgery technical means. Summary of the Invention
[0005] The purpose of the present invention is to provide a deep video forgery detection modeling method that integrates ViT and spatial features, which can improve the accuracy of deep forgery video detection without increasing the computational cost, and at the same time can cope with various forgery technical means.
[0006] The technical solution adopted by the present invention is a deep video forgery detection method that fuses ViT and spatial features, specifically including the following steps:
[0007] Step 1: Construct a spatial feature extraction network
[0008] The spatial feature extraction network consists of a spatial inconsistency module and a channel attention mechanism module, mainly including convolutional layer 1, network module 1, convolutional layer 2, global average pooling layer, and convolutional layer 3. Each global extraction convolutional layer is followed by a normalization layer and a non-linear activation layer.
[0009] Furthermore, the module network plays the role of extracting features and reducing the number of parameters. In the spatial inconsistency module, a 1×3 and a 3×1 convolutional kernel are used for feature extraction operations, and 2 1×1 convolutional kernels and 2 3×3 convolutional kernels are used for detailed feature extraction. At the same time, through the residual idea, they are connected by skip connections. Average pooling and bilinear interpolation operations are added before and after the 1×3 and 3×1 convolutions for dimensionality reduction and upsampling processing to reduce the computational amount. For the splitting caused by the face edge information in the picture, convolutional operations are performed in both the horizontal and vertical directions, and downsampling operations are used to enhance the relevant information in the receptive field. A 3×3 convolutional kernel is used for more detailed feature extraction. Then the original features are retained through the residual connection method. Through this operation, problems such as the degradation of the performance of the convolutional layer can be effectively prevented.
[0010] Furthermore, the feature map is introduced into the 3×3 convolutional layer to further extract and fuse the three-way feature information to ensure the fitting performance of the network.
[0011] Furthermore, a lightweight channel attention mechanism is introduced to perform feature extraction and selection in the channel dimension, and to additionally model the importance of each channel. Global pooling operation is used to change the h and w dimensions into one-dimensional scalars, so as to reduce the computational redundancy brought about by performing convolutional operations in the channel dimension. This enables the network model to strengthen the utilization of channel information while reducing the computational amount and improving the fitting ability of the network model.
[0012] Step 2: Construct a detection network that fuses ViT and spatial features
[0013] Fuse ViT and the designed spatial feature extraction network. After ViT extracts the video frame, slicing operation is performed. The 224×224 pixel frame is split into 16 14×14 pixel patches. After flattening, the channel dimension is 3, and 1D position encoding information is used to encode the 16 pictures in sequence.
[0014] Furthermore, perform embedding position encoding processing. Through linear mapping operation, it is mapped into 16 tokens with a channel dimension of 128, and the 0th token is added to incorporate the position encoding.
[0015] Furthermore, in the Transformer Encoder, a spatial feature extraction network structure is added to the Multi-Head Attention. While the multi-head attention mechanism is processing, spatial feature information is extracted in parallel. The extracted feature weights and channel weights are weighted into the query, key, and value matrices for more detailed feature extraction.
[0016] Step 3: Train and construct a fused ViT and spatial feature detection network model
[0017] The specific steps for training the fused ViT and spatial feature detection network model are as follows: Use the Dlib face extractor to perform face extraction operations on video frames, save the face sampling point information, and crop the face images of the extracted frames; pre-train the network model on the forged video dataset; divide the dataset into a training set, a validation set, and a test set and perform rotation and scaling processing; use the validation set to adjust the hyperparameters, and finally test the model effect through the test set.
[0018] Furthermore, the steps for dividing the forged video dataset into a training set, a validation set, and a test set and performing normalization processing are as follows: Divide the dataset into a training set, a validation set, and a test set in a ratio of 3:1:1.
[0019] Furthermore, in order to expand the dataset and prevent overfitting, image enhancement methods such as random rotation and random cropping are used to perform data augmentation on the training set images. Further, when training the network model, criterion is selected as the loss function to calculate the loss between the model output outputs and the true labels labels. Adam is selected as the optimizer supervised by binary cross-entropy loss for optimization. After multiple rounds of training and verification, it is found that overfitting is likely to occur after 30 epochs. Therefore, the training epoch is set to 30, the initial learning rate is 0.0002, and it decays by 10 times every 10 training epochs.
[0020] Step 4: Input the processed video data into the trained fused ViT and spatial feature detection network. The output of the network model is a probability, which shows the accuracy of the final result in correctly judging the authenticity of the video.
[0021] The present invention has the following advantages compared with the prior art:
[0022] By abandoning the multi-layer convolutional layer operations of CNN and introducing the multi-head self-attention mechanism of ViT, the problems of high hardware pressure and excessive computational complexity caused by overly deep feature extraction levels are effectively solved.
[0023] When using the spatial feature extraction network, splitting the convolutional kernel along the orthogonal direction can effectively reduce the number of model parameters while maintaining the model's extraction ability. By integrating ideas such as channel attention mechanism and residual connection, the accuracy and stability of the network model are also ensured. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 is a schematic flowchart of the deep video forgery detection method that integrates ViT and spatial features provided by the present invention;
[0025] Figure 2 is a schematic structural diagram of the spatial feature extraction network of the present invention;
[0026] Figure 3 is a schematic structural diagram of the deep video forgery detection network that integrates ViT and spatial features of the present invention; DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] The present invention mainly realizes a deep video forgery detection method that integrates ViT and spatial features. The specific method adopted by the present invention will be described in detail below with reference to the accompanying drawings.
[0028] Specifically, the process of a deep video forgery detection method that integrates ViT and spatial features is as Figure 1 shown, including the following steps: S1: Construct a spatial feature extraction network. S2: Construct a detection network that integrates ViT and spatial features. S3: Train the constructed detection network model that integrates ViT and spatial features. S4: Input the processed video data into the trained detection network that integrates ViT and spatial features to determine whether the video is forged.
[0029] For S1: Construct a spatial feature extraction network.
[0030] In the present invention, the network structure design of the spatial feature extraction network is shown in Table 1. The spatial feature extraction network consists of a spatial inconsistency module and a channel attention mechanism module, and mainly includes convolutional layer 1, network module 1, convolutional layer 2, global average pooling layer, and convolutional layer 3. Each global extraction convolutional layer is followed by a normalization layer and a non-linear activation layer.
[0031] Convolutional layer 1: The input layer of the spatial feature extraction network uses a convolutional kernel with a size of 5×5, a stride of 1, and edge padding is added. The output is 256 channels. Its purpose is to perform preliminary detail extraction operations on the input data and retain the picture detail information at a low dimension.
[0032] Network module: The structure of the network module is as Figure 2As shown, it consists of a three-way connection network. The upper path is a skip connection structure; the middle path consists of an average pooling layer, a 1×3 convolution, a 3×1 convolution, and a bilinear interpolation operation in sequence; the lower path is composed of a 1×1 convolution layer, a 3×3 convolution layer, a 3×3 convolution layer, and a 1×1 convolution layer. In the middle path, the average pooling layer is a convolution kernel with a stride of 2. It mainly converts the low-dimensional feature map into a high-dimensional feature map through downsampling operations, providing a larger receptive field, retaining the main information, and ignoring the detailed information. Then, 1×3 and 3×1 convolution operations are performed with a stride of 1. The purpose is to extract texture features from two orthogonal directions, horizontal and vertical, to detect whether there are inconsistencies in the image edges. After that, bilinear interpolation is used for upsampling operations to restore the dimension of the feature map while retaining the weight information. In the lower path, the first 1×1 convolution still performs a dimension expansion operation, further expanding the input data channels. In this invention, it is expanded by 2 times, so that the subsequent feature extraction process can be more in-depth. The following two 3×3 convolution operations are for detailed feature extraction. Since the h and w of the feature map do not change, the detailed features of the feature map can be extracted, and at the same time, the number of channels is reduced, discarding some feature channels to prevent the model from overfitting and maintaining the stability of the model. Finally, another 1×1 convolution operation is performed to perform over-dimension reduction again to keep the output dimension consistent with the input, saving network parameters and facilitating subsequent operations. The upper path is a skip connection residual structure to prevent network degradation, enabling the network to be designed deeper and obtaining stronger fitting ability. The feature maps of the upper path and the middle path are added element by element, and then input into the sigmoid function for normalization to obtain the confidence of the feature map as the attention weight, and then multiplied element by element with the lower path feature map matrix to retain the main feature information.
[0033] Convolution layer 2: Convolve the above-mentioned feature map with a 3×3 convolution kernel with a stride of 1, mainly used to further extract the details of the extracted feature map and fuse the three-way information at the same time.
[0034] Global average pooling layer: The kernel size of the global average pooling layer is 14×14. Take the average operation on the 14×14 resolution matrix output by the front-end network to reduce the dimension to a 1×1 scalar. Then it is concatenated into a new two-dimensional matrix, that is, the two dimensions of the scalar of each layer and the number of channels.
[0035] Convolution layer 3: Extract features in the channel dimension through convolution with a convolution kernel of size 1×3. Since the previous operations are independent operations on the feature maps of each channel layer and do not operate on the correlation between channel layers, channel feature extraction is performed. Then, it is normalized through the sigmoid function and then multiplied element by element with the channel matrix to obtain the weight of the channel attention mechanism as the output of the final feature matrix.
[0036] Table 1 Spatial Feature Extraction Network Structure
[0037] Network layer Convolution kernel size Input channels Output channels Stride Convolution layer 1 <![CDATA[5 2 ,256]]> 256 256 1 Network module - Average pooling layer <![CDATA[2 2 ,512]]> 256 512 2 Network module - Convolution kernel 1 1×3,512 512 512 1 Network module - Convolution kernel 2 3×1,256 512 256 1 Network module - Bilinear interpolation - 256 256 - Network module - Convolution kernel 3 <![CDATA[1 2 ,512]]> 256 512 1 Network module - Convolution kernel 4 <![CDATA[3 2 ,512]]> 512 512 1 Network module - Convolution kernel 5 <![CDATA[3 2 ,512]]> 512 512 1 Network module - Convolution kernel 3 <![CDATA[1 2 ,256]]> 512 256 1 Convolution layer 2 <![CDATA[3 2 ,256]]> 256 256 1 Global average pooling layer <![CDATA[14 2 ,256]]> 256 256 1 Convolution layer 3 1×3,256 256 256 1
[0038] For S2: Construct a network model that fuses ViT and the spatial feature detection network.
[0039] Fuse the spatial feature extraction network with the ViT network. The network structure diagram is as Figure 3 shown. After ViT extracts the video frames, a slicing operation is performed. The face image information cropped in the present invention is all 224×224 pixel values. The face frame is split into 16 patches of 14×14 pixels, with a channel dimension of 3. 1D positional encoding information is adopted, and the 16 pictures are encoded in sequence. Image patch embedding processing is carried out, and through a linear mapping operation, it is mapped into 16 tokens with a channel dimension of 256. The 0th token is added and the positional encoding is incorporated.
[0040] In the Transformer Encoder, add the spatial feature extraction network structure to the Multi-HeadAttention. While processing the multi-head attention mechanism, spatial feature information is extracted in parallel. Since in the spatial feature extraction network, feature attention mechanisms are extracted both in the channel dimension and the spatial dimension, the extracted feature weights and channel weights are weighted into the query matrix for more detailed feature extraction. Then modify the fully connected layer in the Transformer Encoder. Since it is a binary classification problem, the classification head is corrected to 2 dimensions.
[0041] For S3: Train the constructed network model that fuses ViT and the spatial feature detection network.
[0042] Use the Dlib face extractor to perform face extraction operations on the video frames, save the face sampling point information, and crop the face pictures of the extracted frames. Since the main target for deepfake video detection is to detect fake faces, the deepfake video dataset needs to be preprocessed to extract the required face data information.
[0043] The specific steps for training the network model that fuses ViT and the spatial feature detection network are as follows: Pre-train the network model on the fake video dataset; divide the dataset into a training set, a validation set, and a test set and perform rotation and scaling processing; use the validation set to adjust the hyperparameters, and finally test the model effect through the test set.
[0044] Pre-training the network model on a dataset of forged videos means using a deepfake video dataset to pre-train the constructed network. Since the ViT model has a relatively large number of parameters and a relatively long training time, pre-training can be carried out first to obtain better initial values, which is convenient for the subsequent training to converge quickly.
[0045] The steps of dividing the forged video dataset into a training set, a validation set, and a test set and performing normalization are to adjust the network hyperparameters and evaluate the network performance. Since overfitting will eventually occur during the training process, by validating on the validation set and comparing the test results of the test set, a better number of training epochs can be obtained.
[0046] When training the network model, criterion is selected as the loss function to calculate the loss between the model output outputs and the true labels labels. Adam is selected as the optimizer supervised by binary cross-entropy loss for optimization. After multiple rounds of training and validation, it is found that overfitting is likely to occur after 30 epochs. Therefore, the training epoch is set to 30, the initial learning rate is 0.0002, and it decays by 10 times every 10 training epochs. Finally, the test set is used to evaluate the network performance.
[0047] For S4: Input the processed video data into the trained fused ViT and spatial feature detection network to determine whether the video is forged.
[0048] After the processed forged video data is inferred by the trained fused ViT and spatial feature deepfake video detection model, its output is a probability. By setting a threshold of 0.5 to determine whether the judgment is accurate. When the probability is higher than 0.5, the video judgment is accurate, that is, a real video is judged as true and a forged video is judged as false. Otherwise, the judgment is inaccurate. Its display of the final result shows the accuracy of correctly judging true and false videos.
[0049] The above specific implementation manners are only used to illustrate the technical solutions of the present invention, rather than limiting it. Those skilled in the art should understand that the above implementation manners do not limit the present invention in any form. All similar technical solutions obtained by using equivalent replacements or equivalent transformations and the like fall within the protection scope of the present invention.
Claims
1. A deep fake video detection method integrating ViT and spatial features, characterized by: It includes the following steps: Step 1: Build a spatial feature extraction network: The spatial feature extraction network consists of a spatial inconsistency module and a channel attention mechanism module, including convolution layer 1, network module, convolution layer 2, global average pooling layer, and convolution layer 3. Each global extraction convolution layer contains a normalization layer and a nonlinear activation layer. The module network extracts features and reduces the number of parameters. Convolution layer 1 is composed of a 5×5 convolution kernel with a channel dimension of 256 and a stride of 1. The input channel is set to 256 and the output channel is set to 256, which is used to preliminarily extract image detail features; the network module is composed of an average pooling layer, a multi-layer convolution kernel, and a bilinear interpolation layer. The average pooling layer is composed of a 2×2 convolution kernel with a channel dimension of 512 and a stride of 2. The input channel is set to 256 and the output channel is set to 512 for downsampling operations; then, feature extraction is performed by 1 1×3 and 1 3×1, and the input channel dimensions are set to 512 and 512, respectively, and the output channel is set to 512. , 256; add bilinear interpolation operation after 3×1 convolution for dimensionality increase processing to reduce the amount of calculation; then use two 1×1 convolution kernels and two 3×3 convolution kernels to extract detail features, and connect them through skip connections through the residual idea; the average pooling layer is composed of 14×14 convolution kernels with a channel dimension of 256, and the input and output channels are set to 256 to reduce the amount of calculation; then introduce the channel attention mechanism, convolution layer 3 is set to 1×3, the channel dimension is 256, the step size is 1, and the input and output channels are set to 256, which is used for feature extraction and trade-off in the channel dimension; Step 2: Build a fusion ViT and spatial feature detection network: ViT is integrated with the designed spatial feature extraction network. After extracting the video frame, ViT performs a slicing operation to split the 224×224 pixel frame into 16 14×14 pixel patches. After flattening, the channel dimension is 3. Using 1D position encoding information, the 16 images are encoded in sequence, embedded in the position encoding process, and mapped into 16 tokens with a channel dimension of 128 through a linear mapping operation. The 0th token is added and integrated into the position encoding. In the Transformer Encoder, the spatial feature extraction network structure is added to the Multi-HeadAttention, and the spatial feature information is extracted in parallel. The extracted feature weights and channel weights are weighted into the query, key, and value matrices to perform more detailed feature extraction and guide weight updates. Step 3: Train and build a fusion ViT and spatial feature detection network model: The specific steps of training the network model that integrates ViT and spatial feature detection are as follows: use the Dlib face extractor to extract faces from video frames, save face sampling point information, and crop face images of the extracted frames; pre-train the network model on a dataset of forged videos; divide the dataset into training set, validation set, and test set and perform rotation and scaling; Use the validation set to adjust the hyperparameters, and finally use the test set to verify the model effect; the steps of dividing the fake video dataset into training set, validation set and test set and standardizing them are as follows: divide the dataset into training set, validation set and test set in a ratio of 3:1:1; Step 4: Input the processed video data into the trained fusion ViT and spatial feature detection network. The output of the network model is a probability, which shows the final result as the accuracy of correct judgment of true and false videos.
2. According to claim 1, a deep fake video detection method integrating ViT and spatial features is characterized in that: Various image enhancement methods are used to perform data augmentation on the video extraction frames in the training set.
3. According to claim 1, a deep fake video detection method integrating ViT and spatial features is characterized in that: Use Dlib's face 68 sampling point detector to extract face data from video frames.
4. According to claim 1, a deep fake video detection method integrating ViT and spatial features is characterized in that: When training the network model, criterion is selected as the loss function to calculate the loss between the model outputs and the true label labels. Adam is selected as the optimizer supervised by binary cross entropy loss, and multiple rounds of training and verification are performed.
5. According to claim 1, a deep fake video detection method integrating ViT and spatial features is characterized in that: The spatial feature extraction network consists of two parts, namely a spatial inconsistency module network and a channel attention module network, which consists of 2 1×1 convolutions, 2 3×3 convolutions, 1×3 and 3×1 convolutions, up and down sampling, normalization, nonlinear activation, and attention mechanism modules.
6. The deep fake video detection method integrating ViT and spatial features according to claim 1, characterized in that: Each module network uses deep separation convolution to reduce network parameters, and integrates the residual connection idea to make the network design deeper, and the non-linear activation layer uses the sigmoid function.
7. The deep fake video detection method integrating ViT and spatial features according to claim 1, characterized in that: ViT's Encoder extracts the confidence of spatial features and then performs a weighted adjustment strategy.
8. The deep fake video detection method integrating ViT and spatial features according to claim 1, characterized in that: The specific method of inputting the processed data into the trained weather recognition network and outputting the category to which it belongs is as follows: inputting the processed video data into the trained fusion ViT and spatial feature detection network. The output of the network model is a probability. The threshold value of 0.5 is set to determine whether the judgment is accurate. When the probability is higher than 0.5, the video judgment is accurate, that is, the real video is judged to be true and the forged video is judged to be false. Otherwise, the judgment is inaccurate. Finally, the overall accuracy value is output.
Citation Information
Patent Citations
False video detection method and system based on multi-scale convolutional network and ViT
CN114387641A
Deep forgery detection method based on face geometrical relationship reasoning
CN116758604A
Deformable face authentic identification network and time-space consistent face authentic identification model construction method
CN116978103A
Face forgery detection method based on image and video multi-domain fusion
CN118351427A
Deep forgery detection method and system based on graph convolution and multi-scale prompt fusion
CN118941936A
Cited By
Deep fake face image detection method based on double-flow CNN and ViT hybrid model
CN120833637A
Depth counterfeit video detection method based on weighted feature pyramid
CN121236669A
Generated image detection method and device, electronic equipment and readable storage medium
CN121527532A