False Video Detection Method and System Based on Multi-Scale Convolutional Network and ViT

By introducing multi-scale convolutional networks and Vision Transformer into Deepfake fake video detection technology, the problem of low accuracy of low quality video detection is solved, and higher detection accuracy and performance are achieved.

CN114387641BActive Publication Date: 2025-06-10SUN YAT SEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111573856.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-21
Publication Date
2025-06-10
Estimated Expiration
2041-12-21

AI Technical Summary

Technical Problem

When detecting low-quality videos, the existing Deepfake fake video detection technology has low detection accuracy and insufficient detection performance, especially when the boundaries between real videos and fake videos are blurred.

Method used

The false video detection method based on multi-scale convolutional network and Vision Transformer (ViT) is adopted. The edge information of images in low-quality false videos is learned through the multi-scale feature extraction module, combined with the face high-dimensional semantic information extraction module, fuses information at different scales, and uses the ViT module to replace the traditional global average pooling and full connection layer to improve the detection accuracy.

Benefits of technology

It significantly improves the detection accuracy and detection performance of low-quality fake videos, and can more effectively distinguish between real videos and fake videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114387641B_ABST
    Figure CN114387641B_ABST
Patent Text Reader

Abstract

The present invention proposes a method and system for detecting fake videos based on a multi-scale convolutional network and ViT, belonging to the technical field of fake video detection. First, the dataset to be detected is processed to obtain a video frame sequence, and the face regions of the images in the video frame sequence of the dataset to be detected are identified and extracted. Then, a fake video detection model based on a multi-scale convolutional network and ViT is built. Based on this model, face features are accurately extracted, and different scale information of the face regions is fused at the same time. Among them, the multi-scale feature extraction module obtains multi-scale features of the entire face picture by learning the edge information of the images in low-quality fake videos, and uses ViT to replace the global average pooling and fully connected layers for classification, improving the detection accuracy and detection performance of low-quality fake videos.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of fake video detection, and more specifically, to a fake video detection method and system based on a multi-scale convolutional network and ViT. Background Art

[0002] The Deepfake video tampering technology is a tampering technology that generates fake faces by a deep network model and then replaces the faces in the real video with the generated fake faces. With the upgrade of the face-swapping technology and the open source of related applications, the use of face-swapping has gradually evolved from entertainment at the beginning to a criminal tool, posing a potential threat to people's reputations and social stability. Therefore, the detection of Deepfake fake videos is urgent and of great practical significance.

[0003] Currently, most Deepfake fake video detection technologies usually use a single-stream convolutional neural network to extract the face features of fake video frames, obtain their high-dimensional face feature maps, and then use global average pooling and fully connected layers to achieve classification to distinguish real videos from fake videos.

[0004] A video fake face detection method and an electronic device are disclosed in the prior art. The method mentions that: first, face localization is performed on the video to be detected to obtain a face sequence, then the face sequence is preprocessed to obtain a video sampling frame sequence with a specified size and length, and the video sampling frame sequence is input into a trained three-dimensional residual learning convolutional neural network to determine whether the face in the video to be detected is a fake face. Among them, the three-dimensional residual learning convolutional neural network includes one or more convolutional layers and corresponding max-pooling layers, several three-dimensional residual learning layers composed of one or more three-dimensional residual learning modules, an average pooling layer, and an output layer; the three-dimensional residual learning module includes a first branch, a second branch respectively connected to the input of the three-dimensional residual learning module, and an operation layer that adds the output results of the two first branches and the second branch. This method has good detection performance for the detection of high-quality videos, but when detecting highly compressed low-quality videos, the detection accuracy will decrease. Moreover, low-quality videos, as highly compressed or repeatedly compressed videos, blur the boundary between real videos and fake videos, making the detection of fake videos more difficult. Summary of the Invention

[0005] To solve the problems of low detection accuracy and insufficient detection performance of the current Deepfake fake video detection technology when detecting low-quality videos, the present invention proposes a fake video detection method and system based on a multi-scale convolutional network and ViT, which accurately extracts face features, fuses different scale information of the face region, and improves the detection accuracy and detection performance of the fake video detection technology for low-quality fake videos.

[0006] To achieve the above technical effects, the technical solution of the present invention is as follows:

[0007] A method for detecting fake videos based on a multi-scale convolutional network and ViT, characterized by comprising the following steps:

[0008] S1. Determine the video dataset to be detected, decode the videos in the video dataset to be detected into a frame sequence, and randomly sample and select from the frame sequence to obtain the frame sequence S;

[0009] S2. Identify and extract the face regions in the frame sequence S, and then preprocess them to obtain the feature extraction regions;

[0010] S3. Build a fake video detection model based on a multi-scale convolutional network and ViT, including a preprocessing module, a multi-scale feature extraction module, a high-dimensional face semantic information extraction module, and a ViT module;

[0011] S4. Input the RGB images of the feature extraction regions into the preprocessing module for color feature learning to obtain the color feature f p ;

[0012] S5. Extract the multi-scale feature map f p ' of f through the multi-scale feature extraction module, and transform the multi-scale feature map fp' into a high-dimensional multi-scale feature map f p ; 1 ;

[0013] S6. Transform the color feature fp into a high-dimensional face semantic feature f 2 through the high-dimensional face semantic information extraction module, and fuse the high-dimensional multi-scale feature map f 1 with the high-dimensional face semantic feature f 2 into a feature map

[0014] S7. Use the ViT module to learn the global information of the feature map and make a prediction to obtain the classification output results of real and fake videos.

[0015] In this technical solution, the face regions in the image of the video frame sequence of the video dataset to be detected are identified and extracted, a fake video detection model based on a multi-scale convolutional network and ViT is built, and face features are accurately extracted based on this model. At the same time, different scale information of the face regions is fused. Among them, the multi-scale feature extraction module obtains the multi-scale features of the entire face picture by learning the edge information of the images in low-quality fake videos, and uses the ViT module to replace the traditional global average pooling and fully connected layers for classification, improving the detection accuracy and detection performance for low-quality fake videos.

[0016] Preferably, the video dataset to be detected in step S1 includes high-quality real videos, high-quality fake videos, compressed low-quality real videos, and compressed low-quality fake videos. The high-quality real videos and high-quality fake videos are used as high-quality videos and are trained separately from the compressed low-quality real videos and compressed low-quality fake videos during training. After decoding the videos in the video dataset to be detected into frame sequences, each frame sequence is stored in an independent folder to prevent interference between different videos.

[0017] Preferably, in step S2, each video frame image in the frame sequence S is traversed and read, and face region recognition is performed on the video frame image. When preprocessing the face region, the center of the recognized face region is determined, and a face region of a specific size is selected based on the center as the feature extraction region.

[0018] Preferably, the preprocessing module described in step S3 uses EfficientNet-B4 as the baseline convolutional neural network, including a 3*3 convolutional layer and the first ten MBConv Blocks of EfficientNet-B4 connected in sequence;

[0019] The multi-scale feature extraction module is connected to the preprocessing module. The multi-scale feature extraction module includes a dilated convolution unit and a depthwise separable convolution unit connected in sequence. The dilated convolution unit includes L parallel dilated convolutions with different receptive fields. The depthwise separable convolution unit includes Q depthwise separable convolution blocks and P residual separable convolution blocks. Each depthwise separable convolution block consists of a relu non-linear activation function, a depthwise separable convolution, and a normalization layer. Each residual separable convolution block has a linear residual connection in the depthwise separable convolution block;

[0020] The face high-dimensional semantic information extraction module is connected to the preprocessing module. The face high-dimensional semantic information extraction module is based on EfficientNet-B4 as the basic network and is specifically composed of the last 22 MBConv Blocks of EfficientNet-B4 connected in sequence;

[0021] After the output ends of the multi-scale feature extraction module and the face high-dimensional semantic information extraction module are fused, they are connected to the ViT module. The ViT module includes a depthwise separable convolution block and a Vision Transformer module connected in sequence. The Vision Transformer module includes an Embedding layer, a Transformer Encoder layer, and an MLP Head classification layer.

[0022] Preferably, the process of inputting the RGB image of the feature extraction region into the preprocessing module for color feature learning in step S4 is as follows:

[0023] S41. Resize the feature extraction region into an RGB image of (H, W, 3), and perform normalization processing as color feature data, where H represents the height of the RGB image, W represents the width of the RGB image, and 3 represents the number of channels;

[0024] S42. Set the training parameters and loss function of the preprocessing module, and train the preprocessing module to obtain a trained preprocessing module;

[0025] S43. Input the resized RGB image into the trained preprocessing module for color convolution feature learning, and select the tensors output by the first ten MBConv Blocks of EfficientNet-B4 as the color feature f p 。

[0026] Preferably, the process of step S5 is as follows:

[0027] S51. Set the dilation sizes and convolution kernel sizes of the parallel dilated convolutions with L different receptive fields in the dilated convolution unit;

[0028] S52. Input the color feature f p into the parallel dilated convolutions with L different receptive fields of the dilated convolution unit respectively, and use the parallel dilated convolutions with L different receptive fields to extract the face edge feature information respectively to obtain L scale feature maps F 1 ,…,F L ;

[0029] S53. Fuse the L scale feature maps F 1 ,…,F L with the color feature fp to obtain a multi-scale feature map fp';

[0030] S54. Input the multi-scale feature map f p ' into the depthwise separable convolution unit, and use the depthwise separable convolution block and residual separable convolution block of the depthwise separable convolution unit to transform the multi-scale feature map fp' into a high-dimensional multi-scale feature map f 1 。

[0031] Preferably, in step S6, the face high-dimensional semantic information extraction module receives the color feature f output by the preprocessing module p , and the last 22 consecutive MBConv Blocks of EfficientNet-B4 in the face high-dimensional semantic information extraction module transform the color feature f p into a high-dimensional face semantic feature f 2 。

[0032] Preferably, in step S7, use the ViT module to learn the feature map F fuseWhen obtaining the global information of the global information, set the cross-entropy loss function, backpropagate the weight parameters of the ViT module, and obtain a trained fake video detection model;

[0033] The output end of the multi-scale feature extraction module outputs a high-dimensional multi-scale feature map f 1 , and the output end of the high-dimensional face semantic information extraction module outputs high-dimensional face semantic features f 2 The high-dimensional multi-scale feature map f 11 and the high-dimensional face semantic features f 2 are fused into a feature map The feature map is input into the ViT module, and the depthwise separable convolution block is used to learn the information of the feature map from two independent dimensions of space and channel and improve the dimension of the feature map The dimension of the upsampled feature map is divided into several blocks Patches, each block Patch is mapped to a one-dimensional vector through a linear mapping, and then input into the Transformer Encoder layer after PositionEmbedding; in this way, the transformer can not only obtain the information of the entire feature map through the self-attention mechanism, but also understand the structure of the input feature map through the learnable position embedding, so as to combine the local and global information of the feature map F fuse together.

[0034] Preferably, in step S7, the ViT module is used to learn the global information of the feature map and make a prediction. The predicted result is passed through a softmax function to obtain the probability values of the model predicting real and fake videos. Through the probability values, the prediction result (real or fake) of the video is obtained.

[0035] The present invention also proposes a fake video detection system based on a multi-scale convolutional network and ViT, including:

[0036] A video dataset to be detected processing module, which is used to determine the video dataset to be detected, decode the videos in the video dataset to be detected into a frame sequence, and randomly sample and select the frame sequence to obtain a frame sequence S;

[0037] A face region recognition processing module, which recognizes and extracts the face regions in the frame sequence S, and then preprocesses them to obtain a feature extraction region;

[0038] A detection model construction module, which is used to build a fake video detection model based on a multi-scale convolutional network and ViT;

[0039] The false video detection model based on multi-scale convolutional network and ViT includes a preprocessing module, a multi-scale feature extraction module, a high-dimensional face semantic information extraction module, and a ViT module;

[0040] The preprocessing module conducts color feature learning on the RGB image of the feature extraction region to obtain the color feature fp; the multi-scale feature extraction module extracts the multi-scale feature map f p of p ', and transforms the multi-scale feature map f p ' into a high-dimensional multi-scale feature map f 1 ; the high-dimensional face semantic information extraction module transforms the color feature fp into a high-dimensional face semantic feature f 2 , and fuses the high-dimensional multi-scale feature map f 1 with the high-dimensional face semantic feature f 2 into a feature map The ViT module learns the global information of the feature map and makes a prediction to obtain the classification output result of real and false videos.

[0041] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:

[0042] The present invention proposes a false video detection method and system based on a multi-scale convolutional network and ViT. First, the dataset to be detected is processed to obtain a video frame sequence, the face region of the images in the video frame sequence of the dataset to be detected is identified and extracted, and then a false video detection model based on a multi-scale convolutional network and ViT is built. Based on this model, face features are accurately extracted, and different scale information of the face region is fused. Among them, the multi-scale feature extraction module obtains the multi-scale features of the entire face picture by learning the edge information of the images in low-quality false videos, and uses the ViT module to replace the traditional global average pooling and fully connected layers for classification, improving the detection accuracy and performance of low-quality false videos. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 represents the flowchart of the false video detection method based on a multi-scale convolutional network and ViT proposed in Embodiment 1 of the present invention;

[0044] Figure 2 represents the structural block diagram of the false video detection model based on a multi-scale convolutional network and ViT built in Embodiment 1 of the present invention;

[0045] Figure 3 represents the structural schematic diagram of the false video detection system based on a multi-scale convolutional network and ViT proposed in Embodiment 3 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] The accompanying drawings are only for illustrative purposes and should not be construed as limiting the patent;

[0047] To better illustrate this embodiment, some parts of the accompanying drawings will be omitted, enlarged or reduced, which does not represent the actual size;

[0048] For those skilled in the art, it is understandable that some well-known content descriptions in the accompanying drawings may be omitted.

[0049] The descriptions of the positional relationships in the accompanying drawings are only for illustrative purposes and should not be construed as limiting the patent;

[0050] The technical solutions of the present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0051] Embodiment 1

[0052] Considering that the current fake video detection technology has achieved good detection effects on high-quality true and false video sets, but when facing highly compressed videos (such as videos highly compressed by H.264, which are called low-quality videos in this application), since the low-quality videos blur the boundaries between real videos and fake videos and are difficult to detect, the detection performance is not good. To overcome this defect, a fake video detection method based on a multi-scale convolutional network and ViT is proposed in this embodiment. This method uses a multi-scale convolutional network with a Vision transformer (ViT) to detect the authenticity of low-quality videos. For the flowchart, see Figure 1 , including the following steps:

[0053] S1. Determine the video dataset to be detected, decode the videos in the video dataset to be detected into a frame sequence, and randomly sample and select from the frame sequence to obtain the frame sequence S;

[0054] S2. Identify and extract the face regions in the frame sequence S, and then preprocess them to obtain the feature extraction regions;

[0055] S3. Build a fake video detection model based on a multi-scale convolutional network and ViT, including a preprocessing module, a multi-scale feature extraction module, a face high-dimensional semantic information extraction module, and a ViT module;

[0056] S4. Input the RGB images of the feature extraction regions into the preprocessing module for color feature learning to obtain the color feature f p ;

[0057] S5. Extract the multi-scale feature map f p ' of f p through the multi-scale feature extraction module, and transform the multi-scale feature map fp' into a high-dimensional multi-scale feature map f 1 ;

[0058] S6. Transform the color feature fp into a high-dimensional face semantic feature f through the face high-dimensional semantic information extraction module 2 , and combine the high-dimensional multi-scale feature map f 1 with the high-dimensional face semantic feature f 2 to form a feature map

[0059] S7. Use the ViT module to learn the global information of the feature map and make a prediction to obtain the classification output result of real and fake videos.

[0060] In this embodiment, the video data set to be detected in step S1 includes high-quality real videos, high-quality fake videos, compressed low-quality real videos, and compressed low-quality fake videos. High-quality real videos and high-quality fake videos are used as high-quality videos and are separately trained from compressed low-quality real videos and compressed low-quality fake videos when used for training. Use the opencv library of python to decode the videos in the video data set to be detected. After the videos are decoded into frame sequences, each frame sequence is stored in an independent folder to prevent interference between different videos.

[0061] In step S2, traverse and read each video frame image in the frame sequence S, perform face region recognition on the video frame image, use the MTCNN face detection model to recognize the face in the video frame image, extract the face region, and when preprocessing the face region, determine the center of the recognized face region, and select a face region of a specific size based on the center as the feature extraction region. In this embodiment, a size of 320x320 is selected as the feature extraction region.

[0062] After the above processing of the video frames, starting from the idea of using a multi-scale convolutional network with Vision transformer in this application to detect the authenticity of low-quality videos, build a fake video detection model based on the multi-scale convolutional network and ViT. The model structure is shown in Figure 2 , including a preprocessing module, a multi-scale feature extraction module, a face high-dimensional semantic information extraction module, and a ViT module. Specifically:

[0063] The preprocessing module corresponds to Figure 2 the Pre-processing Module shown. The preprocessing module uses EfficientNet-B4 as the benchmark convolutional neural network, including a 3*3 convolutional layer connected in sequence and the first ten MBConv Blocks (MBConv Blocks#1~MBConv Blocks#10) of EfficientNet-B4;

[0064] The multi-scale feature extraction module is connected to the preprocessing module. The multi-scale feature extraction module corresponds to Figure 2 the Stream#1:Multi-scale Module shown. The multi-scale feature extraction module includes a dilated convolution unit and a depthwise separable convolution unit connected in sequence. The dilated convolution unit includes L parallel dilated convolutions with different receptive fields. In this embodiment, the dilated convolution unit uses 4 parallel dilated convolutions with different receptive fields to extract the face edge feature information. The depthwise separable convolution unit includes Q depthwise separable convolution blocks and P residual separable convolution blocks. Refer to Figure 2 , there are two depthwise separable convolution blocks in total, denoted as: SeparableConv Block; there are 4 residual separable convolution blocks in total, denoted as: ResidualSeparableConv Block; more specifically, each depthwise separable convolution block consists of a relu non-linear activation function, a depthwise separable convolution, and a normalization layer. Each residual separable convolution block has a linear residual connection in the depthwise separable convolution block.

[0065] The high-dimensional face semantic information extraction module is connected to the preprocessing module. The high-dimensional face semantic information extraction module corresponds to Figure 2 the Stream#2:MBConv Blocks Module shown. The high-dimensional face semantic information extraction module is based on the EfficientNet-B4 network, and is specifically composed of the last 22 MBConv Blocks of the EfficientNet-B4 connected in sequence.

[0066] After the output ends of the multi-scale feature extraction module and the high-dimensional face semantic information extraction module are fused, they are connected to the ViT module. The ViT module corresponds to Figure 2 the Vision Transformer Module shown. The ViT module includes a depthwise separable convolution block SeparableConv Block and a Vision Transformer module connected in sequence. The Vision Transformer module includes an Embedding layer, a Transformer Encoder layer, and an MLP Head classification layer.

[0067] In this embodiment, the color feature learning performed by inputting the RGB image of the feature extraction area into the preprocessing module in step S4 is implemented based on the constructed fake video detection model. The process is as follows:

[0068] S41. Resize the feature extraction region to an RGB image of (H, W, 3), and perform normalization processing as color feature data, where H represents the height of the RGB image, W is the width of the RGB image, and 3 is the number of channels. At this time, both H and W are 320. The normalization processing adopts conventional operations and will not be elaborated here;

[0069] S42. Set the training parameters and loss function of the preprocessing module, and train the preprocessing module to obtain a trained preprocessing module;

[0070] S43. Input the adjusted RGB image into the trained preprocessing module for color convolution feature learning, and select the tensors output by the first ten MBConv Blocks of EfficientNet - B4 as the color feature f p .

[0071] The process of step S5 is as follows:

[0072] S51. Set the dilation sizes and convolution kernel sizes of the parallel dilated convolutions with L different receptive fields in the dilated convolution unit; based on the specific composition of the above - mentioned fake video detection model, L is taken as 4. The dilations of these 4 parallel dilated convolutions are set to 3, 6, 12, and 18 respectively. When applying the parallel dilated convolutions, in order not to change the size of the original feature map, the convolution kernel sizes of these four parallel dilated convolutions are all set to 1;

[0073] S52. Input the color feature f p into the parallel dilated convolutions with L different receptive fields of the dilated convolution unit respectively, and use the parallel dilated convolutions with L different receptive fields to extract the face edge feature information respectively, obtaining L scale feature maps F 1 , …, F L , that is, F 1 , F 2 , F 3 , F 4 ;

[0074] S53. Fuse the L scale feature maps F 1 , …, F L with the color feature fp to obtain a multi - scale feature map fp'; in this embodiment, fuse F 1 , F 2 , F 3 , F 4 , F RGB . Based on the Figure 2 shown model structure, this fusion process is reflected by the "confluence" of the output of the preprocessing module and the four parallel dilated convolutions in Figure 2 in Stream#1:Multi - scale Module to obtain the multi - scale feature map fp';

[0075] S54. Input the multi-scale feature map fp' into the depthwise separable convolution unit. Use the depthwise separable convolution block and the residual separable convolution block of the depthwise separable convolution unit to learn the information of the multi-scale feature map fp' from two independent dimensions of space and channel, and obtain the edge information of the low-quality video frame image, and transform the multi-scale feature map fp' into a high-dimensional multi-scale feature map f 1 。

[0076] In step S6, the face high-dimensional semantic information extraction module receives the color feature f output by the preprocessing module p , and the last 22 MBConv Blocks of EfficientNet-B4 connected in sequence in the face high-dimensional semantic information extraction module transform the color feature f p into the high-dimensional face semantic feature f 2 。

[0077] In step S7, when using the ViT module to learn the global information of the feature map F fuse , each layer of the TransformerEncoder can obtain the global information of, set the cross-entropy loss function, backpropagate the weight parameters of the ViT module, and obtain the trained false video detection model;

[0078] Specifically: the training period is 20, the optimizer is ADAM, the initial learning rate is 0.0001, and every 10 training periods, the learning rate is adjusted to 1 / 10 of the original. The loss function is designed as cross-entropy, the training batch size is 48, and after the last training period is completed, save the weight parameters of the false video detection model with the minimum loss.

[0079] See Figure 2 , the output end of the multi-scale feature extraction module outputs the high-dimensional multi-scale feature map f 1 , the output end of the face high-dimensional semantic information extraction module outputs the high-dimensional face semantic feature f 2 , the high-dimensional multi-scale feature map f 1 and the high-dimensional face semantic feature f 2 are fused into the feature map Feature map is input into the ViT module, and the information of the feature map is learned from two independent dimensions of space and channel through the depthwise separable convolution block, and the feature map In terms of the dimension, the upsampled feature map is divided into several patches. Each patch is mapped to a one-dimensional vector through linear mapping and then input into the Transformer Encoder layer after Position Embedding. In this way, the transformer can not only obtain the information of the entire feature map through the self-attention mechanism, but also understand the structure of the input feature map through learnable position embedding, thereby combining the local and global information of the feature map F fuse The local and global information is combined. The calculation formula of the attention mechanism is as follows:

[0080]

[0081] Among them, Q, K, and V respectively represent the query, key, and value vectors obtained by a group of mappings of the input feature map The distance of attention increases with the increase of the network depth, thereby combining the local and global information of the feature map The local and global information is combined.

[0082] By introducing the ViT module, the Vision transformer is used to replace the original global average pooling and fully connected layers of the traditional single-stream neural network to achieve classification.

[0083] In this embodiment, step S7 uses the ViT module to learn the global information of the feature map And make a prediction. The predicted result is passed through a softmax function to obtain the probability value of the model predicting real and fake videos. Through the probability value, the prediction result (real or fake) of the video is obtained.

[0084] Embodiment 2

[0085] This embodiment verifies the effectiveness of the method proposed in Embodiment 1 with a specific example.

[0086] The video datasets to be detected use the Deepfake video database Celeb-DF, FaceForensics++ (LQ version), and WildDeepfake as the detection datasets. Celeb-DF is a high-quality Deepfake video dataset with an average video length of approximately 13s and a frame rate of about 30; FaceForensics++ (LQ version) is a video dataset highly compressed by H.264, containing 1000 real videos and 4000 fake videos, and these 4000 fake videos are from four different forgery algorithms; WildDeepfake is a video dataset from the Internet, which may have been compressed once or multiple times and has multiple sources. This experiment is conducted on the Linux system and is mainly implemented based on the deep learning framework pytorch.

[0087] Decode the video to be detected into a frame sequence and randomly select 50 frames. Use the opencv library in python to decode the video into a frame sequence. Each video's frames are in independent folders to prevent interference between different videos. Perform face region recognition detection on the saved 50-frame sequence and use it as the feature extraction region. The specific operations are as follows:

[0088] Traverse and read the frame sequence paths in all folders, perform face recognition on the video frame images through the MTCNN face detection model, extract the face regions, write the paths of the cropped face video frames and the video labels into a csv file, read the csv file, obtain the face regions according to the paths, and adjust their centers to a size of 320x320 as the feature extraction region.

[0089] Input the RGB image of the feature extraction region (such as the picture shown) into the preprocessing module for color convolution feature learning to obtain the color feature f Figure 2 ; use the dilated convolution unit of the multi-scale feature extraction module to extract the multi-scale information of f p to obtain the multi-scale feature map f p '; in order to obtain a higher-dimensional multi-scale feature map, use the depthwise separable convolution unit of the multi-scale feature extraction module to transform the multi-scale feature map f p ' into a high-dimensional multi-scale feature map f p ; at the same time, transform the color feature f 1 into a high-dimensional face semantic feature f p through the high-dimensional face semantic information extraction module, and fuse the high-dimensional multi-scale feature map f 2 with the high-dimensional face semantic feature f p into the feature map 2 ; for the specific flow, please refer to SeeFigure 2 Finally, the ViT module is used to learn the global information of the feature map and make predictions to obtain the classification output results of real and fake videos. In this embodiment, AUC (the area covered by the ROC curve, obviously, the larger the AUC, the better the classification effect) and ACC (detection accuracy) are used as evaluation indicators, and the traditional detection method based on the Xception network is used as a comparison algorithm. The detection results of the method proposed in the present invention and the detection method based on the Xception network for three video datasets are shown in Table 1.

[0090] Table 1

[0091]

[0092] Among the three video datasets, Celeb-DF is a high-quality Deepfake video dataset, which is extremely unbalanced. AUC is used as the evaluation indicator for this dataset (because using ACC will cause large differences, and for unbalanced data, ACC will not be used in model performance measurement). WildDeepfake is a video dataset from the Internet, which may have been compressed once or multiple times and has multiple sources, belonging to a low-quality video dataset. Similarly, WildDeepfake is a relatively balanced dataset, so ACC is used as the evaluation indicator instead of AUC. Many detection datasets are unbalanced (the number of fake videos is much larger than the number of real videos), so in this case, the AUC indicator is used. If it is a balanced dataset (the ratio of fake videos to real videos is close to 1:1), the ACC indicator is used for evaluation. As can be seen from Table 1, when facing the detection of low-quality videos in WildDeepfake, the detection accuracy of the traditional detection method based on the Xception network is only 76.27%, and the detection performance is low. While the detection accuracy of the method proposed in this application reaches 82.63%. In addition, in terms of detection performance, the method proposed in this application is also superior to the traditional method, indicating that the method proposed in this application has high detection accuracy and detection performance when detecting low-quality fake videos, verifying the effectiveness of the method proposed in this application.

[0093] Embodiment 3

[0094] See Figure 3 , the present invention also proposes a fake video detection system based on a multi-scale convolutional network and ViT, including:

[0095] A processing module for the video dataset to be detected, which is used to determine the video dataset to be detected, decode the videos in the video dataset to be detected into a frame sequence, and randomly sample and select the frame sequence to obtain the frame sequence S;

[0096] The face region recognition and processing module recognizes and extracts the face region in the frame sequence S, and then performs preprocessing to obtain the feature extraction region;

[0097] The detection model construction module is used to build a fake video detection model based on a multi-scale convolutional network and ViT; The fake video detection model based on a multi-scale convolutional network and ViT includes a preprocessing module, a multi-scale feature extraction module, a high-dimensional face semantic information extraction module, and a ViT module;

[0098] The preprocessing module performs color feature learning on the RGB image of the feature extraction region to obtain the color feature fp; The multi-scale feature extraction module extracts the multi-scale feature map f p of p ', and transforms the multi-scale feature map f p ' into a high-dimensional multi-scale feature map f 1 ; The high-dimensional face semantic information extraction module transforms the color feature fp into a high-dimensional face semantic feature f 2 , and fuses the high-dimensional multi-scale feature map f 1 with the high-dimensional face semantic feature f 2 into a feature map The ViT module learns the global information of the feature map and makes a prediction to obtain the classification output result of real and fake videos.

[0099] Obviously, the above embodiments of the present invention are only examples for clearly explaining the present invention, and are not intended to limit the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.

Claims

1. A method for detecting fake videos based on a multi-scale convolutional network and ViT, characterized in that, it includes the following steps: S1. Determine the video dataset to be detected, decode the videos in the video dataset to be detected into a frame sequence, and randomly sample and select from the frame sequence to obtain the frame sequence S; S2. Identify and extract the face regions in the frame sequence S, and then preprocess them to obtain the feature extraction regions; S3. Build a fake video detection model based on a multi-scale convolutional network and ViT, including a preprocessing module, a multi-scale feature extraction module, a face high-dimensional semantic information extraction module, and a ViT module; S4. Input the RGB image of the feature extraction region into the preprocessing module for color feature learning to obtain the color feature f p ; S5. Extract the multi-scale feature map $f'$ of $f$ through the multi-scale feature extraction module, and transform the multi-scale feature map $f'$ into a high-dimensional multi-scale feature map $f$; p p 1 ;​​ S6. Transform the color feature fp into a high-dimensional face semantic feature f through the face high-dimensional semantic information extraction module 2 , and fuse the high-dimensional multi-scale feature map f 1 with the high-dimensional face semantic feature f 2 into a feature map S7. Learn the global information of the feature map using the ViT module and make predictions to obtain the classification output results of real and fake videos; ​ The preprocessing module in step S3 uses EfficientNet-B4 as the benchmark convolutional neural network, including a 3*3 convolutional layer and the first ten MBConv Blocks of EfficientNet-B4 connected in sequence; The multi-scale feature extraction module is connected to the preprocessing module. The multi-scale feature extraction module includes a dilated convolution unit and a depthwise separable convolution unit connected in sequence. The dilated convolution unit includes L parallel dilated convolutions with different receptive fields. The depthwise separable convolution unit includes Q depthwise separable convolution blocks and P residual separable convolution blocks. Each depthwise separable convolution block consists of a relu non-linear activation function, a depthwise separable convolution, and a normalization layer. Each residual separable convolution block has a linear residual connection in the depthwise separable convolution block; The face high-dimensional semantic information extraction module is connected to the preprocessing module. The face high-dimensional semantic information extraction module is based on EfficientNet-B4, and is specifically composed of the last 22 MBConvBlocks of EfficientNet-B4 connected in sequence; After the output ends of the multi-scale feature extraction module and the face high-dimensional semantic information extraction module are fused, they are connected to the ViT module. The ViT module includes a depthwise separable convolution block and a Vision Transformer module connected in sequence. The Vision Transformer module includes an Embedding layer, a Transformer Encoder layer, and an MLP Head classification layer; The process of inputting the RGB image of the feature extraction region into the preprocessing module for color feature learning in step S4 is: S41. Resize the feature extraction region to an RGB image of (H, W, 3) and perform normalization processing as color feature data, where H represents the height of the RGB image, W is the width of the RGB image, and 3 is the number of channels; S42. Set the training parameters and loss function of the preprocessing module, train the preprocessing module, and obtain the trained preprocessing module; S43. Input the adjusted RGB image into the trained preprocessing module for color convolution feature learning, and select the tensors output by the first ten MBConv Blocks of EfficientNet-B4 as the color feature f p ; The process of step S5 is: S51. Set the dilation sizes and convolution kernel sizes of the L parallel dilated convolutions with different receptive fields in the dilated convolution unit; S52. Input the color feature f p into the parallel dilated convolutions of L different receptive fields of the dilated convolution unit respectively, and use the parallel dilated convolutions of L different receptive fields to extract the face edge feature information respectively, obtaining L scale feature maps F 1 , …, F L ; S53. Fuse the L scale feature maps F 1 , …, F L with the color feature fp to obtain a multi-scale feature map fp'; S54. Input the multi-scale feature map f p ' into the depthwise separable convolution unit, and use the depthwise separable convolution block and the residual separable convolution block of the depthwise separable convolution unit to transform the multi-scale feature map fp' into a high-dimensional multi-scale feature map f 1 ; In step S6, the face high-dimensional semantic information extraction module receives the color feature f output by the preprocessing module p , and the last 22 MBConv Blocks of EfficientNet-B4 connected in sequence in the face high-dimensional semantic information extraction module transform the color feature f p into the high-dimensional face semantic feature f 2 ; In step S7, when using the ViT module to learn the global information of the feature map F fuse each layer of the Transformer Encoder can obtain the global information. Set the cross-entropy loss function and backpropagate the weight parameters of the ViT module to obtain a trained fake video detection model; The output end of the multi-scale feature extraction module outputs a high-dimensional multi-scale feature map f 1 , the output end of the high-dimensional face semantic information extraction module outputs a high-dimensional face semantic feature f 2 , the high-dimensional multi-scale feature map f 11 and the high-dimensional face semantic feature f 2 are fused into a feature map The feature map is input into the ViT module, and the feature map is learned from two independent dimensions of space and channel through a depthwise separable convolution block 's information, and the dimension of the feature map is increased. The feature map with increased dimension is divided into several blocks Patches, each block Patch is mapped to a one-dimensional vector through a linear mapping, and then input into the Transformer Encoder layer after PositionEmbedding.

2. According to the method for detecting fake videos based on a multi-scale convolutional network and ViT described in claim 1, characterized in that, The video data set to be detected in step S1 includes high-quality real videos, high-quality fake videos, compressed low-quality real videos, and compressed low-quality fake videos. High-quality real videos and high-quality fake videos are regarded as high-quality videos, and they are trained separately from compressed low-quality real videos and compressed low-quality fake videos when used for training. After decoding the videos in the video data set to be detected into frame sequences, each frame sequence is stored in an independent folder.

3. The fake video detection method based on a multi-scale convolutional network and ViT according to claim 2, characterized in that in step S2, each video frame image in the frame sequence S is traversed and read, the face area in the video frame image is recognized, and when preprocessing the face area, the center of the recognized face area is determined, and a face area of a specific size is selected based on the center as the feature extraction area.

4. The fake video detection method based on a multi-scale convolutional network and ViT according to claim 1, characterized in that In step S7, the ViT module is used to learn the global information of the feature map and make a prediction. The prediction result passes through a softmax function to obtain the probability values of the model predicting real and fake videos.

5. A fake video detection system based on a multi-scale convolutional network and ViT, characterized in that comprising: a processing module for the video data set to be detected, configured to determine the video data set to be detected, decode the videos in the video data set to be detected into frame sequences, and randomly sample and select the frame sequences to obtain the frame sequence S; a face area recognition processing module, configured to recognize and extract the face area in the frame sequence S, and then preprocess it to obtain the feature extraction area; a detection model construction module, configured to build a fake video detection model based on a multi-scale convolutional network and ViT; The fake video detection model based on a multi-scale convolutional network and ViT includes a preprocessing module, a multi-scale feature extraction module, a face high-dimensional semantic information extraction module, and a ViT module; The preprocessing module performs color feature learning on the RGB image of the feature extraction area to obtain the color feature fp; The multi-scale feature extraction module extracts the multi-scale feature map f p ', and transforms the multi-scale feature map f p ' into a high-dimensional multi-scale feature map f p ; 1 ; The high-dimensional semantic information extraction module of the human face transforms the color feature fp into the high-dimensional human face semantic feature f 2 , and fuses the high-dimensional multi-scale feature map f 1 with the high-dimensional human face semantic feature f 2 into the feature map The ViT module learns the global information of the feature map and makes a prediction to obtain the classification output results of real and fake videos; The preprocessing module uses EfficientNet-B4 as the benchmark convolutional neural network, and includes a 3*3 convolutional layer and the first ten MBConvBlocks of EfficientNet-B4 connected in sequence; The multi-scale feature extraction module is connected to the preprocessing module. The multi-scale feature extraction module includes a dilated convolution unit and a depthwise separable convolution unit connected in sequence. The dilated convolution unit includes L parallel dilated convolutions with different receptive fields. The depthwise separable convolution unit includes Q depthwise separable convolution blocks and P residual separable convolution blocks. Each depthwise separable convolution block consists of a relu non-linear activation function, a depthwise separable convolution, and a normalization layer. Each residual separable convolution block has a linear residual connection in the depthwise separable convolution block; The face high-dimensional semantic information extraction module is connected to the preprocessing module. The face high-dimensional semantic information extraction module is based on EfficientNet-B4 as the basic network, and is specifically composed of the last 22 MBConvBlocks of EfficientNet-B4 connected in sequence; After the output end of the multi-scale feature extraction module is fused with the output end of the human face high-dimensional semantic information extraction module, it is connected to the ViT module. The ViT module includes a depthwise separable convolution block and a Vision Transformer module connected in sequence. The Vision Transformer module includes an Embedding layer, a Transformer Encoder layer, and an MLP Head classification layer; The process of inputting the RGB image in the feature extraction area into the preprocessing module for color feature learning is as follows: S41. Resize the RGB image in the feature extraction area to an RGB image of (H, W, 3) and perform normalization processing as color feature data, where H represents the height of the RGB image, W is the width of the RGB image, and 3 is the number of channels; S42. Set the training parameters and loss function of the preprocessing module, train the preprocessing module, and obtain the trained preprocessing module; S43. Input the adjusted RGB image into the trained preprocessing module for color convolution feature learning, and select the tensors output by the first ten MBConv Blocks of EfficientNet-B4 as the color feature f p ; The multi-scale feature extraction module extracts the multi-scale feature map f p ', and the process of transforming the multi-scale feature map f p ' into the high-dimensional multi-scale feature map f p is as follows: 1 ​ S51. Set the dilation sizes and convolution kernel sizes of the parallel dilated convolutions with L different receptive fields in the dilated convolution unit; S52. Input the color feature f p into the parallel dilated convolutions of L different receptive fields of the dilated convolution unit respectively, and use the parallel dilated convolutions of L different receptive fields to extract the face edge feature information respectively, obtaining L scale feature maps F 1 , …, F L ; S53. Fuse the L scale feature maps F 1 , …, F L with the color feature fp to obtain a multi-scale feature map fp'; S54. Input the multi-scale feature map f p ' into the depthwise separable convolution unit, and use the depthwise separable convolution block and the residual separable convolution block of the depthwise separable convolution unit to transform the multi-scale feature map fp' into a high-dimensional multi-scale feature map f 1 ; The high-dimensional semantic information extraction module for human faces receives the color feature f output by the preprocessing module p , and the last 22 MBConv Blocks of EfficientNet-B4 connected in sequence in the high-dimensional semantic information extraction module for human faces transform the color feature f p into the high-dimensional human face semantic feature f 2 ; Learn the feature map F using the ViT module fuse When learning the global information, each layer of the Transformer Encoder can obtain the global information, set the cross-entropy loss function, backpropagate the weight parameters of the ViiT module, and obtain a trained fake video detection model; The output end of the multi-scale feature extraction module outputs a high-dimensional multi-scale feature map f 1 , the output end of the high-dimensional face semantic information extraction module outputs high-dimensional face semantic features f 2 , the high-dimensional multi-scale feature map f 11 is fused with the high-dimensional face semantic features f 2 to form a feature map The feature map is input into the ViT module, and the feature map is learned from two independent dimensions of space and channel through a depthwise separable convolution block 's information, and the dimension of the feature map is increased. The feature map with increased dimension is divided into several blocks Patches, each block Patch is mapped to a one-dimensional vector through a linear mapping, and then input into the Transformer Encoder layer after PositionEmbedding.