Face-swapped Video Detection Method and System Based on Predicting Pixel-Level Tampering Probability Values in the Spatiotemporal Domain
Through multi-level spatiotemporal feature extraction network and dual attention mechanism, pixel-level tampering probability values are generated, which solves the problems of the existing Deepfake detection algorithm's cross-border performance and high computational complexity in high-quality face swap videos, and realizes efficient and accurate face swap video detection.
Patent Information
- Application Number
- CN202211391297.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-08
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-11-08
AI Technical Summary
The existing Deepfake detection algorithm based on image segmentation idea is difficult to effectively detect high-quality face-changing videos, and there are problems such as degradation of cross-base performance and high model calculation complexity.
A multi-level spatiotemporal feature extraction network and dual attention mechanism are adopted to build tamper masks and random sampling coordinate positions, combined with DenseNet and ConvLSTM modules, pixel-level tamper probability values are generated, and auxiliary supervision is performed to improve detection performance and generalization.
It realizes efficient detection of face-switching videos in library and cross-store testing, reduces the complexity of model calculations, and improves the accuracy and generalization capabilities of detection.
Smart Images

Figure CN115719462B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital video tampering detection, and particularly relates to a face-swapping video detection method and system based on predicting pixel-level tampering probability values in the spatio-temporal domain. Background Art
[0002] With the rapid development of face forgery methods, more and more open-source face-swapping software (such as DeepFakes, DeepFaceLab2.0) has emerged, which is easily accessible to the public for making various face-swapping videos. The abuse of face-swapping software has attracted wide public attention. Therefore, there is an urgent need to develop effective detection technologies for face-swapping tampered videos.
[0003] Currently, most Deepfake detection algorithms based on the idea of image segmentation regard face-swapping videos as special splicing tampering problems, believing that the most discriminative features of tampered images are local rather than global. They more often use popular neural segmentation networks, such as fully convolutional network models that stack a large number of convolutional layers and transposed convolutional layers to generate large-scale tampering area prediction maps to detect local tampering traces, rather than collecting global statistical features from different levels globally. However, the forged faces produced by the latest face-swapping technologies have very high quality, and it is often difficult to extract subtle tampering information from the local part of the image. Detection algorithms that only extract tampering features from the local part are prone to the problem of insufficient domain generalization where the cross-library performance drops sharply. Therefore, from the perspective of global forensics, the authenticity should be detected by leveraging the inconsistency between the background and the face region. In addition, most fully convolutional network models have relatively large parameters, high model computational complexity, and the network is prone to overfitting. Summary of the Invention
[0004] In order to overcome the defects and deficiencies existing in the prior art, the present invention provides a face-swapping video detection method and system based on predicting pixel-level tampering probability values in the spatio-temporal domain. The present invention uses a model with a multi-level spatio-temporal feature extraction network and a dual attention mechanism to extract spatio-temporal inconsistency features. Spatio-temporal inconsistency features are extracted from the perspective of multi-resolution in time, and for the tampering mask, multiple groups of coordinate position-tampering value pairs (x coordinate, y coordinate, tampering value) are constructed by sampling. The spatio-temporal inconsistency features are concatenated with the xy-axis coordinates of several randomly sampled pixel points, and the tampering probability prediction value of the pixel point is obtained through coordinate position reconstruction, for auxiliary supervision of the pixel-level tampering probability value. It has good detection performance in both in-library and cross-library tests and has good generalization.
[0005] To achieve the above object, the present invention adopts the following technical solutions:
[0006] The present invention provides a face-swapping video detection method based on predicting pixel-level tampering probability values in the spatio-temporal domain, including the following steps:
[0007] Frame the video to be tested and detect and extract face bounding boxes for each frame of the image;
[0008] Construct a tampering mask based on facial key points;
[0009] Randomly sample the tampering mask of the image to be detected to obtain the coordinate positions of multiple pixel points and the corresponding tampering values;
[0010] Build a shallow convolutional layer module based on the DenseNet module to extract spatial domain features from the input face image and output spatial domain features;
[0011] Build a multi-level temporal domain feature extraction module based on ConvLSTM, input the spatial domain features into the multi-level temporal domain feature extraction module to output spatio-temporal features corresponding to different temporal resolutions;
[0012] Construct a dual attention mechanism module. The dual attention mechanism module is equipped with a spatio-temporal attention mechanism and a temporal resolution attention mechanism. In the spatio-temporal attention mechanism, generate a three-dimensional probability attention map to enhance the features of multiple spatio-temporal features. In the temporal resolution attention mechanism, input the spatio-temporal features after feature enhancement, calculate the temporal correlation between multiple spatio-temporal features, adaptively calculate the weight values of different spatio-temporal features, and weighted obtain spatio-temporal inconsistent features;
[0013] Concatenate the spatio-temporal inconsistent features with the coordinate channels of the pixel points, generate the tampering probability prediction value corresponding to the pixel coordinate position after passing through multiple convolutional layers, use the tampering value channel as the tampering value label, and calculate the cross-entropy loss function by comparing the pixel-level tampering probability prediction value and the tampering value label;
[0014] Set a threshold, and perform binary classification on the spatio-temporal inconsistent features based on the set threshold to obtain the face-swapping video tampering detection result.
[0015] As a preferred technical solution, the steps of constructing the tampering mask according to the facial key points specifically include:
[0016] Generate a convex polygon as the tampering area based on the point set of the facial key points to construct the tampering mask. The tampering mask is set with binary values of 0 or 1. When it is determined that the fake face image is tampered in the facial contour area, all values in the facial contour area are set to 1, and all values in the background area are set to 0; when it is determined that the real face image is not tampered in the whole image, the whole mask is set to 0;
[0017] Crop the face area image according to the face bounding box position, and use bilinear interpolation to resample the image into a unified preset resolution as the image input.
[0018] As a preferred technical solution, the steps of generating a three-dimensional probability attention map in the spatio-temporal attention mechanism to enhance the features of multiple spatio-temporal features specifically include:
[0019] In the spatio-temporal attention mechanism, two three-dimensional attention maps are obtained by using max-pooling and average-pooling operations on the spatio-temporal features along different input feature direction axes. After concatenating the two three-dimensional attention maps, a three-dimensional probability feature map is generated through 3D convolution and the Sigmoid function. The three-dimensional probability feature map is multiplied element-wise with the spatio-temporal features to obtain the enhanced spatio-temporal features.
[0020] As a preferred technical solution, generating a three-dimensional probability attention map in the spatio-temporal attention mechanism to enhance the features of multiple spatio-temporal features is specifically expressed as:
[0021]
[0022]
[0023]
[0024]
[0025] Among them, represents element-wise multiplication, σ represents the Sigmoid function, Conv3d represents the 3D convolutional layer, ST t,C,H,W represents the spatio-temporal feature input, represents the output value of max-pooling along the input feature direction, represents the output value of average-pooling along the input feature direction, represents the enhanced spatio-temporal features, t0 represents the t0-th frame of the image, N represents the total number of frames of the image, and C, H, and W respectively represent the channels, height, and width of the image.
[0026] As a preferred technical solution, inputting the enhanced spatio-temporal features in the temporal resolution attention mechanism, calculating the temporal correlation between multiple spatio-temporal features, and adaptively calculating the weight values of different spatio-temporal features, the specific steps include:
[0027] The enhanced spatio-temporal features go through three-dimensional max-pooling and three-dimensional average-pooling operations to obtain the max-pooling features and average-pooling features. The max-pooling features and average-pooling features are respectively input into a multi-layer perceptron to adaptively generate weight vectors and extract the temporal correlation of spatio-temporal features at different time resolutions.
[0028] As a preferred technical solution, inputting the enhanced spatio-temporal features in the temporal resolution attention mechanism, calculating the temporal correlation between multiple spatio-temporal features, and adaptively calculating the weight values of different spatio-temporal features, and weighted to obtain the spatio-temporal inconsistency features, which is specifically expressed as:
[0029]
[0030]
[0031]
[0032]
[0033] Among them, denotes element-wise multiplication, where C, H, and W represent the channels, height, and width of the image respectively. denotes the spatio-temporal feature after feature enhancement. denotes the max-pooling feature. denotes the average-pooling feature, and MLP denotes the multi-layer perceptron, F STW denotes the probability feature vector output by the Softmax function, ST STW denotes the spatio-temporal feature of multi-scale temporal correlation, where t0 represents the t0-th frame of the image and N represents the total number of frames of the image.
[0034] As a preferred technical solution, the cross-entropy loss function is calculated for the contrast pixel-level tampering probability prediction value and the tampering value label, and the specific calculation formula is expressed as:
[0035]
[0036] Among them, p′ represents the tampering probability prediction value, p represents the tampering value label, and p c represents the set of tampering value labels.
[0037] As a preferred technical solution, the spatio-temporal inconsistent features are binary-classified based on a set threshold, and the specific steps include:
[0038] The spatio-temporal inconsistent features are flattened into one-dimensional features and input into the classifier of two fully connected layers to output the tampering probability value of the frame to be detected. The tampering probability value is compared with the set threshold. If it is higher than the set threshold, the face image is determined to be a tampered image; if it is lower than the set threshold, the face image is determined to be a real image.
[0039] As a preferred technical solution, it further includes the step of constructing a network training loss function. The network training loss function includes the cross-entropy loss function L1 between the binary-classified true / false prediction and the true label, and the cross-entropy loss function L2 between the reconstructed pixel-level tampering probability prediction value and the tampering value label, which is specifically expressed as:
[0040] L1 = -(ylog(y′)+(1 - y)log(1 - y′))
[0041]
[0042]
[0043] Among them, p' represents the predicted value of the tampering probability, p represents the tampering value label, and p c represents the set of tampering value labels, y' represents the tampering probability value of the frame to be detected in the binary classification output, y represents the binary classification label, and L represents the network training loss function.
[0044] The present invention also provides a face-swapped video detection system for predicting pixel-level tampering probability values based on the spatio-temporal domain, including: a video preprocessing module, a tampering mask construction module, a coordinate position and tampering value pair construction module, a spatial domain feature extraction module, a multi-level temporal domain feature extraction module, a dual attention mechanism construction module, a pixel-level position tampering probability value reconstruction module, a binary classification module, and a detection result output module;
[0045] The video preprocessing module is used to frame the video to be tested and detect and extract the face bounding box;
[0046] The tampering mask construction module is used to construct a tampering mask according to the face key points;
[0047] The coordinate position and tampering value pair construction module is used to randomly sample the tampering mask of the image to be detected to obtain the coordinate positions of multiple pixel points and the corresponding tampering values;
[0048] The spatial domain feature extraction module is used to construct a shallow convolutional layer module based on the DenseNet module to extract spatial domain features from the input face image and output spatial domain features;
[0049] The multi-level temporal domain feature extraction module is used to construct a multi-level temporal domain feature extraction module based on ConvLSTM, input the spatial domain features into the multi-level temporal domain feature extraction module to output spatio-temporal features corresponding to different time resolutions;
[0050] The dual attention mechanism construction module is used to construct a dual attention mechanism module. The dual attention mechanism module is provided with a spatio-temporal attention mechanism and a temporal resolution attention mechanism. In the spatio-temporal attention mechanism, a three-dimensional probability attention map is generated to enhance the features of multiple spatio-temporal features. In the temporal resolution attention mechanism, the spatio-temporal features after feature enhancement are input, the temporal correlation between multiple spatio-temporal features is calculated, the weight values of different spatio-temporal features are adaptively calculated, and the spatio-temporal inconsistent features are obtained by weighted summation;
[0051] The pixel-level position tampering probability value reconstruction module is used to splice the spatio-temporal inconsistent features with the coordinate channels of the pixel points, generate the predicted tampering probability value corresponding to the pixel coordinate position after passing through multiple convolutional layers, use the tampering value channel as the tampering value label, and calculate the cross-entropy loss function by comparing the pixel-level tampering probability prediction value and the tampering value label;
[0052] The binary classification module is used to perform binary classification on the spatio-temporal inconsistency features based on a set threshold;
[0053] The detection result output module is used to output the face-swapping video tampering detection result after binary classification.
[0054] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0055] (1) The present invention uses the method of reconstructing the pixel-level position tampering probability value as auxiliary supervision. By sampling the coordinate information, the problem of high computational complexity in the previous method of generating the tampering probability map using the fully convolutional network model is solved, and the overfitting problem caused by the large-scale prediction map is solved, improving the domain generalization of the network.
[0056] (2) The present invention constructs a dual attention mechanism module and applies the attention module to the 3D convolutional network, which can better combine the spatial information and the temporal information with different temporal resolutions to improve the detection ability of the model, analyze the tampering information in a multi-scale manner from local frames to global frames, and make the network more focused on the most discriminative parts in the spatio-temporal domain. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 It is a framework schematic diagram of the face-swapping video detection method for predicting the pixel-level tampering probability value based on the spatio-temporal domain in this embodiment;
[0058] Figure 2 It is a schematic diagram of the training phase process of the face-swapping video detection method for predicting the pixel-level tampering probability value based on the spatio-temporal domain in this embodiment;
[0059] Figure 3 It is a schematic diagram of the network structure for extracting shallow spatial domain features in the face-swapping video detection method for predicting the pixel-level tampering probability value based on the spatio-temporal domain in this embodiment;
[0060] Figure 4 It is a schematic diagram of the network structure of the spatio-temporal attention mechanism in the dual attention module in this embodiment;
[0061] Figure 5 It is a schematic diagram of the network structure of the temporal resolution attention mechanism in the dual attention module in this embodiment;
[0062] Figure 6 It is a schematic diagram of the network structure for reconstructing the pixel-level position tampering probability value in this embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0063] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0064] Embodiment
[0065] This embodiment takes training on the FaceForensics++ (FF++) (C23) database, cross-testing different forgery methods on the FF++ (C23) database, and cross-database testing on the CelebDF, DeepfakeDetection (DFD), and DeepfakeDetection Challenge (DFDC) databases as examples. The FaceForensics++ database uses the H.264 encoder to synthesize videos with three different compression levels: compression rate 0 (C0), compression rate 23 (C23), and compression rate 40 (C40). Among them, there are 1000 real videos and 3000 face-swapped videos. The FaceForensics++ contains forged videos generated by four forgery methods, including Deepfakes (DF), NeuralTextures (NT), FaceSwap (FS), and Face2Face (F2F). The videos in the DeepfakeDetection database include uncompressed rate (C0), compression rate 23 (C23), and compression rate 40 (C40), among which there are 363 real videos and 3068 face-swapped videos.
[0066] First, divide the FF++ database into a training set, a validation set, and a test set in a ratio of 7:2:1. At the same time, to ensure the balance of the positive and negative sample ratios, the ratio of real video frames to face-swapped video frames in the selected dataset is ensured to be about 1:1. The experiment is carried out on the Ubuntu 18.04 system, using Python language version 3.8 and the Pytorch artificial neural network library version 1.7.0. The CUDA version is 11.0, and the cudnn version is 7.6.5.
[0067] As Figure 1 、 Figure 2 shown, this embodiment provides a face-swapped video detection method based on predicting the pixel-level tampering probability value in the spatio-temporal domain. It uses a multi-level spatio-temporal feature extraction module to extract spatio-temporal features with different temporal resolutions, and then uses a dual attention mechanism to extract temporal correlations to obtain spatio-temporal inconsistent features after feature enhancement. By splicing coordinate information (x coordinate, y coordinate) on the spatio-temporal inconsistent features, the tampering probability value of this coordinate point is accurately predicted for auxiliary supervision. Finally, the spatio-temporal inconsistent features are input into a binary classifier to obtain the tampering probability of the frame to be detected for true / false face judgment. The specific steps of the network training part include:
[0068] S1: Video preprocessing. The videos of each dataset are framed, saved as a frame sequence F, and the get_frontal_face_detector module in the Dlib library is used to detect and extract face bounding boxes for each frame image.
[0069] S2: Construct a tampering mask based on facial key points.
[0070] In this embodiment, for each video VX, on the frame sequence F i and the corresponding tampering mask, partial region pictures are cropped according to the face bounding box position and sampled to the same resolution to obtain training set pictures I = {I1, I2,..., I NF}, and the tampering masks of the training set images are Y = {Y1, Y2,..., Y NF};
[0071] In this embodiment, for each video VX in the training set, the video frames are framed to obtain F = {F1, F2,..., F NF}, where NF is the number of extracted video frames. The Dlib library is used to extract 68 key points of the face from each frame's face bounding box. A convex polygon is generated based on the 68 key point sets of the face as the tampering region (i.e., the face contour region) to construct the tampering mask. This tampering mask only has binary values of 0 or 1. It is determined that for a fake face image, only when the face contour region has been tampered with, the values within the face contour region are set to 1, and the background region is set to 0; it is determined that for a real face image, the entire image has not been tampered with, and the entire mask is set to 0;
[0072] In this embodiment, the cropped picture region is 1.3 times the detected face bounding box region. After cropping the original video frame image and the tampering mask, they are resampled to a resolution of 224×224 as the input image of the model of the present invention, and the sampling method is bilinear interpolation;
[0073] S3: Construct multiple groups of coordinate position - tampering value pairs;
[0074] In this embodiment, 14×14 pixel points are randomly sampled from the tampering mask of the picture to be detected, and their x - coordinates, y - coordinates, and tampering values are taken, that is, coordinate position - tampering value pairs of size 3×14×14 are obtained. The three channels represent the x - coordinate, y - coordinate, and tampering value respectively. Among them, the x and y coordinate positions will be used for splicing with spatio - temporal inconsistent features, and the tampering value is used as the label value for comparison with the tampering probability prediction value;
[0075] S4: Construct a multi - level spatio - temporal feature extraction module CNN - ConvLSTM. Input N consecutive images into the multi - level spatio - temporal feature extraction module CNN - ConvLSTM to obtain spatio - temporal features with different time resolutions, and obtain N spatio - temporal features;
[0076] In this embodiment, the cropped face image of the t0-th frame is selected and the subsequent N-1 consecutive frames of images are read to obtain N consecutive frames of face images It is judged whether t0+N-1≤NF. If not, the t0-th frame is skipped and the next video is re-detected
[0077] In this embodiment, a shallow convolutional layer module is constructed based on the DenseNet module to extract spatial domain features from the input N frames of face images, and N spatial domain features are output; a multi-level temporal domain feature extraction module is constructed based on ConvLSTM. This module inputs N spatial domain features and outputs N spatio-temporal features corresponding to different time resolutions
[0078] Input N consecutive frames, and extract N spatial domain features S through the shallow convolutional layer module t 。The N spatial domain features S t are respectively sent to the multi-level temporal domain feature extraction module. Each time a frame of spatial domain feature is input, a spatio-temporal feature ST with a different time resolution is output t There are N in total, and the specific calculation formula is
[0079] S t =Conv(I t ),t∈[t0,t0+N-1]
[0080]
[0081] Among them, where I t represents the cropped face image of the t-th frame, which is a six-color channel spliced by RGB and YUV. S t is the spatial domain feature after extraction. ST t represents the spatio-temporal features with different time resolutions, all of which have N. Conv represents the shallow convolutional layer module, and ConvLSTM represents the multi-level temporal domain feature extraction module
[0082] Such as Figure 3 shown, the input of the shallow spatial domain feature extraction network is the cropped face image with a size of 224×224, which is a six-color channel spliced by RGB and YUV. The shallow spatial domain feature extraction network includes three convolutional modules (the first convolutional module and the second convolutional module): the first convolutional module includes a 3×3 first convolutional layer with a 6-channel input and a stride of 1, a 7×7 second convolutional layer with a 3-channel input and a stride of 3, a BN layer, a Relu activation function, and a 2×2 max pooling layer with a stride of 2. The second convolutional module includes a first DenseBlock module with 6 convolutional layers, and the third convolutional module includes a second DenseBlock module with 12 convolutional layers
[0083] S5: Construct a dual attention mechanism module, input the N spatio-temporal features into the dual attention mechanism module to extract the three-dimensional attention map and the temporal correlation weight value, and obtain the spatio-temporal inconsistency feature;
[0084] In this embodiment, the input is N spatio-temporal features, which pass through two attention mechanisms. The first attention mechanism is the spatio-temporal attention mechanism, and the second attention mechanism is the temporal resolution attention mechanism.
[0085] As Figure 4 shown, in the spatio-temporal attention mechanism, the N spatio-temporal features obtained in step S4 Use max pooling and average pooling operations along different input feature direction axes to obtain two three-dimensional attention maps: and Then concatenate the two and send them into a 7×7×7 3D convolution and the Sigmoid function to finally generate a three-dimensional probability feature map The three-dimensional probability feature map generated after concatenation performs feature enhancement on each spatio-temporal feature, suppresses useless information, and multiplies it element-wise with the N input spatio-temporal feature maps ST t to obtain The specific calculation formula is:
[0086]
[0087]
[0088]
[0089]
[0090] where represents element-wise multiplication, σ represents the Sigmoid function, Conv3d represents the 3D convolutional layer, preferably a 3D convolutional layer with a convolutional kernel of 7×7×7, represents the input of N spatio-temporal features, represents the output value of max pooling along the input feature direction, represents the output value of average pooling along the input feature direction, that is, represents two three-dimensional spatial attention maps, represents the spatio-temporal features after feature enhancement, with a total of N.
[0091] The second attention mechanism module is the time-domain resolution attention mechanism. For each spatio-temporal feature map, three-dimensional max pooling and three-dimensional average pooling operations are used, and through a multi-layer perceptron (MLP), weight vectors can be adaptively generated for multi-level spatio-temporal features, extracting the time-domain correlations of spatio-temporal features at different time resolutions, and capturing tampering information multi-scale from the globally adjacent frames to the locally adjacent frames in the time domain.
[0092] As Figure 5 shown, in the time-domain resolution attention mechanism, N spatio-temporal feature maps are input After two three-dimensional pooling operations, we respectively obtain and represent the max pooling feature and the average pooling feature respectively, and are respectively fed into an MLP with a hidden layer for summation. To reduce the parameter overhead, two 3D convolutions are used instead of fully connected layers in the MLP, and the dimension size of the hidden layer is set to where r is the expansion rate, and then the probability feature vector is output by Softmax Multiplied by T spatio-temporal features and summed to obtain the spatio-temporal features of multi-scale time-domain correlation The specific calculation formula is as follows:
[0093]
[0094]
[0095]
[0096]
[0097] Among them, and represent the weight values after three-dimensional max pooling and three-dimensional average pooling of the entire feature map, represents the weight vector after multi-scale time-domain correlation, is the spatio-temporal inconsistency feature after calculating the sum of the products of multiple spatio-temporal features and the weight vector.
[0098] S6: Construct a pixel-level position tampering probability value reconstruction for auxiliary supervision;
[0099] After concatenating the spatio-temporal inconsistent features (256×14×14) after the dual attention mechanism module with the x and y coordinate two channels (2×14×14) in the coordinate position-tampering value pairs sampled in step S3, input them into two layers of 1×1 convolutional layers to obtain the tampering probability values predicted according to the pixel coordinate positions. Take the tampering value channel in the coordinate position-tampering value pairs as the label, calculate the cross-entropy loss function L2 for the predicted tampering probability values and the tampering value labels corresponding to the pixel coordinates for auxiliary supervision. The loss function formula is as follows:
[0100]
[0101] Among them, p′ represents the predicted tampering probability value, and p represents the tampering value label.
[0102] As Figure 6 shown, concatenate the spatio-temporal inconsistent features and the coordinate positions (x coordinate, y coordinate), and input them into two layers of convolutional layers to calculate and output the predicted tampering probability values of the corresponding pixels. Among them, the convolutional layers both use 1×1 convolutional kernels with a stride of 1.
[0103] S7: Perform binary classification on the spatio-temporal inconsistent features to obtain the detection result;
[0104] In this embodiment, the spatio-temporal inconsistent features are flattened into one-dimensional features and input into the classifier of two layers of fully connected layers to output the tampering probability value of a frame to be detected, and distinguish between real and fake faces. Set the threshold to 0.5. If it is higher than the threshold, it is determined that the face image is a tampered image, otherwise it is a real image. The output of the first layer of fully connected layer is 1×64, and the output of the second layer of fully connected layer is 1×1;
[0105] S8: Construct the loss function for network training;
[0106] In this embodiment, the loss function includes two parts: the cross-entropy loss function L1 of the binary classification true and false prediction and the real label, and the cross-entropy loss function L2 of the pixel-level tampering probability prediction value reconstructed and the tampering value label. In order to balance L1 and L2, during the training process, the weight values are dynamically allocated according to their sizes to calculate the total loss function L. The specific calculation formula is:
[0107] L1 = -(ylog(y′)+(1 - y)log(1 - y′))
[0108]
[0109]
[0110] Among them, p′ represents the predicted tampering probability value, p represents the tampering value label, p c represents the set of tampering value labels, y′ represents the tampering probability of the frame to be detected, and y represents the binary classification label.
[0111] S9: Set the network parameter optimization algorithm;
[0112] In this embodiment, the Adam algorithm is used for parameter optimization, and the learning rate is set to 1x10 -4 , the first-order smoothing parameter β1 = 0.9, the second-order smoothing parameter β2 = 0.999, and the constant e to prevent the denominator from being 0 is 1x10 -8 , of course, other gradient optimization algorithms such as SGD and RMSprop can be used for the optimization algorithm; the training period is 30, and the training batch size is 32.
[0113] S10: Train the network;
[0114] S11; Application of the model tested between different forgery methods;
[0115] In this embodiment, the model structure and parameters saved in the model training step are loaded as the background module of the detection system; for each video in the test set, 10 consecutive frames are selected, 10 single-frame features are extracted, input into the detection system, and the prediction classification result is obtained.
[0116] FaceForensics++ is the most widely used database among many Deepfake detection methods. It contains 1000 original real videos from the Internet, and each real video corresponds to forged videos generated by 4 forgery methods, including four forgery methods: Deepfakes (DF), NeuralTextures (NT), FaceSwap (FS), and Face2Face (F2F).
[0117] The present invention loads the model and weights of the network trained using the training set of one of the forgery methods in the FF++ database, and then uses the test set of the database of the four forgery methods in FF++ for testing.
[0118] As shown in Table 1 below, the test results between different forgery methods are obtained;
[0119] Table 1 Test results between different forgery methods
[0120]
[0121] S12: Application of the model tested from FF++ to other databases;
[0122] In this embodiment, the model structure and parameters saved in the model training step are loaded as the background module of the detection system; for each video in the test set, 10 consecutive frames are selected, 10 single-frame features are extracted, input into the detection system, and the prediction classification result is obtained.
[0123] The present invention loads the model and weights of the network trained using the training set of the FF++ (HQ) database, then uses the test sets of the DFD, DFDC, and Celeb-DF databases for cross-database testing, and compares the results of the cross-database testing with the latest detection algorithms.
[0124] As shown in Table 2 below, the test results from FF++ to other databases are obtained:
[0125] Table 2 Test Results from FF++ to Other Databases
[0126]
[0127] This embodiment also provides a face-swapped video detection system for predicting the pixel-level tampering probability value based on the spatio-temporal domain, including: a video preprocessing module, a tampering mask construction module, a coordinate position and tampering value pair construction module, a spatial domain feature extraction module, a multi-level temporal domain feature extraction module, a dual attention mechanism construction module, a pixel-level position tampering probability value reconstruction module, a binary classification module, and a detection result output module;
[0128] In this embodiment, the video preprocessing module is used to frame the video to be tested and detect and extract the face bounding boxes;
[0129] In this embodiment, the tampering mask construction module is used to construct a tampering mask according to the face key points. Specifically, 68 key points of the face are detected and extracted from the face bounding box, and a convex polygon (i.e., the face contour) is generated according to the key point set as the tampering area to construct the tampering mask. The tampering mask has only two values, 0 or 1. It is considered that the fake face is tampered in the convex polygon area and set to 1, and the background is not tampered and set to 0; while the real face is not tampered in the whole image and set to 0;
[0130] In this embodiment, the coordinate position and tampering value pair construction module is used to randomly sample the coordinate positions (x coordinate, y coordinate) of multiple pixels and the corresponding tampering values from the tampering mask to form a coordinate position and tampering value pair, which is used to assist in pixel-level position tampering value reconstruction during supervision. Among them, the x and y coordinate positions will be used for splicing with the spatio-temporal inconsistent features, and the tampering value is used as the label value for comparison with the tampering probability prediction value;
[0131] In this embodiment, the spatial domain feature extraction module is used to construct a shallow convolutional layer module based on the DenseNet module to extract the spatial domain features of the input face image and output the spatial domain features;
[0132] In this embodiment, the multi-level temporal domain feature extraction module is used to construct a multi-level temporal domain feature extraction module based on ConvLSTM, and input the spatial domain features into the multi-level temporal domain feature extraction module to output spatio-temporal features corresponding to different time resolutions;
[0133] In this embodiment, the dual attention mechanism construction module is used to construct a dual attention mechanism module. The dual attention mechanism module is provided with a spatio-temporal attention mechanism and a temporal resolution attention mechanism. In the spatio-temporal attention mechanism, a three-dimensional probability attention map is generated to enhance the features of multiple spatio-temporal features. In the temporal resolution attention mechanism, the spatio-temporal features after feature enhancement are input, the temporal correlation between multiple spatio-temporal features is calculated, the weight values of different spatio-temporal features are adaptively calculated, and the spatio-temporal inconsistent features are obtained by weighted summation.
[0134] In this embodiment, the pixel-level position tampering probability value reconstruction module is used to splice the finally generated spatio-temporal inconsistent features with the x and y coordinate channels of the pixel points in the coordinate position-tampering value pair. After passing through two 1×1 convolutional layers, the tampering probability prediction value corresponding to the pixel coordinate position is generated. The tampering value channel in the coordinate position-tampering value pair is used as a label, and the cross-entropy loss function is calculated by comparing the reconstructed pixel-level tampering probability prediction value and the tampering value label.
[0135] In this embodiment, the binary classification module is used to perform binary classification on the spatio-temporal inconsistent features based on a set threshold, and the detection result output module is used to output the binary classification result.
[0136] The above embodiments are preferred embodiments of the present invention. However, the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. A face-swapping video detection method based on predicting pixel-level tampering probability values in the spatio-temporal domain, characterized in that, Including the following steps: Frame the video to be tested and detect and extract face bounding boxes for each frame of the image; Construct a tampering mask based on face key points; Randomly sample the tampering mask of the image to be detected to obtain the coordinate positions of multiple pixel points and the corresponding tampering values; Construct a shallow convolutional layer module based on the DenseNet module to extract spatial domain features from the input face image and output spatial domain features; Construct a multi-level temporal domain feature extraction module based on ConvLSTM, input the spatial domain features into the multi-level temporal domain feature extraction module to output spatio-temporal features corresponding to different temporal resolutions; Construct a dual attention mechanism module, the dual attention mechanism module is provided with a spatio-temporal attention mechanism and a temporal resolution attention mechanism. In the spatio-temporal attention mechanism, generate a three-dimensional probability attention map to enhance the features of multiple spatio-temporal features. In the temporal resolution attention mechanism, input the spatio-temporal features after feature enhancement, calculate the temporal correlation between multiple spatio-temporal features, adaptively calculate the weight values of different spatio-temporal features, and obtain spatio-temporal inconsistency features by weighted summation; Concatenate the spatio-temporal inconsistency features with the coordinate channels of the pixel points, generate a tampering probability prediction value corresponding to the pixel coordinate position after passing through multiple convolutional layers, use the tampering value channel as the tampering value label, and calculate the cross-entropy loss function by comparing the pixel-level tampering probability prediction value and the tampering value label; Set a threshold, and perform binary classification on the spatio-temporal inconsistency features based on the set threshold to obtain the face-swapping video tampering detection result.
2. The face-swapping video detection method based on predicting pixel-level tampering probability values in the spatio-temporal domain according to claim 1, wherein The specific steps of constructing the tampering mask according to the face key points include: Generate a convex polygon as the tampering area based on the point set of the face key points to construct the tampering mask. The tampering mask is provided with binary values of 0 or 1. When it is determined that the fake face image is tampered in the face contour area, all values in the face contour area are set to 1, and all background areas are set to 0. When it is determined that the real face image is not tampered in the whole image, the whole mask is set to 0; Crop the face region image according to the face bounding box position, and use bilinear interpolation to resample the image into a unified preset resolution as the image input.
3. The face-swapping video detection method based on predicting pixel-level tampering probability values in the spatio-temporal domain according to claim 1, characterized in that, The specific steps of generating a three-dimensional probability attention map in the spatio-temporal attention mechanism to enhance the features of multiple spatio-temporal features include: In the spatio-temporal attention mechanism, use max pooling and average pooling operations on the spatio-temporal features along different input feature direction axes to obtain two three-dimensional attention maps. After concatenating the two three-dimensional attention maps, generate a three-dimensional probability feature map through 3D convolution and the Sigmoid function, and multiply the three-dimensional probability feature map element-wise with the spatio-temporal features to obtain the spatio-temporal features after feature enhancement.
4. The face-swapping video detection method based on predicting pixel-level tampering probability values in the spatio-temporal domain according to claim 1 or 3, characterized in that, The generation of a three-dimensional probability attention map in the spatio-temporal attention mechanism to enhance the features of multiple spatio-temporal features is specifically expressed as: Among them, denotes element-wise multiplication, σ denotes the Sigmoid function, Conv3d denotes the 3D convolutional layer, and ST t,C,H,W denotes the spatio-temporal feature input, denotes the maximum pooling output value along the input feature direction, denotes the average pooling output value along the input feature direction, denotes the spatio-temporal feature after feature enhancement, t0 denotes the t0-th frame of the image, N denotes the total number of frames of the image, and C, H, and W denote the channels, height, and width of the image, respectively.
5. The face-swapping video detection method based on predicting pixel-level tampering probability values in the spatio-temporal domain according to claim 1, wherein Input the spatio-temporal features after feature enhancement into the temporal resolution attention mechanism, calculate the temporal correlation between multiple spatio-temporal features, and adaptively calculate the weight values of different spatio-temporal features. The specific steps include: After the spatio-temporal features enhanced by features are subjected to three-dimensional max pooling and three-dimensional average pooling operations, max pooling features and average pooling features are obtained. The max pooling features and average pooling features are respectively input into a multi-layer perceptron to adaptively generate a weight vector and extract the temporal domain correlation of the spatio-temporal features at different time resolutions.
6. The face-swapping video detection method based on predicting pixel-level tampering probability values in the spatio-temporal domain according to claim 1 or 5, characterized in that, In the temporal resolution attention mechanism, the spatio-temporal features after feature enhancement are input, the temporal domain correlation between multiple spatio-temporal features is calculated, the weight values of different spatio-temporal features are adaptively calculated, and the spatio-temporal inconsistent features are obtained by weighting, which is specifically expressed as: Among them, represents element-wise multiplication, where C, H, and W represent the number of channels, height, and width of the image respectively. represents the spatio-temporal feature after feature enhancement. represents the max-pooling feature. represents the average-pooling feature, MLP represents the multi-layer perceptron, and F STW represents the probability feature vector output by the Softmax function, and ST STW represents the spatio-temporal feature of multi-scale temporal correlation, t0 represents the t0-th frame of the image, and N represents the total number of frames of the image.
7. The face-swapping video detection method for predicting pixel-level tampering probability values based on spatio-temporal domain according to claim 1, wherein The cross-entropy loss function is calculated by comparing the predicted value of the pixel-level tampering probability with the tampering value label, and the specific calculation formula is expressed as: Among them, p' represents the predicted value of the tampering probability, p represents the tampering value label, and p c represents the set of tampering value labels.
8. The face-swapping video detection method based on predicting pixel-level tampering probability values in the spatio-temporal domain according to claim 1, wherein, Based on a set threshold, the spatio-temporal inconsistent features are classified into two categories, and the specific steps include: The spatio-temporal inconsistent features are flattened into one-dimensional features and input into the classifier of the two-layer fully connected layer to output the tampering probability value of the frame to be detected. The tampering probability value is compared with the set threshold. If it is higher than the set threshold, the face image is determined to be a tampered image; if it is lower than the set threshold, the face image is determined to be a real image.
9. The face-swapping video detection method for predicting pixel-level tampering probability values based on spatio-temporal domain according to claim 1, characterized in that It also includes the step of constructing the network training loss function. The network training loss function includes the cross-entropy loss function L1 between the binary classification true / false prediction and the true label, and the cross-entropy loss function L2 between the predicted value of the pixel-level tampering probability reconstructed and the tampering value label, which is specifically expressed as: L1 = -(y log(y')+(1 - y) log(1 - y')) Among them, p' represents the predicted value of the tampering probability, p represents the tampering value label, p c represents the set of tampering value labels, y' represents the tampering probability value of the frame to be detected in the binary classification output, y represents the binary classification label, and L represents the network training loss function.
10. A face-swapping video detection system based on predicting pixel-level tampering probability values in the spatio-temporal domain, characterized in that, including: a video preprocessing module, a tampering mask construction module, a coordinate position and tampering value pair construction module, a spatial domain feature extraction module, a multi-level temporal domain feature extraction module, a dual attention mechanism construction module, a pixel-level position tampering probability value reconstruction module, a binary classification module, and a detection result output module; The video preprocessing module is used to frame the video to be measured and detect and extract the face frame; The tampering mask construction module is used to construct a tampering mask according to the face key points; The coordinate position and tampering value pair construction module is used to randomly sample the tampering mask of the image to be detected to obtain the coordinate positions of multiple pixel points and the corresponding tampering values; The spatial domain feature extraction module is used to construct a shallow convolutional layer module based on the DenseNet module, extract the spatial domain features of the input face image, and output the spatial domain features; The multi-level temporal domain feature extraction module is used to construct a multi-level temporal domain feature extraction module based on ConvLSTM. The spatial domain features are input into the multi-level temporal domain feature extraction module to output spatio-temporal features corresponding to different time resolutions; The dual attention mechanism construction module is used to construct a dual attention mechanism module. The dual attention mechanism module is provided with a spatio-temporal attention mechanism and a temporal resolution attention mechanism. In the spatio-temporal attention mechanism, a three-dimensional probability attention map is generated to enhance the features of multiple spatio-temporal features. In the temporal resolution attention mechanism, the spatio-temporal features after feature enhancement are input, the temporal domain correlation between multiple spatio-temporal features is calculated, the weight values of different spatio-temporal features are adaptively calculated, and the spatio-temporal inconsistent features are obtained by weighting; The pixel-level position tampering probability value reconstruction module is used to splice the spatio-temporal inconsistency features with the coordinate channels of the pixel points, generate the tampering probability prediction values corresponding to the pixel coordinate positions after passing through multiple convolutional layers, use the tampering value channel as the tampering value label, and calculate the cross-entropy loss function by comparing the pixel-level tampering probability prediction values and the tampering value labels; The binary classification module is used to perform binary classification on the spatio-temporal inconsistency features based on a set threshold; The detection result output module is used to output the face-swapping video tampering detection result after binary classification.
Citation Information
Patent Citations
Method and system for passively detecting AI face-changing video based on joint features
CN111539272A
Face depth tampered image detection method based on multi-scale depth feature fusion
CN111539942A