A Deepfake Video Detection Method Based on Semi-Supervised Learning
Through the deep fake video detection method based on semi-supervised learning, the negative sample generator and data augmentation module are used to train a lightweight semi-supervised classifier, which solves the problem of relying on a large amount of labeled data in the existing technology, and realizes efficient and robust deep fake video detection.
Patent Information
- Application Number
- CN202310362080.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-30
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2043-03-30
AI Technical Summary
Existing deep fake video detection methods rely on a large amount of labeled data, and high-quality data collection requires a lot of resources, which affects detection performance, especially in the face of insufficient negative samples or unknown fake generation technology.
A deep fake video detection method based on semi-supervised learning is proposed. A negative sample generator is used to generate a balanced number of negative samples, and combined with a data augmentation module and a lightweight semi-supervised classifier, the classifier is trained to complete the authenticity identification task.
It realizes the more robust feature representation while reducing the consumption of data collection resources, improves the generalization performance of the detection model, and is suitable for practical application scenarios of insufficient negative samples or unknown forgery generation technology.
Smart Images

Figure CN116824430B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a deepfake video detection method based on semi-supervised learning, belonging to the technical field of deepfake video forensics processing. Background Art
[0002] In recent years, face-swapping technologies represented by Deepfake have begun to appear in various types of digital images and videos. The so-called deepfake technology comes from the combination of deep-learning and fake. Such technologies manipulate image, video, and audio content based on deep learning or graphics methods. Based on this technology, it is possible to replace the identity of the target person and the original person in digital media, or make the target person perform some specific actions in a set scene, such as changing the mouth movement and expression of the person, so as to achieve the purpose of confusing the public. While the Deepfake technology is used in positive applications such as entertainment and film and television, the threat of being manipulated and abused by malicious users has gradually emerged. From large-scale political smear, election interference, and military deception between countries to small-scale video fraud, false news, and the production of pornographic movies, the abuse of deepfake technology has brought many risks to various aspects such as national politics, military security, social security, judicial forensics, and public opinion guidance. Therefore, how to effectively detect deepfake media has broad application prospects and urgent research necessity.
[0003] However, currently mainstream detection methods basically rely on a large number of high-quality positive and negative labeled samples, and train a classifier through a fully supervised method to complete the detection task. However, in real application scenarios, high-quality labeled information costs a high price to collect, and the number of forged samples is far less than that of real samples. The unbalanced positive and negative sample ratio will also affect the detection performance of the detector. Therefore, in order to reduce the resource consumption in the process of collecting data and learn more robust discriminative feature representations of Deepfake videos at a lower cost.
[0004] The generation of deepfake videos can generally be divided into four core steps: target face image extraction, model training, face-swapped image generation, and post-processing. In the target face image extraction stage, face images of the target person are collected, and these images should contain as rich expressions as possible, such as head postures like opening the mouth, closing the mouth, raising the head, lowering the head, opening the eyes, and closing the eyes. In the model training stage, generative adversarial networks (GANs) and variational autoencoders (VAEs) are generally used to train the model. The encoder is used to analyze and extract the latent feature information of the image, share parameter information, and continuously update the model parameters. In the image generation stage, the decoder is often used to reconstruct the face image to obtain the face-swapped image. In the final post-processing stage, the generated face image is restored to the original frame through affine transformation. There are often obvious splicing boundaries in this process, so post-processing operations such as sharpening and blurring are required. Finally, the generated face-swapped frame is converted into a forged video.
[0005] Correspondingly, based on the research and analysis of the generation mechanism of deepfake videos, deepfake video detection methods are mainly divided into the following three types: deepfake video detection based on inconsistent biological signals; deepfake video detection based on data-driven; deepfake video detection based on traditional manually extracted features. Detection methods based on biological signals mainly utilize the inconsistencies in biological signals such as blink frequency, head posture, eye color, and heart rate existing in early deepfake videos to identify the authenticity of videos. For example, LSTM is used to analyze the blink frequency of the person in the video; relevant features of facial expressions and head movements in two types of videos are extracted to train a binary classification SVM; facial key points of the human face are extracted, and then these facial region markers are used as discriminant feature vectors, etc. Detection methods based on data-driven mainly utilize the powerful learning ability of deep neural networks to analyze a large number of real and deepfake videos, and extract features to train a detection model. For example, CapsuleNet is trained to extract discriminant features in forged images for detection by using the difference in the positions of the facial features of the person in the deepfake video and the real video; XceptionNet is trained using the Faceforensic++ dataset to detect the tampering traces existing in forged videos; the more general "Face X-ray" is used to extract the tampering traces brought about by the post-processing operations in the synthesis process of deepfake videos, etc. Detection methods based on traditional manual features mainly analyze the differential image features between genuine and fake video frames. For example, PRNU is used to analyze the difference in the average normalized correlation coefficient between real videos and deepfake videos; feature vectors extracted by different feature point detectors are used as discriminant features.
[0006] First, for the deepfake video detection method based on inconsistent biosignals, as the generation technology level continues to improve, deepfake content becomes more and more realistic, so the detection effect of this type of method is greatly reduced; the data-driven deepfake video detection method relies on a large amount of labeled data and often performs well on data from a unified dataset source, but has poor model generalization when facing unknown forgery generation technologies, and collecting high-quality data requires a large amount of resources; when the traditional handcrafted feature-based deepfake video detection method faces post-processing operations such as cropping, compression, blurring, and noise on digital content, the original image features are damaged, and the traditional feature extractor has poor effect in extracting discriminative features, and the discrimination ability of the detection model will be greatly reduced.
[0007] For these general deepfake detection methods, they mostly start from the defects generated in the process of generating a new face and a part of the surrounding area by the AI face swap algorithm, and determine the authenticity of the video to be detected by detecting these tampering traces existing in the image or between video frames. However, the tampering traces in deepfake video synthesis are not uniquely determined, and the defects caused by different face swap algorithms in synthesis vary greatly. Therefore, simply through a data-driven approach, the classification model trained relying on a large number of training samples does not have good generalization performance, and even has the problem of overfitting on data from the same dataset source. At the same time, as the deepfake algorithm is gradually optimized, these defects are difficult to be discovered and even no longer exist. In addition, in actual application scenarios, collecting high-quality fake videos requires a large amount of manpower and material resources, and the number of fake videos is far less than that of real videos.
[0008] Therefore, in order to obtain a more robust feature representation, while reducing the consumption of resources such as computing power and manpower brought by data collection, and keeping the balance of positive and negative samples, this method proposes a deepfake video detection method based on semi-supervised learning, which only needs to provide a certain number of real samples, and a negative sample generator constructed by a simple graphics method is used to generate the same number of negative samples. The positive and negative samples together form the data for training a semi-supervised classifier. The semi-supervised classifier learns a robust feature representation by mining the information existing in the spatial and temporal domains of deepfake content, so as to realize the discrimination of real and fake videos. Summary of the Invention
[0009] In order to overcome the deficiencies of existing research, the present invention proposes a semi-supervised deepfake video detection method. The input of this scheme only requires real samples, and a negative sample generator is proposed to generate an equal number of negative samples, and a uniquely designed semi-supervised classifier is trained to complete the final authenticity discrimination task.
[0010] The generation of forged videos must go through the final image fusion process. Therefore, the present invention utilizes the characteristic that fixed tampering artifacts will be generated in this stage, and adopts a simpler and more effective graphics method to generate negative samples to create general forensic "clues". The graphics-based face replacement steps mainly include: face key point extraction, nearest search, affine transformation, and post-processing. And in order to make the generated negative samples more realistic, the idea of GAN is borrowed, and a discriminator is introduced to discriminate the generated negative samples, and iteratively generate negative samples that are closer to real samples. Finally, all positive and negative samples go through a data augmentation module, including five augmentation forms: random cropping, random grayscaling, random erasing, random flipping, and random variation, to obtain more diverse data for the training of the downstream semi-supervised classifier. In the design of the semi-supervised classifier, the present invention designs a more lightweight feature extraction network based on XceptionNet, combines RGB images, low-level spatial noise maps, and inter-frame correlated temporal noise maps, and mines forgery traces from multiple scales. The uniquely designed semi-supervised classifier is mainly responsible for solving the situation where there are insufficient or no labeled negative samples, and is more suitable for actual deepfake detection scenarios.
[0011] A deepfake video detection method based on semi-supervised learning realizes the identification of genuine and fake videos through three steps: negative sample generation, data augmentation, and semi-supervised classifier construction;
[0012] The negative sample generation includes constructing a real video set, obtaining a real face set through preprocessing, randomly sampling a real face from the real face set, and extracting the key point information of the face; calculating the distances between the key points of this face and the key points of the remaining faces in the set, finding the face closest to this face, and using it as the foreground face; the sampled real face is used as the background face; performing an affine transformation on the foreground face; fitting the processed foreground face on the background face, and performing color correction, Gaussian blur, and edge smoothing processing to obtain a synthesized forged face image;
[0013] The generated forged face and real face are sent to the discriminator for judgment, and the generated negative samples are fine-tuned according to the correct or incorrect judgment of the discriminator. If the discriminator makes a correct judgment, it means that the generated negative samples are not realistic enough, and random sampling needs to be performed again to generate negative samples. Iterate like this until the discriminator cannot make a correct judgment. Its objective function can be expressed by the following formula:
[0014]
[0015] where E(*) represents the expected value of the distribution function, P real (x) represents the distribution of real face samples, P pseudo-fake(z) represents the distribution of forged face samples after data augmentation, D(x) represents the discriminator's discrimination process, where x is a real face sample, G(z) represents the negative sample generation process, and z is the face to be tampered with randomly sampled from the real face set;
[0016] The data augmentation includes a data augmentation combination for deepfake video detection, including five augmentation forms: erasing, cropping, flipping, block recombination, and color jitter;
[0017] The construction of the semi-supervised classifier includes a training stage and a testing stage. In the training stage, only real samples are used as the input of the negative sample generator. After the negative sample generation process, a large number of forged samples are generated. The real samples with balanced quantity and the generated forged samples are jointly input into the data augmentation module to obtain diversified augmented data. The augmented data is used as training data and input into the semi-supervised classifier for binary classification training.
[0018] The data augmentation combination includes
[0019] Erasing: A local area of the face image is randomly erased, that is, the pixel values of the randomly selected area are set to 0;
[0020] Cropping: A part of the facial image is randomly cropped, and then the cropped area is rescaled to the original size;
[0021] Flipping: Horizontally flip the original face image;
[0022] Block recombination: Cut the original face into blocks of the same size, shuffle the order and recombine them;
[0023] Color jitter: Random changes in the attributes of the face image, including brightness, contrast, saturation, and hue.
[0024] The semi-supervised classifier adopts a three-layer flow structure design, including an entrance layer, eight intermediate layers, and an exit layer, to summarize and organize features and hand them over to the fully connected layer for expression:
[0025] Before being fed into the entrance layer, the processed image is normalized RGB image, the noise map in the spatial domain, and the noise map in the temporal domain. For an image, the spatial domain noise map refers to a Gaussian filter with a kernel size of 5;
[0026] The generation of the temporal domain noise map goes through six steps: spatial Gaussian filtering; per-pixel temporal high-pass filtering; batch normalization; suppressing the amplitude less than the threshold t; calculating the temporal gradient; temporal low-pass filtering;
[0027] After convolution, pooling and other operations in the three layers of input, middle and output, the logistic regression value for predicting the true or false category is finally output, and is trained by backpropagation with cross-entropy loss. The definition of the cross-entropy loss function is as follows:
[0028]
[0029] y i represents the label of sample i, where the forged class is 1 and the real class is 0, and p i represents the logistic regression value of sample i predicted as the forged class.
[0030] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0031] Based on the semi-supervised learning method, the present invention realizes the detection of deepfake videos only by relying on real videos as input. The negative sample generator proposed in the invention effectively simulates the tampering traces left in the last step of deepfake video synthesis. The data augmentation module emphasizes the differential information between genuine and forged images, and accelerates the learning process of the detection model from the data level. The lightweight semi-supervised classifier combines RGB images, low-level spatial noise maps, and inter-frame correlated temporal noise maps to find forensic clues from multiple scales. The three complement each other and jointly form an efficient and robust semi-supervised Deepfake video detection framework.
[0032] The present invention proposes a semi-supervised method that can complete the deepfake video forensics task only by relying on real videos, which is applicable to the situation where there are insufficient or no negative samples in actual application scenarios. Through the research on the generation mechanism of deepfake videos, more general tampering traces in the forgery synthesis process are discovered. According to this discovery, a simpler and more effective forged sample generation method is proposed, saving the cost of collecting and generating forged data. The designed data augmentation module effectively increases the diversity of data, while enhancing the representation of tampering traces at the image level, emphasizing the differential information between real images and forged images, which is beneficial to the fast and accurate learning of the detector. A more lightweight feature extraction network is designed based on XceptionNet, which combines RGB images, low-level spatial noise maps, and inter-frame correlated temporal noise maps to mine tampering traces from multiple scales and learn more robust feature representations. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0034] Figure 1 This is a flowchart of a deepfake video detection method based on semi-supervised learning according to the present invention;
[0035] Figure 2 This is an architecture diagram of a semi-supervised classifier for a deepfake video detection method based on semi-supervised learning according to the present invention. Specific embodiments
[0036] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0037] A deepfake video detection method based on semi-supervised learning includes three steps: negative sample generation, data augmentation, and semi-supervised classifier construction;
[0038] The negative sample generation includes constructing a real video set, obtaining a real face set through preprocessing, randomly sampling a real face from the real face set, and extracting the key point information of the face; calculating the distances between the key points of the face and the key points of the remaining faces in the set, finding the face closest to this face as the foreground face; the sampled real face is used as the background face; performing an affine transformation on the foreground face; pasting the processed foreground face on the background face, and performing color correction, Gaussian blur, and edge smoothing processing to obtain a synthesized forged face image; the original data only requires real person videos,
[0039] In the discriminative fine-tuning stage: The generated forged faces and real faces are sent to the discriminator for judgment. According to whether the discriminator's judgment is correct, the generated negative samples are fine-tuned. If the discriminator discriminates correctly, it means that the generated negative samples are not realistic enough, and random sampling needs to be performed again to generate negative samples. Iterate like this until the discriminator cannot make a correct judgment. Its objective function can be expressed by the following formula:
[0040]
[0041] where E(*) represents the expected value of the distribution function, P real (x) represents the distribution of real face samples, P pseudo-fake (z) represents the distribution of forged face samples after data augmentation, D(x) represents the discriminator's discrimination process, x is a real face sample, G(z) represents the negative sample generation process, and z is the face to be tampered with randomly sampled from the real face set;
[0042] In the process of deepfake video synthesis, regardless of the face-swapping algorithm used, the generated new face must be fused into the original image to replace the facial area in the original image. Therefore, the present invention utilizes the characteristics of this step to simulate the process of replacing a face with the source face in a simpler and more effective graphics method. By this means, a large number of negative samples are generated. And to make the generated negative samples more realistic, the present invention introduces a discriminator to discriminate the generated negative samples, compares the synthesized face with the original face, and iteratively generates negative samples closer to real samples.
[0043] The data augmentation includes:
[0044] In the image classification task, data augmentation of image data is a commonly used regularization method, mainly used to increase the amount of training data and make the dataset as diverse as possible, so that the trained model has stronger generalization ability, and is commonly used in scenarios where the amount of data is insufficient or there are many model parameters. In the Deepfake forensics task, the purpose of data augmentation is not only to increase the diversity of data, but also to enhance the representation of forgery traces at the image level and emphasize the differential information between real images and forged images. Therefore, it is necessary to destroy the high-level semantic features of the face to a certain extent. Therefore, general data augmentation methods cannot be directly applied to handle Deepfake detection problems.
[0045] The present invention designs a data augmentation combination for deepfake video detection according to the uniqueness of the target task. The combination includes five augmentation forms. 1) Erasing: A local area of the face image is randomly erased, that is, the pixel values of the randomly selected area are set to 0. 2) Cropping: A part of the facial image is randomly cropped, and then the cropped area is rescaled to the original size. 3) Flipping: The original face image is flipped horizontally. 3) Block recombination: The original face is cut into blocks of the same size, shuffled and recombined; 4) Color jitter: Random changes in the attributes of the face image, including brightness, contrast, saturation and hue.
[0046] The present invention divides the construction of the semi-supervised classifier into two stages, the training stage and the testing stage. In the training stage, only real samples are used as the input of the negative sample generator in (1). After the negative sample generation process described in (1), a large number of forged samples are generated. Then, an equal number of real samples and the generated forged samples are jointly input into the data augmentation module to obtain diversified augmented data. The augmented data is used as training data and input into the semi-supervised classifier for binary classification training. The present invention designs a classifier network more suitable for deepfake video forensics based on XceptionNet.
[0047] XceptionNet is a general image classifier, which is linearly stacked by depthwise separable convolution modules with residual connections. This network realizes the complete decoupling of spatial information in the length and width directions and cross-channel information. Therefore, its architecture is very easy to define and modify, and this feature is beneficial for us to design a flexible classifier architecture. Among them, the original XceptionNet adopts a three-layer flow structure design, including an entry layer (Entry Flow): mainly used for continuous downsampling to reduce the spatial dimension; eight middle layers (Middle flow): continuously learning correlation relationships and optimizing features; an exit layer (Exit flow): summarizing and organizing features and handing them over to the fully connected layer for expression. However, this network is not designed for deepfake detection tasks. And due to various post-processing operations during the synthesis of deepfake videos, post-processing processes such as blurring and compression will result in very few significant visual artifacts that can be easily extracted by the original XceptionNet. Therefore, in order to find more significant artifact features from the multi-scale space, while enhancing the model's ability to detect deepfake videos, without affecting its powerful feature extraction ability. Therefore, the present invention combines the normalized RGB image, the low-level spatial noise map, and the inter-frame correlated temporal noise map, retains the entry layer and the exit layer, and only uses one middle layer to perform convolution operations on the feature map, and can also mine forgery traces from different scales.
[0048] Before feeding into the entry layer, the processed image is normalized to obtain an RGB image, a noise map in the spatial domain, and a noise map in the temporal domain. For an image, the spatial domain noise map refers to a Gaussian filter with a kernel size of 5.
[0049] The generation of the temporal noise map goes through six steps: spatial Gaussian filtering; per-pixel temporal high-pass filtering; batch normalization; suppressing amplitudes less than the threshold t; calculating the temporal gradient; temporal low-pass filtering.
[0050] After convolution, pooling and other operations through the entry, middle, and exit layers, the logical regression value predicting the true or false category is finally output, and it is trained by backpropagation with cross-entropy loss. The definition of the cross-entropy loss function is as follows:
[0051]
[0052] y i represents the label of sample i, the forgery class is 1, and the real class is 0, p i represents the logical regression value of sample i predicted as the forgery class.
[0053] In the testing stage, the video to be detected is preprocessed and then sent into the trained classifier model. A Softmax layer is added to output the normalized prediction probabilities of the two categories of genuine and fake for each frame image. The detection result of the video is obtained by calculating the average value of the prediction probabilities of all frames of the video, that is, the final prediction confidence of the video is returned. If the confidence is higher than the set threshold p (p = 0.5), the video to be detected is judged as a genuine sample; if the confidence is less than the threshold p, it is classified as a forged sample.
[0054] The embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings, but the present invention is not limited to the described embodiments. For those skilled in the art, without departing from the principle and spirit of the present invention, various changes, modifications, substitutions and variations to these embodiments still fall within the protection scope of the present invention.
Claims
1. A deepfake video detection method based on semi-supervised learning, characterized in that: The authenticity discrimination of genuine and fake videos is achieved through three steps: negative sample generation, data augmentation, and semi-supervised classifier construction; The negative sample generation includes constructing a real video set, obtaining a real face set through preprocessing, randomly sampling a real face from the real face set, and extracting the key point information of the face; Calculate the distances between the key points of this face and the key points of the remaining faces in the set, find the face with the closest distance to this face, and use it as the foreground face; the sampled real face is used as the background face; Perform an affine transformation on the foreground face; paste the processed foreground face on the background face, and perform color correction, Gaussian blur, and edge smoothing to obtain a synthesized forged face image; Send the generated forged face and real face into the discriminator for judgment, and fine-tune the generated negative samples according to the correct or incorrect judgment of the discriminator. If the discriminator makes a correct judgment, it means that the generated negative samples are not realistic enough, and random sampling needs to be performed again to generate negative samples. Iterate until the discriminator cannot make a correct judgment. Its objective function is expressed by the following formula: where E(*) represents the expected value of the distribution function, P real (x) represents the distribution of real face samples, P pseudo-fake (z) represents the distribution of forged face samples after data augmentation, D(x) represents the discriminator's discrimination process, x is a real face sample, G(z) represents the negative sample generation process, and z is the face to be tampered with randomly sampled from the real face set; The data augmentation includes a data augmentation combination for deepfake video detection, including five augmentation forms: erasing, cropping, flipping, block recombination, and color jitter; The construction of the semi-supervised classifier includes a training stage and a testing stage. In the training stage, only real samples are used as the input of the negative sample generator. After the negative sample generation process, a large number of forged samples are generated. The balanced real samples and the generated forged samples are jointly input into the data augmentation module to obtain diversified augmented data. The augmented data is used as training data and input into the semi-supervised classifier for binary classification training.
2. The semi-supervised learning-based deepfake video detection method according to claim 1, characterized in that: The data augmentation combination includes: Erasing: A local area of the face image is randomly erased, that is, the pixel values of the randomly selected area are set to 0; Cropping: Randomly crop a part of the facial image, and then resize the cropped area to the original size; Flipping: Horizontally flip the original face image; Block recombination: Cut the original face into blocks of the same size, shuffle the order and recombine them; Color jitter: Random changes in the attributes of the face image, including brightness, contrast, saturation, and hue.
3. A deepfake video detection method based on semi-supervised learning according to claim 1, characterized in that: The semi-supervised classifier adopts a three-layer flow structure design, including an entrance layer, eight intermediate layers, and an exit layer, which summarizes and organizes features and is expressed by the fully connected layer. Before being sent into the entrance layer, the processed image is normalized RGB image, the noise map in the spatial domain, and the noise map in the temporal domain. For an image, the spatial domain noise map refers to a Gaussian filter with a kernel size of 5; The generation of the temporal domain noise map goes through six steps: spatial Gaussian filtering; Per-pixel temporal high-pass filtering; batch normalization; suppressing the amplitude less than the threshold t; Calculating the temporal gradient; temporal low-pass filtering; After the convolutional and pooling operations of the entrance, intermediate, and exit layers, the logical regression value predicting the authenticity category is finally output, and is trained by backpropagation using the cross-entropy loss. The definition of the cross-entropy loss function is as follows: y i represents the label of sample i, where the forged class is 1 and the genuine class is 0, p i represents the logistic regression value of sample i predicted as the forged class.
Citation Information
Patent Citations
False face video identification method and system and readable storage medium
CN111967427A
Deep counterfeit video detection method based on double fine-grained artifacts
CN115019370A