Counterfeit face detection method, device and system, and storage medium
By using YOLOv8 backbone network and blocked random masking technology in deep forgery detection, combined with the domain invariant feature learning module, the problem of insufficient overfitting and generalization performance in the existing technology is solved, and higher robustness and generalization are achieved.
Patent Information
- Application Number
- CN202411986416.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2044-12-31
AI Technical Summary
Existing deep forgery detection methods are prone to insufficient overfitting and generalization performance due to data deviations, making it difficult to effectively identify cross-domain forgery images.
The YOLOv8 backbone network is used to combine the blocked random mask and the domain invariant feature learning module to reduce the network's overfitting of irrelevant features through the blocked random mask, and learn global generalizable features and local key features through the domain invariant feature learning module.
It improves the robustness and generalization of the model, and can more accurately detect deep fake face pictures, showing good generalization effect across data sets.
Smart Images

Figure CN120032408A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and in particular relates to a forged face detection method and device, a system, and a storage medium. Background Art
[0002] Advances in deep learning and computer vision technologies have enabled artificial intelligence to better empower human life, but at the same time have also brought many threats. The detection of deep fake content has become one of the hot issues facing individuals, businesses, and governments around the world.
[0003] Existing deep fake detection methods are often based on deep neural networks. By carefully designing various feature learning modules and utilizing multiple learning strategies, they have demonstrated good performance in detection on some public datasets. However, this "illusion" of good performance often hides the risk of overfitting and the hidden danger of insufficient generalization performance. Due to the complex data bias caused by different data sources and different forgery methods, when faced with forged images from different domains, the model can easily fall into the "data trap" of early learning, making it difficult to identify cross-domain forged images. Summary of the invention
[0004] The technical problem to be solved by the present invention is to provide a method and device, system, and storage medium for detecting fake faces, which jointly mine more subtle artifact clues at the data level and the feature level, explore more generalized essential features, and thus effectively improve the robustness and generalization of the feature detection model.
[0005] To achieve the above object, the present invention adopts the following technical solution:
[0006] A forged face detection method, comprising:
[0007] Obtain a public dataset of fake faces;
[0008] Based on the public face forgery dataset, the enhanced face images are obtained by using block random mask operation;
[0009] The enhanced face image is input into the YOLOv8 backbone network, and the preliminary understanding of face image features is completed through the learning of the P1-P5 stages;
[0010] The domain-invariant feature learning module is used to learn global generalizable features and local key features respectively;
[0011] The obtained features are fused and input into the classification head of the target network;
[0012] The binary cross entropy loss function is used to update and optimize the network parameter weights to obtain the trained model weights;
[0013] Verify the trained model weights on multiple public datasets;
[0014] Use the trained model for reasoning to determine the authenticity of deep fake face images.
[0015] Preferably, the YOLOv8 backbone network is composed of CBS and C2f superimposed together, wherein CBS is composed of a combination of Conv2d, BatchNorm2d and SiLU, and C2f is CSPLayer_2Conv. In stages P1 to P5, the data passes through one CBS in stage P1 and four CBS-C2f combinations in stages P2 to P5 to complete the preliminary feature learning of the face image.
[0016] Preferably, the domain-invariant feature learning module is composed of a global generalized feature learning module and a local key feature learning module in parallel; wherein, the global generalized feature learning module enhances the model's learning of the high-frequency part by perturbing the low-frequency part of the face image, and the local key feature learning module emphasizes the imperceptible key category features in the forged samples through the encoding mapping of the individual inherent codebooks.
[0017] The present invention also provides a forged face detection device, comprising:
[0018] The first processing module is used to obtain a public dataset of forged faces;
[0019] The second processing module is used to obtain an enhanced face image by using a block random mask operation based on a public face forgery dataset;
[0020] The third processing module is used to input the enhanced face image into the backbone network of YOLOv8, and complete the preliminary understanding of face image features through learning in the P1-P5 stages;
[0021] a fourth processing module, for learning global generalizable features and local key features respectively through a domain invariant feature learning module;
[0022] The fifth processing module is used to fuse the obtained features and input them into the classification head of the target network;
[0023] The sixth processing module is used to update and optimize the network parameter weights using a binary cross entropy loss function to obtain a trained model weight;
[0024] The seventh processing module is used to verify the trained model weights on multiple public data sets;
[0025] The eighth processing module is used to use the trained model for reasoning to determine the authenticity of deep fake face images.
[0026] Preferably, the YOLOv8 backbone network is composed of CBS and C2f superimposed together, wherein CBS is composed of a combination of Conv2d, BatchNorm2d and SiLU, and C2f is CSPLayer_2Conv. In stages P1 to P5, the data passes through one CBS in stage P1 and four CBS-C2f combinations in stages P2 to P5 to complete the preliminary feature learning of the face image.
[0027] Preferably, the domain-invariant feature learning module is composed of a global generalized feature learning module and a local key feature learning module in parallel; wherein, the global generalized feature learning module enhances the model's learning of the high-frequency part by perturbing the low-frequency part of the face image, and the local key feature learning module emphasizes the imperceptible key category features in the forged samples through the encoding mapping of the individual inherent codebooks.
[0028] An embodiment of the present invention further provides a forged face detection system, comprising: a memory and a processor, wherein the memory stores a computer program executed by the processor, and the computer program executes a forged face detection method when executed by the processor.
[0029] An embodiment of the present invention further provides a storage medium, on which a computer program is stored, and the computer program executes the forged face detection method when running.
[0030] Compared with the prior art, the present invention has the following advantages and technical effects:
[0031] 1. Use the object detection network YOLOv8 as the backbone network for face forgery detection. Compared with the conventional classification backbone network, the YOLOv8 backbone network has higher basic detection performance.
[0032] 2. The block random masking strategy is adopted to effectively alleviate the overfitting problem of the network for irrelevant features and specific patterns at the data level.
[0033] 3. The domain-invariant feature learning module (DICL) is used to better learn discriminative domain-invariant features by sampling and reconstructing the low-frequency distribution in the image and combining it with local key features. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0035] Figure 1This is a flow chart of a forged face detection method according to an embodiment of the present invention;
[0036] Figure 2 A schematic diagram of data processing of a forged face detection method according to an embodiment of the present invention;
[0037] Figure 3 This is a sample diagram of block random mask processing according to an embodiment of the present invention;
[0038] Figure 4 A heat map is shown for the experimental results of the embodiments of the present invention. DETAILED DESCRIPTION
[0039] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0040] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0041] Embodiment 1:
[0042] like Figure 1 As shown, this embodiment provides a method for counting pedestrian flow by integrating semantic information, including the following steps:
[0043] S1: Obtain a public dataset of forged faces and perform preprocessing such as frame cutting and face cropping.
[0044] S11: Download the original video data of FF++, Celeb-DF, DFDC, DFDC-P and DFD respectively, and randomly extract frames from the video data to obtain frame image datasets. The FF++ dataset is used for training and verification, and the Celeb-DF, DFDC, WDF and DFD datasets are only used for verification.
[0045] FF++ contains a total of 1,000 videos generated by four forgery methods: DeepFakes, Face2Face, FaceSwap, and NeuralTextures. For each forgery method, the dataset provides videos in three resolutions: Raw, high quality (HQ), and low quality (LQ).
[0046] Celeb-DF contains 590 original videos and 5639 corresponding deep fake videos, which are derived from 59 celebrity interview videos of different genders, ages and races on YouTube. The average length of all videos is about 13 seconds, the standard frame rate is 30 frames per second, and the final video format is MPEG4.0.
[0047] DFDC is a more difficult face forgery detection dataset derived from the competition, consisting of 124,000 video clips, 244 real clips shot by various actors in different scenarios, and deeply forged using 8 different techniques.
[0048] The WildDeepfake (WDF) dataset consists of 3,805 real video clips and 3,509 fake video clips collected from the Internet. The videos in this dataset are uniquely challenging because they are manipulated using undisclosed techniques and presented in different contexts. The DFD dataset contains more than 363 original videos of 28 paid actors in 16 different scenes, and more than 3,000 fake videos using DeepFakes.
[0049] S12: For the frame image data, use the RtinaFace tool to locate the face area in the image, obtain the face position information, and then crop the face image data.
[0050] S13: resize and normalize the face image to obtain a preprocessed face image. The processed face image is resized to 224×224 pixels.
[0051] S2: Divide the FF++ dataset into a training set, a validation set, and a test set, and input the face images in the training set into the target network. Specifically, the FF++ dataset is divided into a ratio of 36:7:7.
[0052] S3: Use block random masking to reduce the network's overfitting of irrelevant information in the image and obtain an enhanced face image. Figure 3 Shows examples of face images enhanced with random mask blocks.
[0053] S31: Block mask generation: The mask block is set to a square with an aspect ratio of 1. In order to take into account the "disturbance" of the background and face areas, five random and discrete mask operations are performed on each image. The erasure ratio of each mask block is 2% of the entire image. That is, the overall perturbation ratio will be controlled below 10% of the entire image, avoiding excessive erasure that inadvertently covers important local artifacts, thereby misleading the model;
[0054] S32: Adding random noise: Compared with the "all black" operation of setting all pixels to 0, the "all white" operation of setting all pixels to 255, and the operation of filling with the ImageNet average value, adding random pixel values to the mask can better reduce the error rate of the network during testing. Therefore, further adding random noise to the random mask block generated in the previous step can achieve a more robust data enhancement effect.
[0055] S4: Input the backbone network of yolov8, and complete the preliminary understanding of facial image features after learning in the P1-P5 stages;
[0056] S5: Learn global generalizable features and local key features respectively through domain invariant feature learning module;
[0057] S51: Since the high-frequency part of the image contains more global structural features, and the low-frequency spectrum of the image contains most of the style information of domain changes, the global generalized feature learning module enhances the model's learning of the high-frequency part by fully perturbing the low-frequency part of the face image, and finally obtains the output feature X GCFL ;
[0058] S52: In order to obtain global generalization features while retaining the easily ignored local region features, a local key feature learning (LCFL) module is added to the domain invariant feature learning module in parallel with GGFL. The purpose is to emphasize the key category features that are not easily perceived in the forged samples through the encoding mapping of an inherent codebook, and finally obtain the output feature X DIFL .
[0059] S6: The obtained features are fused through the Concat operation and input into the classification head of the target network;
[0060] S7: Use the binary cross entropy loss function to update and optimize the network parameter weights to obtain the trained model weights;
[0061] S8: Verify the trained model weights on multiple public datasets.
[0062] S9: Use the trained model for inference to determine the authenticity of deep fake face images.
[0063] As a further limitation of the technical solution, the specific steps of S4 are:
[0064] The YOLOv8 backbone network used is mainly composed of CBS and C2f. CBS is a combination of Conv2d+BatchNorm2d+SiLU; C2f, namely CSPLayer_2Conv, allows YOLOv8 to obtain richer gradient flow information while ensuring lightweight. In the P1 to P5 stages, the data passes through a CBS (P1) and four CBS-C2f combinations (P2 to P4) in turn to complete the preliminary feature learning of the face image, and the learned features are input into step S5.
[0065] Furthermore, the specific steps of S51 are:
[0066] First, for x∈R H×W×D , and perform DCT along the spatial dimension to convert x into the frequency domain:
[0067] X=D[x]∈C H×W×D (1)
[0068] Where D[·] represents the DCT transform. Note that X is a complex tensor representing the frequency domain space of x;
[0069] Secondly, for a given batch of samples Extract the low-frequency space in the image separately and high frequency space According to the characteristics of DCT, the extracted low-frequency part is located in the upper left corner of the image;
[0070] Then, the obtained low-frequency part is modeled by Gaussian distribution. The mean is the original value of the corresponding pixel, and the variance is calculated by the value of the element in different samples. This process can be expressed by mathematical formula:
[0071]
[0072] ∑ 2 The size of reflects the change in the potential domain shift of the corresponding element. For each element in the low frequency spectrum, resample from the estimated Gaussian distribution to obtain its new pixel value:
[0073]
[0074] Where ε∈[0,1] represents the intensity of the perturbation. Through this method, the model can not only retain some detailed features carried by the low-frequency part of the original image, but also introduce other noises in a similar distribution, thereby enhancing the network's learning of the semantic information carried by the high-frequency part.
[0075] Finally, the perturbed low-frequency part and the original high frequency part Combined into a complete new frequency domain, and further combined with a global filter Q∈C H×W×C Perform a dot product operation:
[0076]
[0077] The global filter Q can be viewed as a set of learnable frequency filters of different hidden dimensions. This method can further help the network remove features that are irrelevant to global structural information, thereby enhancing the generalization of global features. It is then mapped back to the spatial domain through inverse DCT transform, normalized, and passed into a set of multi-layer perceptrons.
[0078]
[0079] The low-frequency perturbed features are subjected to residual operations with the original features to complete global generalizable feature learning.
[0080]
[0081] Furthermore, the specific steps of S52 are:
[0082] First, a set of convolutions are used to transform the input feature X in The encoded features are further processed using CBR blocks. CBR blocks consist of 3×3 convolutions with BN layers and ReLU activation functions. Through the above steps, the processed encoded features are input into the Codebook and combined with a set of learnable scaling factors. Specifically, a set of scaling factors c is used in sequence. and b m Map the corresponding position information. The information about the mth codeword in the entire image can be calculated as follows:
[0083]
[0084] in, is the i-th pixel, b m is the mth learnable visual codeword, c m is the mth scaling factor and is also a learnable parameter set. is the information about the position of each pixel relative to the codeword. M is the total number of visual centers. In order to fuse all s m , use f to calculate the complete image information with M codewords:
[0085]
[0086] Among them, f is a BN layer containing a ReLu layer and an average layer. Subsequently, the complete features after codebook mapping are sent to a fully connected layer and a 1×1 convolutional layer to predict the prominent key class features, and the input feature X in Channel multiplication of local angular area features processed with the scale factor coefficient:
[0087]
[0088] is the channel multiplication, w(·) is the Sigmoid activation function, Conv 1×1 (·) is a 1×1 convolution operation. Finally, the input feature X in With local features X l Add channel by channel to get the final local key feature X LCFL :
[0089]
[0090] It is channel level addition.
[0091] Furthermore, the specific steps of S6 are:
[0092] The features learned through global generalized feature learning (GGFL) and local key feature learning (LCFL) are concat-concatenated to obtain the final domain-invariant features:
[0093] X DIFL =Concat(X GGFL ,X LCFL ) (11)
[0094] Furthermore, the specific steps of S7 are:
[0095] The loss function of the model is a binary cross entropy loss function. The calculation formula of the binary cross entropy loss function is:
[0096]
[0097] Where y is the true label, which can be 0 or 1 for a binary classification problem. is the probability predicted by the model, that is, the probability that the model predicts the positive class.
[0098] Furthermore, the specific steps of S8 are:
[0099] The model trained on the FF++ dataset is tested for generalization on the FF++, DFDC, Celeb-DF, WDF, and DFD datasets, and compared with other existing face forgery detection methods. The evaluation indicators are ACC (Accuracy) and AUC (Area Under the Receiver Operating Characteristic Curve). Figure 4 The heat maps generated by the GradCam technology on different forged images are compared between the embodiments of the present invention and other existing networks.
[0100] The present invention obtains a public dataset of face forgery and performs preprocessing work such as frame cutting and face cutting; divides the dataset, and inputs the face images in the training set into the target network; uses block random mask operation to obtain enhanced face images; inputs the backbone network of YOLOv8, and completes the preliminary understanding of face image features through learning in the P1-P5 stage; respectively learns global generalizable features and local key features through the domain invariant feature learning module; fuses the obtained features and inputs them into the classification head of the target network; uses the binary cross entropy loss function to update and optimize the network parameter weights to obtain the trained model weights; verifies the trained model weights on multiple public datasets; uses the trained model for reasoning to achieve the true and false judgment of deep forged face images. The present invention innovatively uses the target detection network as the backbone of the face forgery detection model, alleviates the network's overfitting learning of features and specific patterns unrelated to forgery at the data level, and at the same time, reconstructs the low-frequency distribution in the image by sampling and combining it with local key features to better learn the discriminative domain invariant features, so as to achieve more accurate and more generalized face forgery detection performance.
[0101] Experiments have shown that the ACC of this embodiment on the C23 and C40 subsets of FF++ achieved 98.76% and 93.71% detection values, and the AUC achieved 99.53% and 96.85% detection values, respectively, surpassing most existing fake face detection solutions. Through cross-dataset generalization experiments, this example was trained on the FF++ dataset and tested on the Celeb-DF, DFDC, WDF, and DFD datasets, achieving AUC values of 73.03%, 71.18%, 68.42%, and 78.26%, respectively, which can achieve good generalization effects. This shows that the solution adopted in this article can effectively improve the accuracy and robustness of the model.
[0102] Embodiment 2:
[0103] The embodiment of the present invention also provides a forged face detection device, comprising:
[0104] The first processing module is used to obtain a public dataset of forged faces;
[0105] The second processing module is used to obtain an enhanced face image by using a block random mask operation based on a public face forgery dataset;
[0106] The third processing module is used to input the enhanced face image into the backbone network of YOLOv8, and complete the preliminary understanding of face image features through learning in the P1-P5 stages;
[0107] a fourth processing module, for learning global generalizable features and local key features respectively through a domain invariant feature learning module;
[0108] The fifth processing module is used to fuse the obtained features and input them into the classification head of the target network;
[0109] The sixth processing module is used to update and optimize the network parameter weights using a binary cross entropy loss function to obtain a trained model weight;
[0110] The seventh processing module is used to verify the trained model weights on multiple public data sets;
[0111] The eighth processing module is used to use the trained model for reasoning to determine the authenticity of deep fake face images.
[0112] As an implementation method of an embodiment of the present invention, the YOLOv8 backbone network is composed of CBS and C2f superimposed together, wherein CBS is composed of a combination of Conv2d, BatchNorm2d and SiLU, C2f is CSPLayer_2Conv, and in stages P1 to P5, the data passes through one CBS in stage P1 and four CBS-C2f combinations in stages P2 to P5 in turn to complete the preliminary feature learning of the face image.
[0113] As an implementation method of an embodiment of the present invention, the domain-invariant feature learning module is composed of a global generalized feature learning module and a local key feature learning module in parallel; wherein, the global generalized feature learning module enhances the model's learning of the high-frequency part by perturbing the low-frequency part of the face image, and the local key feature learning module emphasizes the imperceptible key category features in the forged sample through the encoding mapping of an inherent code book.
[0114] Embodiment 3:
[0115] An embodiment of the present invention further provides a forged face detection system, comprising: a memory and a processor, wherein the memory stores a computer program executed by the processor, and the computer program executes a forged face detection method when executed by the processor.
[0116] Embodiment 4:
[0117] An embodiment of the present invention further provides a storage medium, on which a computer program is stored, and the computer program executes the forged face detection method when running.
[0118] The embodiments described above are only descriptions of the preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the design spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by ordinary technicians in this field should all fall within the protection scope determined by the claims of the present invention.
Claims
1. A method for detecting fake faces, characterized in that: include: Obtain a public dataset of fake faces; Based on the public face forgery dataset, the enhanced face images are obtained by using block random mask operation; The enhanced face image is input into the YOLOv8 backbone network, and the preliminary understanding of face image features is completed through the learning of the P1-P5 stages; The domain-invariant feature learning module is used to learn global generalizable features and local key features respectively; The obtained features are fused and input into the classification head of the target network; The binary cross entropy loss function is used to update and optimize the network parameter weights to obtain the trained model weights; Verify the trained model weights on multiple public datasets; Use the trained model for reasoning to determine the authenticity of deep fake face images.
2. The forged face detection method according to claim 1, characterized in that: The YOLOv8 backbone network is composed of CBS and C2f superimposed together. CBS is composed of a combination of Conv2d, BatchNorm2d and SiLU, and C2f is CSPLayer_2Conv. In stages P1 to P5, the data passes through one CBS in stage P1 and four CBS-C2f combinations in stages P2 to P5 to complete the preliminary feature learning of face images.
3. The forged face detection method according to claim 2, characterized in that: The domain-invariant feature learning module is composed of a global generalized feature learning module and a local key feature learning module in parallel; the global generalized feature learning module enhances the model's learning of the high-frequency part by perturbing the low-frequency part of the face image, and the local key feature learning module emphasizes the subtle key category features in the forged samples through the encoding mapping of the individual inherent codebooks.
4. A forged face detection device, characterized in that: include: The first processing module is used to obtain a public dataset of forged faces; The second processing module is used to obtain an enhanced face image by using a block random mask operation based on a public face forgery dataset; The third processing module is used to input the enhanced face image into the backbone network of YOLOv8, and complete the preliminary understanding of face image features through learning in the P1-P5 stages; a fourth processing module, for learning global generalizable features and local key features respectively through a domain invariant feature learning module; The fifth processing module is used to fuse the obtained features and input them into the classification head of the target network; The sixth processing module is used to update and optimize the network parameter weights using a binary cross entropy loss function to obtain a trained model weight; The seventh processing module is used to verify the trained model weights on multiple public data sets; The eighth processing module is used to use the trained model for reasoning to determine the authenticity of deep fake face images.
5. The forged face detection device according to claim 4, characterized in that: The YOLOv8 backbone network is composed of CBS and C2f superimposed together. CBS is composed of a combination of Conv2d, BatchNorm2d and SiLU, and C2f is CSPLayer_2Conv. In stages P1 to P5, the data passes through one CBS in stage P1 and four CBS-C2f combinations in stages P2 to P5 to complete the preliminary feature learning of face images.
6. The forged face detection device according to claim 5, characterized in that: The domain-invariant feature learning module is composed of a global generalized feature learning module and a local key feature learning module in parallel; the global generalized feature learning module enhances the model's learning of the high-frequency part by perturbing the low-frequency part of the face image, and the local key feature learning module emphasizes the subtle key category features in the forged samples through the encoding mapping of the individual inherent codebooks.
7. A fake face detection system, characterized in that: include: A memory and a processor, wherein the memory stores a computer program executed by the processor, and when the computer program is executed by the processor, the forged face detection method according to any one of claims 1 to 3 is executed.
8. A storage medium, characterized in that: The storage medium stores a computer program, which executes the forged face detection method according to any one of claims 1 to 3 when running.
Citation Information
Patent Citations
Generalized face forgery detection method based on domain invariant features
CN114692741A
Living body detection method based on self-supervised domain clustering and domain generalization
CN116403290A
Face forgery detection method based on global context structure difference
CN117037290A
Cross-modal cross-domain universal face forgery positioning method
CN117292442A
Social network deep counterfeit video detection method and system based on spatial-temporal characteristics
CN117523439A