A Face Forgery Detection Method Based on Reconstruction Error Distribution
By building a face forgery detection network based on the dual-stream reconstruction strategy, combining local and global reconstruction error distribution, the problem of insufficient robustness and generalization capabilities of face forgery detection in the existing technology is solved, and more efficient face forgery detection is achieved.
Patent Information
- Application Number
- CN202211648476.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-21
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-12-21
AI Technical Summary
The prior art cannot achieve robust extraction of real and forged facial features in face forgery detection, which limits the detection performance and generalization capabilities.
A face forged detection network model including local transformer reconstruction module, global convolution reconstruction module and forged detection module is constructed based on the dual-stream reconstruction strategy. The data is enhanced by using the random mask strategy, and feature fusion is performed through local and global reconstruction error distributions, and finally prediction is performed using the forged detection module.
It improves the generalization ability of novel unknown forgery types and low-quality damaged scenes of face data, enhances the resolution and robustness of features, and realizes more robust face forgery detection.
Smart Images

Figure CN116206347B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of face forgery detection, and particularly relates to a face forgery detection method based on the distribution of reconstruction errors. Background Art
[0002] Face recognition technology is a biometric identification technology based on the facial feature information of a person, and has currently been widely applied in fields such as finance, justice, military, public security, border control, government, aerospace, electric power, factories, education, medical care, and numerous enterprises and institutions. With the rapid development of deepfake technology and face forgery models, a large number of realistic face models have emerged, which are difficult to effectively identify using traditional identification methods. At the same time, face forgery presents a development trend of concealment and diversity, bringing severe challenges to generalized face forgery detection.
[0003] Considering that early face forgery technology has its uncontrollable characteristics, forged images often exhibit abnormal manifestations in the spatial domain and frequency domain. Specifically, many forgery detection methods analyze the obvious visual artifacts that appear in the facial region or the inconsistencies in the frequency or timeline.
[0004] Recently, noticing the inconsistency in the distribution of genuine and fake face data, the distribution of reconstruction errors has begun to be applied to face forgery detection. This method does not rely on specific forgery clues and has better generalization performance. For example, Document 1, "J. Cao, C. Ma, T. Yao, S. Chen, S. Ding, and X. Yang. End-to-end reconstruction-classification learning for face forgery detection. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pages 4113–4122, 2022" proposed a forgery detection framework based on reconstruction-classification learning. The reconstruction learning on real images can enhance the learned representations and even discover unknown forgery patterns; while the classification learning is responsible for mining the essential differences between real and fake images and promoting the understanding of forgery. Document 2, "Y. He, N. Yu, M. Keuper, and M. Fritz. Beyond the spectrum: Detecting deepfakes via re-synthesis. arXiv preprint arXiv:2105.14376, 2021" proposed a re-synthesis strategy for super-resolution, denoising, and coloring to extract robust features and isolate fake images, and extract residual visual clues for forgery detection.
[0005] Although the above reconstruction error method can amplify the differences in the distributions between forged and real faces, it cannot robustly extract the features of real and forged faces, which makes the learned face features incomplete and limits the performance and robustness of face forgery detection. Summary of the Invention
[0006] To solve the above problems existing in the prior art, the present invention provides a face forgery detection method based on the reconstruction error distribution. The technical problems to be solved by the present invention are achieved through the following technical solutions:
[0007] A face forgery detection method based on the reconstruction error distribution, comprising:
[0008] Constructing a face forgery detection network model including a local transformer reconstruction module, a global convolutional reconstruction module, and a forgery detection module based on a two-stream reconstruction strategy;
[0009] Using a random masking strategy to perform data augmentation on the image to be detected to obtain a pixel sequence;
[0010] Based on the pixel sequence, using the local transformer reconstruction module to perform image reconstruction to obtain a local reconstructed image; at the same time, using the global convolutional reconstruction module to reconstruct the image to be detected to obtain a global reconstructed image;
[0011] Based on the reconstruction error distributions corresponding to the local reconstructed image and the global reconstructed image, merging the local reconstructed image and the global reconstructed image to obtain a fused image;
[0012] Using the forgery detection module to predict the fused image to obtain a face forgery detection result.
[0013] In an embodiment of the present invention, using a random masking strategy to perform data augmentation on the image to be detected to obtain an enhanced pixel sequence, comprising:
[0014] Dividing the image to be detected into a number of non-overlapping pixel blocks;
[0015] Randomly generating a 01 mask corresponding to and uniformly distributed over the pixel blocks;
[0016] Selecting the pixel blocks corresponding to the mask value of 1 to generate a pixel sequence.
[0017] In an embodiment of the present invention, the local transformer reconstruction module includes a linear projection unit, a first encoder, and a first decoder; then, based on the pixel sequence, using the local transformer reconstruction module to perform image reconstruction to obtain a local reconstructed image, comprising:
[0018] Perform linear projection on the pixel sequence using the linear projection unit;
[0019] Extract features from the output image after linear projection using the first encoder to obtain a first feature map;
[0020] Reconstruct the first feature map using the first decoder to obtain a local reconstructed image.
[0021] In an embodiment of the present invention, the global convolutional reconstruction module includes a second encoder, a DMFB unit, and a second decoder; then, reconstruct the image to be detected using the global convolutional reconstruction module to obtain a global reconstructed image, including:
[0022] Extract features from the image to be detected using the second encoder to obtain a second feature map;
[0023] Perform multi-scale feature extraction on the second feature map using the DMFB unit, and fuse the extracted multi-scale features to obtain a multi-scale feature map;
[0024] Reconstruct the multi-scale feature map using the second decoder to obtain a global reconstructed image.
[0025] In an embodiment of the present invention, based on the reconstruction error distributions corresponding to the local reconstructed image and the global reconstructed image, merge the local reconstructed image and the global reconstructed image to obtain a fused image, including:
[0026] Perform error analysis on the local reconstructed image and the global reconstructed image respectively based on the image to be detected, and correspondingly obtain a local reconstruction error distribution and a global reconstruction error distribution;
[0027] Add the local reconstruction error distribution to the local reconstructed image, and add the global reconstruction error distribution to the global reconstructed image;
[0028] Fuse the local reconstructed image and the global reconstructed image with the added reconstruction errors to obtain a fused image.
[0029] In an embodiment of the present invention, the forgery detection module includes a third encoder and a linear layer; then, use the forgery detection module to predict the fused image to obtain a face forgery detection result, including:
[0030] Extract features from the fused image using the third encoder;
[0031] Based on the extracted features, use the linear layer to perform prediction to obtain a face forgery detection result.
[0032] In one embodiment of the present invention, after constructing the face forgery detection network model, the following steps are further included:
[0033] Construct a loss function to train the network model; where the expression of the loss function is:
[0034]
[0035]
[0036] where W represents the width of the image to be detected, H represents the height of the image to be detected, || ||2 represents the L2 norm, I represents the image to be detected, Output a and Output a represent the local reconstructed image and the global reconstructed image respectively, y represents the true label, represents the predicted label.
[0037] Advantages of the present invention:
[0038] The face forgery detection method based on the reconstruction error distribution provided by the present invention, on the one hand, applies the random mask strategy to the reconstruction process of genuine and forged images, increasing the robustness and generality of the reconstruction error distribution of the reconstructed images; on the other hand, uses a two-stream reconstruction network based on the error distribution to analyze the local and global reconstruction errors, enhancing the generalization ability for novel unknown forgery types and low-quality damaged scenarios of face data; this method simultaneously considers the extraction of robust genuine and forged features and the characterization features of local and global robustness, making the finally obtained features more discriminative and robust, thus realizing more robust face forgery detection.
[0039] The following will further elaborate on the present invention in conjunction with the accompanying drawings and embodiments. Brief Description of the Drawings
[0040] Figure 1 is a schematic diagram of a face forgery detection method based on the reconstruction error distribution provided by an embodiment of the present invention;
[0041] Figure 2 is a block diagram of a face forgery detection network model provided by an embodiment of the present invention;
[0042] Figure 3 is an overall block diagram of a face forgery detection method based on the reconstruction error distribution provided by an embodiment of the present invention. Detailed Embodiments
[0043] The following further describes the present invention in detail with specific embodiments, but the implementation manners of the present invention are not limited thereto.
[0044] Embodiment 1
[0045] Please refer to Figure 1 , Figure 1 which is a schematic diagram of a face forgery detection method based on reconstruction error distribution provided by an embodiment of the present invention, and includes:
[0046] Step 1: Construct a face forgery detection network model including a local transformer reconstruction module, a global convolutional reconstruction module, and a forgery detection module based on a dual-stream reconstruction strategy, as Figure 2 shown.
[0047] In this embodiment, a reconstruction network based on error distribution dual-stream is used, which can analyze the reconstruction errors of local and global reconstructions simultaneously, increasing the generalization ability of the model and solving the problem of insufficient generalization ability of the model for novel unknown forgery types and low-quality damaged face data scenarios.
[0048] Step 2: Use a random masking strategy to perform data augmentation on the image to be detected to obtain a pixel sequence.
[0049] Please refer to Figure 2 , Figure 2 which is the overall block diagram of a face forgery detection method based on reconstruction error distribution provided by an embodiment of the present invention.
[0050] Specifically, for the face image to be detected, a random masking strategy needs to be adopted to continue data augmentation, including:
[0051] Divide the image to be detected into several non-overlapping pixel blocks;
[0052] Randomly generate 0-1 masks corresponding to the pixel blocks and uniformly distributed;
[0053] Select the corresponding pixel blocks with a mask value of 1 to generate a pixel sequence.
[0054] For example, first, this embodiment can divide the image to be detected I with a size of C×H×W into non-overlapping pixel blocks with a size of d×d Then, randomly generate a uniformly distributed 0, 1 mask with a length of . Finally, the corresponding pixel blocks with a mask value of 1 can be selected to generate a pixel sequence.
[0055] It can be understood that the corresponding pixel blocks with a mask value of 0 can also be selected to generate a pixel sequence according to needs.
[0056] Step 3: Based on the pixel sequence, use the local transformer reconstruction module to perform image reconstruction to obtain a local reconstructed image; at the same time, use the global convolutional reconstruction module to reconstruct the image to be detected to obtain a global reconstructed image.
[0057] In this embodiment, the local transformer reconstruction module includes a linear projection unit, a first encoder, and a first decoder. As Figure 3 shown, using the local transformer reconstruction module for image reconstruction, the obtained local reconstructed image includes:
[0058] a) Performing a linear projection on the pixel sequence using the linear projection unit, and its expression is:
[0059] Input a ={p k |mask[k]=1}
[0060] where Input a represents the output image after linear projection, that is, the input image of the first encoder, and p k represents the pixel sequence in the image to be detected.
[0061] b) Extracting features from the output image after linear projection using the first encoder to obtain the first feature map, and its expression is:
[0062] latent a =Encoder a (f concat (p l ,Input a )+pos)
[0063] where p l is a learnable feature vector, pos is the position encoding, f concat represents an operation of combining two vectors with the same shape, and latent a represents the robust face feature image output by the first encoder, that is, the first feature map.
[0064] c) Reconstructing the first feature map using the first decoder to obtain the local reconstructed image.
[0065] Specifically, inputting the robust face feature and the learnable feature vector into the first decoder, and predicting the d×d pixel values of each pixel block through the last linear layer. The calculation formula is:
[0066] Output a =Decoder a (f concat (tokens,latent a )+pos)
[0067] where tokens is the learnable feature vector, and Output a represents the output image of the first decoder, that is, the local reconstructed image.
[0068] Further, please continue to refer to Figure 3 , where the global convolutional reconstruction module includes a second encoder, a DMFB unit, and a second decoder. The global convolutional reconstruction module is used to reconstruct the image to be detected to obtain a globally reconstructed image, including:
[0069] a) Use the second encoder to extract features from the image to be detected to obtain a second feature map, and its expression is:
[0070] latent a1 = Encoder a (I)
[0071] where I represents the image to be detected, represents the output image of the second encoder, that is, the second feature map.
[0072] b) Use the DMFB unit to perform multi-scale feature extraction on the second feature map, and fuse the extracted multi-scale features to obtain a multi-scale feature map, and its expression is:
[0073]
[0074] where, represents the output image of the DMFB unit, that is, the multi-scale feature map.
[0075] c) Use the second decoder to reconstruct the multi-scale feature map to obtain a globally reconstructed image, and its expression is:
[0076]
[0077] where Output b represents the output image of the second decoder, that is, the globally reconstructed image.
[0078] Step 4: Based on the reconstruction error distributions corresponding to the local reconstructed image and the globally reconstructed image, merge the local reconstructed image and the globally reconstructed image to obtain a fused image.
[0079] First, perform error analysis on the local reconstructed image and the globally reconstructed image respectively based on the image to be detected, and obtain the local reconstruction error distribution and the global reconstruction error distribution correspondingly. Their expressions are respectively:
[0080] Error a = f abs (I - Output a )
[0081] Error b = f abs (I - Output b )
[0082] Among them, f abs is the absolute value operation on the vector, Error a and Error b respectively represent the local reconstruction error distribution and the global reconstruction error distribution.
[0083] Then, the local reconstruction error distribution is added to the local reconstructed image, and the global reconstruction error distribution is added to the global reconstructed image. The formula is expressed as:
[0084]
[0085]
[0086] Finally, the local reconstructed image and the global reconstructed image with the added reconstruction error are fused to obtain a fused image, and its expression is:
[0087]
[0088] Step 5: Use the forgery detection module to predict the fused image to obtain the face forgery detection result.
[0089] First, use the third encoder to extract features from the fused image; then, based on the extracted features, use the linear layer for prediction to obtain the face forgery detection result.
[0090] Specifically, similar to the local transformer reconstruction module, the input is divided into non-overlapping pixel blocks of size 2C×d×d. Then, the feature vectors of all pixel blocks are obtained through the third encoder, and finally, the linear layer is connected to predict whether the image is forged.
[0091] It should be noted that in this embodiment, during testing, a post-processing module of test-time augmentation (TTA) is added, and its calculation formula is:
[0092]
[0093] latent c =Encoder c (Input c )
[0094] It can be understood that after constructing the face forgery detection network model, it also includes: constructing a loss function to train the network model.
[0095] Specifically, considering the loss of dual-stream self-supervised face reconstruction and the loss of binary classification based on a supervised forgery discriminator, this embodiment uses the L2 norm to ensure the consistency between the reconstructed image and the original image. At the same time, the cross-entropy loss is used to classify real and fake images. The above three loss functions are weighted to obtain the overall loss function, and its expression is:
[0096]
[0097]
[0098] where W represents the width of the image to be detected, H represents the height of the image to be detected, || ||2 represents the L2 norm, I represents the image to be detected, Output a and Output a represent the local reconstructed image and the global reconstructed image respectively, y represents the true label, represents the predicted label.
[0099] This embodiment adopts a training method similar to the generative adversarial network. First, only the face reconstruction module is trained, and then the forgery discriminator is alternately trained to ensure that the reconstructed face does not produce a large deformation. The detailed process can refer to the prior art and will not be introduced here.
[0100] On the one hand, the face forgery detection method based on the reconstruction error distribution provided by the present invention applies the random mask strategy to the reconstruction process of real and fake images, increasing the robustness and generality of the reconstruction error distribution of the reconstructed images. On the other hand, the dual-stream reconstruction network based on the error distribution is used to analyze the local and global reconstruction errors, enhancing the generalization ability for novel unknown forgery types and low-quality damaged scenarios of face data. This method simultaneously considers the extraction of robust real and fake features and the characterization features of local and global robustness, making the finally obtained features more discriminative and robust, thus realizing more robust face forgery detection.
[0101] Embodiment 2
[0102] In this embodiment, simulation experiments are carried out on the commonly used face forgery detection datasets WildDeepfake and Celeb-DF (v2), and compared with several existing methods to verify the effectiveness and robustness of the above method.
[0103] Among them, WildDeepfake contains 7,314 facial sequences extracted from 707 Deepfake videos. The videos are all collected from bilibili and YouTube, and their face-swapping videos are synthesized by various methods, so the detection difficulty is greater.
[0104] Celeb-DF (v2) is a large-scale dataset proposed based on Celeb-DF v1, containing 590 original real videos collected from YouTube and 5,639 corresponding fake videos, with diverse distributions in terms of gender, age, and race.
[0105] Experiment 1:
[0106] In this experiment, the effects of 8 existing methods and the method of the present invention were compared on the WildDeepfake dataset. Among them, the 8 methods are as follows:
[0107] Method 1 is the method provided in the literature "D. Afchar, V. Nozick, J. Yamagishi, and I. Echizen. Mesonet: a compact facial video forgery detection network. In 2018 IEEE international workshop on information forensics and security (WIFS), pages 1–7. IEEE, 2018"; Method 2 is the method provided in the literature "A. Rossler, D. Cozzolino, L. Verdoliva, C. Riess, J. Thies, and M. Nieβner. Faceforensics++: Learning to detect manipulated facial images. In Proceedings of the IEEE / CVF international conference on computer vision, pages 1–11, 2019"; Method 3 is the method provided in the literature "H. H. Nguyen, J. Yamagishi, and I. Echizen. Capsule forensics: Using capsule networks to detect forged images and videos. In ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 2307–2311. IEEE, 2019"; Method 4 is the method provided in the literature "B. Zi, M. Chang, J. Chen, X. Ma, and Y.-G. Jiang. Wilddeepfake: A challenging real-world dataset for deepfake detection. In Proceedings of the 28th ACM international conference on multimedia, pages 2382–2390, 2020"; Method 5 is the method provided in the literature "J. Hu, X. Liao, W. Wang, and Z. Qin.The method provided in "Detecting compressed deepfake videos in social networks using frame-temporality two-stream convolutional network. IEEE Transactions on Circuits and Systems for Video Technology, 32(3): 1089–1102, 2021"; Method 6 is the method provided in the literature "J. Hu, X. Liao, J. Liang, W. Zhou, and Z. Qin. Infer: Frame inference-based deepfake detection for high-visual-quality videos. IEEE Trans. Pattern Anal. Mach. Intell, pages 1–9, 2022"; Method 7 is the method provided in the literature "Y. Zhou, A. Luo, X. Kang, and S. Lyu. Face forgery detection based on segmentation network. In 2021 IEEE International Conference on Image Processing (ICIP), pages 3597–3601. IEEE, 2021"; Method 8 is the method provided in the literature "Y. Qian, G. Yin, L. Sheng, Z. Chen, and J. Shao. Thinking in frequency: Face forgery detection by mining frequency-aware clues. In European conference on computer vision, pages 86–103. Springer, 2020".
[0108] The test results are shown in Table 1 below.
[0109] Table 1 Performance comparison of different methods on the WildDeepfake dataset
[0110] Method Accuracy(%) Method 1: MesoNet 73.95 Method 2: Xception 74.32 Method 3: Capsule 77.60 Method 4: ADDNet 76.25 Method 5: FT-two-stream 68.78 Method 6: FInfer 75.88 Method 7: Segmentation Network 79.87 <![CDATA[Method 8: F 2 -Net]]> 80.66 Method of the present invention: RED 79.44
[0111] As can be seen from Table 1, the method of the present invention achieves a relatively high accuracy of 79.44%. Compared with other frame-level methods, the accuracy of the present invention is 1.84% higher than that of Method 3 Capsule, 5.12% higher than that of Method 2 Xception, and 6.49% higher than that of Method 1 MesoNet. Compared with the existing methods, the simple baseline of the method of the present invention is also competitive.
[0112] Experiment 2
[0113] This test was trained on the WildDeepfake dataset and tested on the Celeb-DF (v2) dataset, and the results are shown in Table 2 below.
[0114] Table 2 Performance comparison of different methods on the Celeb-DF (v2) dataset
[0115] Method Accuracy(%) Method 1: MesoNet 49.11 Method 2: Xception 51.87 Method 3: Capsule 53.0 Method 4: ADDNet 62.12 <![CDATA[Method 8: F 3 -Net]]> 60.88 Method of the present invention: RED 65.25
[0116] As can be seen from Table 2, the generalization performance of the method of the present invention is better than that of the other five modern detection networks.
[0117] Combining Table 1 and Table 2, it can be seen that although the existing method 8F 3 -Net is superior to the method of the present invention in terms of in-dataset evaluation, the generalization ability of the method of the present invention is 4.37% higher than that of this method, indicating that the baseline proposed by the present invention has good generalization ability.
[0118] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention belongs, without departing from the concept of the present invention, several simple deductions or substitutions can be made, and all should be regarded as belonging to the protection scope of the present invention.
Claims
1. A face forgery detection method based on the distribution of reconstruction errors, characterized in that Including: Construct a face forgery detection network model including a local transformer reconstruction module, a global convolution reconstruction module, and a forgery detection module based on a dual-stream reconstruction strategy; Use a random masking strategy to perform data augmentation on the image to be detected to obtain a pixel sequence; Based on the pixel sequence, use the local transformer reconstruction module to perform image reconstruction to obtain a local reconstructed image; meanwhile, use the global convolution reconstruction module to reconstruct the image to be detected to obtain a global reconstructed image; Based on the reconstruction error distributions corresponding to the local reconstructed image and the global reconstructed image, merge the local reconstructed image and the global reconstructed image to obtain a fused image; Use the forgery detection module to predict the fused image to obtain a face forgery detection result; Among them, the local transformer reconstruction module includes a linear projection unit, a first encoder, and a first decoder; then, based on the pixel sequence, using the local transformer reconstruction module to perform image reconstruction to obtain a local reconstructed image includes: Use the linear projection unit to perform a linear projection on the pixel sequence; Use the first encoder to extract features from the output image after linear projection to obtain a first feature map; Use the first decoder to reconstruct the first feature map to obtain a local reconstructed image.
2. The face forgery detection method based on the reconstruction error distribution according to claim 1, characterized in that Using a random masking strategy to perform data augmentation on the image to be detected to obtain an enhanced pixel sequence includes: Divide the image to be detected into several non-overlapping pixel blocks; Randomly generate 01 masks corresponding to the pixel blocks and uniformly distributed; Select the pixel blocks with a mask value of 1 to generate a pixel sequence.
3. The face forgery detection method based on the reconstruction error distribution according to claim 1, wherein, The global convolution reconstruction module includes a second encoder, a DMFB unit, and a second decoder; then, using the global convolution reconstruction module to reconstruct the image to be detected to obtain a global reconstructed image includes: Use the second encoder to extract features from the image to be detected to obtain a second feature map; Use the DMFB unit to perform multi-scale feature extraction on the second feature map and fuse the extracted multi-scale features to obtain a multi-scale feature map; Use the second decoder to reconstruct the multi-scale feature map to obtain a global reconstructed image.
4. The face forgery detection method based on the reconstruction error distribution according to claim 1, characterized in that Based on the reconstruction error distributions corresponding to the local reconstructed image and the global reconstructed image, merge the local reconstructed image and the global reconstructed image to obtain a fused image includes: Perform error analysis on the local reconstructed image and the global reconstructed image respectively based on the image to be detected, and correspondingly obtain a local reconstruction error distribution and a global reconstruction error distribution; Add the local reconstruction error distribution to the local reconstructed image and add the global reconstruction error distribution to the global reconstructed image; Fuse the local reconstructed image and the global reconstructed image with the added reconstruction errors to obtain a fused image.
5. The face forgery detection method based on the reconstruction error distribution according to claim 1, characterized in that The forgery detection module includes a third encoder and a linear layer; then, using the forgery detection module to predict the fused image to obtain a face forgery detection result includes: Use the third encoder to extract features from the fused image; Based on the extracted features, use the linear layer for prediction to obtain the face forgery detection result.
6. The face forgery detection method based on reconstruction error distribution according to claim 1, wherein After constructing the face forgery detection network model, it further includes: Construct a loss function to train the network model; wherein, the expression of the loss function is: Among them, W represents the width of the image to be detected, H represents the height of the image to be detected, || ||2 represents the L2 norm, I represents the image to be detected, Output a and Output a represent the local reconstructed image and the global reconstructed image respectively, y represents the true label, represents the predicted label.
Citation Information
Patent Citations
Method and system for detecting face forgery of double-stream video based on multiple clues
CN114596608A
Deep counterfeit video detection method based on double fine-grained artifacts
CN115019370A