Face forgery detection method based on multi-domain clue reconstruction
Through the face forgery detection method of multi-domain clue reconstruction, combined with multi-domain feature fusion and attention mechanism, the problem of difficulty in capturing complex forgery features is solved in the existing technology, and more efficient and general forgery detection capabilities are achieved.
Patent Information
- Application Number
- CN202510235171.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-03
AI Technical Summary
Existing face forgery detection methods based on deep learning are difficult to capture complex forgery features and patterns, especially when facing forgery parts of unknown patterns, detection performance is degraded.
The face forgery detection method based on multi-domain clue reconstruction is adopted, and the network is constructed through the multi-domain feature fusion enhancement module, the encoder-decoder reconstruction network, the adaptive multi-scale feature fusion module and the dual-feature attention module. The high-frequency image, inter-frame residual image and inter-frame consistency information are used to perform modular feature fusion and attention mechanisms to achieve accurate detection of the forged image.
It significantly improves the generalization ability and accuracy of forgery detection, and can better cope with the challenges brought by different forgery methods, especially when dealing with unseen forgery patterns and still work effectively without large amounts of forgery samples.
Smart Images

Figure CN120088833A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deepfake detection, and particularly to a face forgery detection method based on multi-domain clue reconstruction. Background Art
[0002] With the rapid development of deepfake technology, it has become increasingly simple to generate highly realistic and imperceptible forged content. These forged contents not only pose challenges to the human perception system, but also may interfere with the facial recognition system, bringing huge information security risks. Especially in the fields of social media, news dissemination, and virtual reality, the spread of forged images and videos can affect public trust, interfere with public opinion, and be maliciously used to create fake news or cybercrimes.
[0003] For early traditional face image forgery techniques such as splicing and copy-move, various detection methods have been proposed by researchers. For example, specific regions in the image are distinguished whether being tampered with through handcrafted features such as illumination color, color filter array pattern, and blur type inconsistency. These handcrafted features amplify the subtle differences between real images and fake images, and have a high detection accuracy for specific forgery techniques. However, with the development of Deepfake based on deep learning, fake face images have become more diverse and complex, and these early detections based on simple image processing and feature extraction techniques often fail to capture these complex features, resulting in a decline in detection performance. Therefore, in order to capture more advanced forgery features and patterns of fake faces, researchers have turned to deep learning-based methods.
[0004] Deep learning-based detection methods usually identify whether a face image is real or fake by using specific artifacts generated during the deepfake technology process. When generating a forged face image, it is usually necessary to fuse the generated face with the original face background. Due to the defects of the forgery algorithm, it is difficult for facial features such as the head pose, eye color, and teeth of the fake face image to perfectly match the original face, resulting in the generation of inconsistent features such as RGB statistical features, frequency domain information, and facial texture. Current detection methods detect the features of forged face images based on the defects under specific forgery patterns, such as noise features, local texture, and frequency information. Although good results have been achieved on multiple datasets, they rely on the defects generated by a certain specific forgery technology in the training set. Therefore, in practical applications, due to the emergence of new forgery technologies and various types of perturbations, forged parts with unknown patterns can easily cause existing detection methods to fail. Therefore, there is an urgent need for a more efficient and general face forgery detection method. Summary of the Invention
[0005] The purpose of the embodiments of the present invention is to provide a face forgery detection method based on multi-domain clue reconstruction, which comprehensively utilizes high-frequency images, inter-frame residual images, and inter-frame consistency information, and through modular feature fusion and attention mechanisms, realizes accurate detection of forged images.
[0006] To solve the above technical problems, the technical solution adopted by the present invention is a face forgery detection method based on multi-domain clue reconstruction, including the following steps:
[0007] S1. Collect face video data, extract and crop the face images of the X i -th frame and the X i+n -th frame therein, where X i represents the i-th frame image in the face video, and n represents the number of frames in the video interval;
[0008] S2. Construct a convolutional neural network model, including a multi-domain feature fusion enhancement module, an encoder-decoder reconstruction network, an adaptive multi-scale feature fusion module, and a dual-feature attention module, and initialize the parameters of the convolutional neural network model;
[0009] S3. Pass the face images of the X i -th frame and the X i+n -th frame through the multi-domain feature fusion enhancement module to generate a feature-enhanced image;
[0010] S4. Input the feature-enhanced image into the encoder-decoder reconstruction network to generate latent space features and generate a multi-scale reconstructed image through multi-scale, and extract the multi-scale features of the latent space features;
[0011] S5. Generate a fusion feature by passing the latent space features and the multi-scale features through the adaptive multi-scale feature fusion module;
[0012] S6. Perform pixel-level subtraction on the reconstructed image and the face image of the X i+n -th frame to generate a residual image, and generate a global attention feature by passing the fusion feature and the residual image through the dual-feature attention module;
[0013] S7. Calculate the loss function and backpropagate to update the parameters of the convolutional neural network model until the loss function converges or reaches the preset number of iterations, and perform face forgery detection to determine whether there are forgery traces.
[0014] Further, in step S1, the number of frames in the interval is 10 frames.
[0015] Further, in step S3, the generation of the feature-enhanced image is as follows:
[0016] Pass the face image of the X i -th frame through the SRM filter to convert it into a high-frequency image, and calculate X iand X i+n The pixel-level difference between and X is counted as the inter-frame residual image. The result after performing element-wise multiplication on the high-frequency image and the inter-frame residual image is added to X i+n element-wise to output the feature-enhanced image.
[0017] Furthermore, the generation of the feature-enhanced image is expressed as:
[0018] X r = |X i - X i+n |
[0019]
[0020] In the formula, X r represents the residual image, X out represents the feature-enhanced image, f srm represents the convolutional layer containing the SRM filter, represents element-wise multiplication, X i_h represents the high-frequency image, represents element-wise addition, X i+n represents the (i + n)-th frame face image.
[0021] Furthermore, in step S4, in the encoder, the latent space feature Z is extracted through a depthwise separable convolutional network, and the reconstructed image X is generated through the encoder-decoder trained reconstruction network rec by Z, which is expressed as:
[0022]
[0023] The multi-scale features of the latent space feature are extracted, expressed as D 1 , D 2 ,..., D j , where D j represents the feature at the j-th scale.
[0024] Furthermore, in step S5, the generated fusion feature is:
[0025] The latent space feature is mapped to the shared space through the first convolutional neural network to generate the latent space feature representation, and the multi-scale feature is mapped to the shared space through the second convolutional neural network to generate the multi-scale feature representation. The weight parameter is calculated, expressed as:
[0026]
[0027] In the formula, α j represents the weight parameter at the j-th scale in the decoder, e represents the base of the natural logarithm, and f(·) represents the Relu non-linear transformation, For multi-scale feature representation, For latent space feature representation, J represents the total number of scales in the decoder;
[0028] And the fused feature is output through the weight parameter, expressed as:
[0029]
[0030] In the formula, F fusion represents the fused feature, represents element-wise multiplication, and Z represents the latent space feature.
[0031] Furthermore, in step S6, the generation of the global attention feature is expressed as:
[0032]
[0033] In the formula, F att represents the global attention feature, f DFA represents the dual feature attention convolution space, h 1 (R) represents mapping R to f DFA convolution operation, R represents the residual image, h 2 (F fusion ) represents mapping F fusion to f DFA convolution operation, F fusion represents the fused feature.
[0034] Furthermore, in step S7, the loss function includes reconstruction loss, metric learning loss, and cross-entropy loss.
[0035] Furthermore, the reconstruction loss is expressed as:
[0036]
[0037] In the formula, L rec represents the reconstruction loss, N represents the number of samples in the true sample set of the face video data, Y k represents the k-th sample in the true sample set, and X rec represents the reconstructed image;
[0038] The metric learning loss is expressed as:
[0039]
[0040] In the formula, L ml represents the metric learning loss, S Real true sample set in the face video data, S Fake represents the false sample set in the face video data, l represents the l-th false sample, Npositive Denotes the total number of real-real sample pairs, N negative Denotes the total number of real-fake sample pairs, D k,l Denotes the Euclidean distance between k and l.
[0041] The cross-entropy loss is expressed as:
[0042]
[0043] In the formula, L cls Denotes the cross-entropy loss, log() denotes the logarithmic operation, y denotes the actual label of the input image including forged face or real face, Denotes the probability value predicted by the convolutional neural network model for a real face.
[0044] The beneficial effects of the present invention are:
[0045] (1) The present invention uses a multi-domain feature fusion enhancement module to combine high-frequency images, inter-frame residual images, and inter-frame consistency information for multi-domain feature extraction, significantly improving the generalization ability and accuracy of forgery detection. Through the fusion of these information, the method of the present invention can better cope with the challenges brought by different forgery methods, especially performing excellently when dealing with unseen forgery patterns.
[0046] (2) The present invention uses an unsupervised encoder-decoder reconstruction network. By training on real face data, the system can still work effectively without a large number of forged samples and can accurately amplify the difference between forged and real samples.
[0047] (3) The present invention optimizes the expression of global features through a dual-feature attention mechanism by calculating the residual image between the input image and the reconstructed image and based on this residual image and the fused features. Especially in complex forged images, this mechanism can effectively focus on the forged area and improve the classification accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0049] Figure 1 Is the overall framework structure diagram of the present invention.
[0050] Figure 2 Is the flowchart of the adaptive multi-scale feature fusion module of the present invention.
[0051] Figure 3 It is the flow chart of the dual-feature attention module of the present invention.
[0052] Figure 4 It is the Grad-CAM diagram of the present invention. Specific embodiments
[0053] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0054] As Figure 1 shown, the present invention provides a face forgery detection method based on multi-domain clue reconstruction. Through modular design and multi-domain feature fusion, efficient detection and accurate judgment of face forgery are achieved, and it has broad application prospects. The present invention includes the following steps:
[0055] S1. Collect video data in the face video datasets FF++, DFDC, and Celeb-V2, extract the face images of the X i th frame and the X i+n th frame, and crop them using the dlib method. X i represents the i-th frame image in the face video, and n represents the number of frames in the video interval.
[0056] S2. Build a convolutional neural network model, including a multi-domain feature fusion enhancement module (MDFBE), an encoder-decoder reconstruction network, an adaptive multi-scale feature fusion module (AMSFF), and a dual-feature attention module (DFA), and initialize the parameters of the convolutional neural network model.
[0057] S3. Input the face images of the X i th frame and the X i+n th frame into the multi-domain feature fusion enhancement module. Convert the face image of the X i th frame into a high-frequency image X i_h using the SRM filter, calculate the pixel-level difference between X i and X i+n and record it as the inter-frame residual image X r . Perform element-wise addition on the result of element-wise multiplication of X i_h and X r and X i+n . Mix X i_h , X r and X i+n and output the feature enhancement image X out , which is expressed as:
[0058] X r = |X i - X i+n |
[0059]
[0060] In the formula, f srm represents the convolutional layer containing the SRM filter, represents element-wise multiplication, represents element-wise addition, X i+n represents the face image of the (i + n)-th frame.
[0061] In the present invention, X out is kept to match the size of the input face image to ensure that the integrity of the image is not affected.
[0062] S4. Input the feature-enhanced image X out into the encoder-decoder reconstruction network. In the encoder, extract the latent space feature Z through the depthwise separable convolution (Xception) network, and through the reconstruction network trained by the encoder-decoder regenerate Z into the reconstructed image X rec , which is expressed as:
[0063]
[0064] Extract the multi-scale feature representation of Z as D 1 , D 2 ,..., D j , where D j represents the feature at the j-th scale.
[0065] In the present invention, the task of the decoder is to gradually restore and generate the reconstructed image from the latent space feature generated by the encoder. This process requires regenerating the details of the image at different scales, and the features at different scales in the image represent different levels of information in the image. Low-level features usually capture edge textures and other details, while high-level features pay more attention to semantic information or global structure, and these features are complementary in the information they convey.
[0066] S5. Input Z and D j into the AMSFF, and map them to the shared space through the first convolutional neural network f 1 to generate the latent space feature representation Map D j to the shared space through the second convolutional neural network f 2 to generate the multi-scale feature representation Calculate the weight parameter, which is expressed as:
[0067]
[0068] In the formula, α j represents the weight parameter at the j-th scale in the decoder, e represents the base of the natural logarithm, f(·) represents the Relu non-linear transformation, and J represents the total number of scales in the decoder.
[0069] And output the fused feature through the weight parameter, which is expressed as:
[0070]
[0071] In the formula, F fusion represents the fused feature.
[0072] In order to make full use of the features captured by the decoder in the reconstruction network for final classification, the present invention proposes an adaptive multi-scale feature fusion module, which fuses the multi-scale feature representation of the decoder and the latent space feature representation of the encoder. By fusing these features, the information that may be missed when relying on a single scale can be compensated, so as to ensure that the model makes decisions based on a more complete and richer information source. This complementarity enhances the ability of the model to capture detailed and semantic information in the image.
[0073] S6. Subtract X rec from X i+n at the pixel level to obtain the residual image R, and establish the correlation between F fusion and R, which is expressed as:
[0074] F att = f DFA (h 1 (R), h 2 (F fusion ))
[0075] In the formula, F att represents the global attention feature, f DFA represents the dual-feature attention convolution space, h 1 (R) represents mapping R to f DFA convolution operation, and h 2 (F fusion ) represents mapping F fusion to f DFA convolution operation.
[0076] S7. Calculate the loss function, which includes the reconstruction loss, the metric learning loss, and the cross-entropy loss. Among them, the reconstruction loss is expressed as:
[0077]
[0078] In the formula, L recdenotes the reconstruction loss, N denotes the number of samples in the real sample set of the face video data, and Y k denotes the k-th sample in the real sample set.
[0079] The metric learning loss is expressed as:
[0080]
[0081] In the formula, L ml denotes the metric learning loss, S Real the real sample set in the face video data, S Fake denotes the fake sample set in the face video data, l denotes the l-th fake sample, and N positive denotes the total number of real-real sample pairs, N negative denotes the total number of real-fake sample pairs, D k,l denotes the Euclidean distance between k and l.
[0082] The cross-entropy loss measures the difference between the predicted probability distribution of the convolutional neural network model and the actual label distribution, and is expressed as:
[0083]
[0084] In the formula, L cls denotes the cross-entropy loss, y denotes the actual label of the input image including 0 (forged face) or 1 (real face), log() denotes the logarithmic operation, denotes the probability value that the convolutional neural network model predicts a real face.
[0085] Calculate the overall loss, which is expressed as:
[0086] L = L cls + β rec L rec + β ml L ml
[0087] In the formula, L denotes the overall loss, β rec denotes the hyperparameter of the reconstruction loss, β ml denotes the hyperparameter of the metric learning loss.
[0088] Repeat training the convolutional neural network model until the loss function converges or reaches the preset number of iterations, and perform face forgery detection to determine whether there are forgery traces.
[0089] In the present invention, metric learning is incorporated into the latent space features output by the encoder to ensure that the features of similar real faces are mapped more closely in the latent space, while further separating real faces from forged faces. During training, the model selects any two real human faces to form a (real, real) sample pair, calculates their Euclidean distance in the latent space, and minimizes this distance through a loss function. This ensures that the features of real faces cluster together in the latent space, forming a compact feature distribution. For each pair of real face and forged face (real, fake) sample pairs, their Euclidean distance in the latent space is calculated, and this distance is maximized through a loss function to ensure their effective separation in space. A joint training strategy is used to combine metric learning with the reconstruction task. During training, the model not only minimizes the image reconstruction error, but also optimizes the feature distribution in the latent space through the metric learning loss function. And during the training phase, real face data is used to update the model parameters to ensure that the network learns as many real face features as possible.
[0090] The present invention uses two video frames with an interval of 10 frames as input. If the interval number of frames is too large, it will lead to inconsistencies in facial expressions and movements, thus hindering the subsequent reconstruction process and affecting the performance of the model. The present invention sets the interval number of frames at 5, 10, 15, and 20 respectively in the FF++ and Celeb-DF datasets to detect the impact on the model performance. As shown in Table 1, it can be seen that when the interval number of frames is 10, the model prediction accuracy is better.
[0091] Table 1 Influence of video interval number of frames on model performance
[0092] Inter-frame distance FF++ Celeb-DF 5 0.991 0.978 10 0.997 0.999 15 0.985 0.999 20 0.961 0.977
[0093] As Figure 2 shown is the adaptive multi-scale feature fusion module proposed by the present invention. This module adaptively fuses the multi-scale outputs of the decoder with the output of the encoder, promoting the complementary integration of information across different scales.
[0094] As Figure 3 shown is the diagram of the dual feature attention module proposed by the present invention. This module integrates the input residual image with the fused features output by the AMSFF module through a cross-attention mechanism and further processes them using the attention mechanism.
[0095] As Figure 4 shown is the visualization heat map (Grad-CAM) diagram of the present invention. The Grad-CAM visualization shows the regions of interest of the method of the present invention. The highlighted regions indicate that the present invention has higher attention to them, showing subtle artifacts in expressions and textures.
[0096] To better analyze and compare the results of face forgery detection methods, the present invention uses ACC (accuracy) and AUC (area under the curve) to evaluate the forgery image detection method. As shown in Table 2, the detection comparison results of the method of the present invention and existing algorithms on the FF++, Celeb-DF, and DFDC data sets are presented. It can be seen that the method of the present invention can better detect forged face images and achieve the highest detection accuracy within the data sets.
[0097] Table 2 Results of Face Forgery Detection Methods
[0098]
[0099] To evaluate the generality of the method of the present invention on different data sets, the present invention conducts cross-data set experiments by training and testing on different data sets. As shown in Table 3, the method of the present invention has an average improvement of 9.3% on the Celeb-DF data set compared to other methods and an average improvement of 4.5% on the DFDC data set compared to other methods. It can be seen that the method of the present invention is applicable to various data sets.
[0100] Table 3 Cross-Data Set Experiments
[0101]
[0102] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment.
[0103] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.
Claims
1. A face forgery detection method based on multi-domain clue reconstruction, characterized in that: The following steps are involved: S1. Collect face video data and extract the Xth i Frame with X i+n Frame face image and crop it, X i represents the i-th frame image in the face video, and n represents the number of video interval frames; S2. Construct a convolutional neural network model, including a multi-domain feature fusion enhancement module, an encoder-decoder reconstruction network, an adaptive multi-scale feature fusion module, and a dual-feature attention module, and initialize the convolutional neural network model parameters; S3, the X i Frame with X i+n The facial image of the frame is used to generate a feature enhanced image through a multi-domain feature fusion enhancement module; S4, inputting the feature enhanced image into an encoder-decoder reconstruction network to generate latent space features and generate a reconstructed image through multi-scale, and extracting multi-scale features of the latent space features; S5, generating fusion features by combining latent space features and multi-scale features through an adaptive multi-scale feature fusion module; S6. Combine the reconstructed image with the Xth i+n The facial image of the frame is subtracted at the pixel level to generate a residual image, and the fusion feature and the residual image are used to generate a global attention feature through a dual-feature attention module; S7. Calculate the loss function and back-propagate to update the convolutional neural network model parameters until the loss function converges or reaches a preset number of iterations, and perform face forgery detection to determine whether there are traces of forgery.
2. The face forgery detection method based on multi-domain clue reconstruction according to claim 1 is characterized in that: In step S1, the interval frame number is 10 frames.
3. The face forgery detection method based on multi-domain clue reconstruction according to claim 1 is characterized in that: In step S3, the feature enhanced image is generated as follows: The X i The frame face image is converted into a high-frequency image through the SRM filter, and X is calculated. i and X i+n The pixel-level difference between the two is counted as the inter-frame residual image, and the result of element-wise multiplication of the high-frequency image and the inter-frame residual image is added to X i+n Perform element-wise addition and output a feature-enhanced image.
4. The face forgery detection method based on multi-domain clue reconstruction according to claim 3 is characterized in that: The generated feature enhanced image is expressed as: X r =|X i -X i+n | In the formula, X r represents the residual image, X out represents the feature enhanced image, f srm represents a convolutional layer containing an SRM filter, represents element-wise multiplication, X i_h represents a high-frequency image, represents element-wise addition, X i+n Represents the i+nth frame of the face image.
5. The face forgery detection method based on multi-domain clue reconstruction according to claim 1 is characterized in that: In step S4, the latent space feature Z is extracted by a deep separable convolutional network in the encoder and reconstructed by a reconstruction network trained by the encoder-decoder Generate the reconstructed image X through multi-scale Z rec , expressed as: Extract multi-scale features of latent space features, denoted as D1, D2, ..., D j , where D j Represents the features at the jth scale.
6. The face forgery detection method based on multi-domain clue reconstruction according to claim 1 is characterized in that: In step S5, the generated fusion features are: The latent space features are mapped to the shared space through the first convolutional neural network to generate the latent space feature representation, and the multi-scale features are mapped to the shared space through the second convolutional neural network to generate the multi-scale feature representation, and the weight parameters are calculated, which is expressed as: In the formula, α j represents the weight parameter on the jth scale in the decoder, e represents the base of the natural logarithm, f(·) represents the ReLU nonlinear transformation, is a multi-scale feature representation, is the latent space feature representation, J represents the total number of scales in the decoder; And the fusion features are output through the weight parameters, expressed as: In the formula, F fusion represents the fusion feature, represents element-wise multiplication, and Z represents the latent space features.
7. The face forgery detection method based on multi-domain clue reconstruction according to claim 1 is characterized in that: In step S6, the generated global attention feature is expressed as: In the formula, F att represents the global attention feature, f DFA represents the dual-feature attention convolution space, h1(R) represents mapping R to f DFA Convolution operation, R represents the residual image, h2(F fusion ) indicates that F fusion Mapping to f DFA Convolution operation, F fusion Represents fusion features.
8. The face forgery detection method based on multi-domain clue reconstruction according to claim 1 is characterized in that: In step S7, the loss function includes reconstruction loss, metric learning loss and cross entropy loss.
9. The face forgery detection method based on multi-domain clue reconstruction according to claim 8 is characterized in that: The reconstruction loss is expressed as: Where, L rec represents the reconstruction loss, N represents the number of samples in the real sample set in the face video data, and Y k represents the kth sample in the real sample set, X rec represents the reconstructed image; The metric learning loss is expressed as: Where, L ml represents the metric learning loss, S Real The real sample set in face video data, S Fake represents the set of false samples in the face video data, l represents the lth false sample, N positive Represents the total number of true-true sample pairs, N negative represents the total number of real-fake sample pairs, D k,l represents the Euclidean distance between k and l; The cross entropy loss is expressed as: Where, L cls represents the cross entropy loss, log() represents the logarithmic operation, y represents the actual label of the input image, including fake faces or real faces, Represents the probability value of the convolutional neural network model predicting a real face.
Citation Information
Cited By
Deep counterfeit image detection method, device, equipment and medium
CN120746983A
Deepfake image detection method, apparatus, device and medium
CN120746983B
Diffusion reconstruction error enhanced diffusion model forged image positioning method and device
CN120894287A
Face image depth forgery detection method and system
CN121600582A
A method and system for detecting deepfake facial images
CN121600582B