Deep fake face detection method and device based on space-frequency domain reconstruction

By combining space-frequency domain reconstruction and attention learning in deep forgery detection, the problem of insufficient detection capabilities of the prior art under low resolution and high compression conditions is solved, and higher robustness and detection accuracy are achieved.

CN120148082APending Publication Date: 2025-06-13ZHEJIANG UNIV OF TECH

Patent Information

Application Number
CN202510039375.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Existing deep forgery detection technologies are difficult to effectively detect forgery traces under low resolution and high compression conditions, and the model's dependence on global information leads to poor generalization capabilities.

Method used

The method based on space frequency domain reconstruction is adopted, combining spatial domain and frequency domain characteristics, and the spatial domain learning is guided through the frequency domain to capture the forged traces at different compression rates, and the artifacts are accurately captured through reconstruction attention learning.

Benefits of technology

It improves the robustness and generalization ability of the model, and can accurately detect depth forgery in low resolution and high compression environments, improving detection accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148082A_ABST
    Figure CN120148082A_ABST
Patent Text Reader

Abstract

The invention discloses a deep fake face detection method and device based on space-frequency domain reconstruction, and the method comprises the steps: firstly, training a model through a known data set until a detection result is kept good in a verification set, and storing a weight parameter of the model as a detection model; and taking new video data to be detected as the input of the model, and obtaining a detection result after model detection. In order to solve the problem that the model detection accuracy is reduced due to the fact that quality compression exists in social media transmission in deep fake face detection, the method still achieves a good detection effect under the condition that data are compressed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of deepfake detection, and particularly relates to a face forgery detection method and device based on spatio-temporal frequency domain reconstruction. Background Art

[0002] With the rapid development of technologies such as deep learning and generative adversarial networks, the face synthesis technology based on deepfake has been continuously improved, and the methods for generating forged faces have become increasingly diverse. The tampered multimedia content is widely spread through the network, misleading the public's perception and posing a new threat to privacy and even national security. Therefore, using deepfake detection to combat Deepfake has become an important issue in the current research field.

[0003] The current deepfake detection technology faces multiple challenges. First, traditional methods (such as Xception, EfficientNet) mainly rely on global features for detection, which makes the detector prone to overfitting, resulting in poor generalization ability and easy misjudgment when detecting unseen data. Second, in the compressed deepfake images, the global features are not obvious, making it difficult for the model to capture the forgery traces, thus reducing the detection effect of the detector.

[0004] To address this situation, deepfake detection methods need to possess robust capabilities, that is, to ensure a certain level of accuracy even under low-resolution or compressed conditions. The deepfake face detection method based on spatio-frequency domain reconstruction is used to detect forgery cases such as lossy compression on social media. For example, "A Method and Related Devices for Detecting Deepfake Face Images" disclosed in CN202310047277.9 provides a method and related devices for detecting deepfake face images. This method provides a deepfake face image detection technology by calculating the similarity of feature vectors and generating class label vectors; inputting the feature vectors into a classification network for training, and optimizing the model parameters until the loss approaches a preset threshold for forgery detection. This method relies on extracting multi-dimensional feature vectors from images for training and detection. However, in low-resolution and highly compressed environments in real life, the image quality deteriorates and details are lost, resulting in the inability to accurately extract key features. Especially in compressed images, due to the encoding method of the compression algorithm and the interference of noise, forgery traces may be confused with noise and compression artifacts in the image, making it difficult for the model to effectively distinguish between real and forged content. For example, "A Deepfake Face Detection Method Based on Generative Adversarial Network" disclosed in CN117831137A provides a deepfake face detection method based on generative adversarial network (GAN), mainly used in the field of big data AI image anti-counterfeiting. During the training process, the method only uses real face images, and by fitting the spatial distribution of real samples, it ensures that the generator can reconstruct the noiseless input to achieve the effect of defending against forgery. During the training process of this method, only real face images are used, and the variability of images in different propagation media is not fully considered. In practical applications, images and videos often undergo processing such as compression, noise, or resolution reduction in different media, and these factors may cause differences between the reconstructed images generated by the generator and the actual input images, thereby affecting the detection accuracy. In addition, since the generator relies on deviation as the rule for judging forgery, this simple anomaly judgment may not be able to handle complex and diverse forgery techniques, especially under low-resolution and high-compression conditions, which may lead to misjudgments.

[0005] In view of the deficiencies of existing methods, the present invention proposes a deepfake face detection method and device based on spatio-frequency domain reconstruction. This method combines spatial domain and frequency domain features, enabling the model to capture global information while effectively capturing forgery traces at different compression rates by learning high-frequency features in the frequency domain, thereby enhancing its robustness and generalization ability. In addition, through reconstruction attention learning, the model can more accurately capture the artifacts in deepfakes, further improving the accuracy of detection. Summary of the Invention

[0006] Aiming at the deficiencies of existing forgery detection methods in terms of robustness and generalization ability, the present invention proposes a forgery detection method and device based on the combination of spatial and frequency domains and reconstruction. The present invention guides the spatial domain learning through the frequency domain, enabling the model to maintain good accuracy under different compression rates and enhancing the generalization ability of the model. And through error reconstruction, the dependence of the model on global information is significantly reduced, thereby achieving higher accuracy in deep forgery detection.

[0007] A forgery detection method based on the combination of spatial and frequency domains and reconstruction of the present invention includes the following steps:

[0008] S1 Preprocess the input face video data, extract video frames and perform face detection and alignment, and at the same time perform data augmentation on each frame;

[0009] S2 Extract the global spatial features and high-frequency features in the frequency domain of each frame, and use the high-frequency features to guide the learning of spatial features to capture fine-grained forgery traces.

[0010] S3 Reconstruct by combining the high-dimensional features after the fusion of the spatial domain and the frequency domain, strengthen the learning of forgery features through the residual attention mechanism, and perform deep forgery detection.

[0011] The second aspect of the present invention relates to a deep forgery face detection device based on spatial and frequency domain reconstruction, including a memory and one or more processors. Executable code is stored in the memory. When the one or more processors execute the executable code, it is used to implement a deep forgery face detection method based on spatial and frequency domain reconstruction of the present invention.

[0012] The third aspect of the present invention relates to a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements a deep forgery face detection method based on spatial and frequency domain reconstruction of the present invention.

[0013] The technical concept of the present invention is as follows: First, obtain a deep forgery face dataset for training a neural network. During the training process, preprocess the video data, including extracting video frames, extracting binary masks for each frame, and generating a configuration file based on the file path. Subsequently, send the video frames into the spatial domain and frequency domain backbone networks respectively to extract global features and high-frequency features. Next, perform multi-scale fusion on the global features and high-frequency features to guide the learning of global features. Then, reconstruct the fused high-dimensional features and use the reconstruction attention to learn specific forgery regions. Finally, send the result into the classification module to output the forgery detection result. After the model training is fitted, the compressed videos in social media can be subjected to the same preprocessing operations and input into the trained neural network model. The model will give the corresponding true or false detection results.

[0014] The beneficial effects of the present invention are mainly manifested in:

[0015] The present invention uses a frequency-domain enhancement method to guide the model's learning of high-frequency components by means of multi-stage fusion of the learned high-frequency features. The model not only learns global features but also emphasizes the importance of high-frequency features in global features, improving the robustness and generalization ability of deepfake face detection. Secondly, the model performs reconstruction error attention on the high-dimensional fusion features, making the model pay more attention to the real forgery area, enabling the model to accurately capture the fine-grained feature differences between genuine and fake images and improving the model's detection accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It is a schematic flow chart of a deepfake face detection method based on spatio-frequency domain reconstruction according to the present invention;

[0017] Figure 2 It is a schematic diagram of the model structure of the deepfake face detection method based on spatio-frequency domain reconstruction according to the present invention;

[0018] Figure 3 It is a schematic diagram of the device of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] The technical solution of the present invention will be further described below with reference to the accompanying drawings.

[0020] Embodiment 1

[0021] Refer to Figure 1 , this embodiment relates to a deepfake face detection method based on spatio-frequency domain reconstruction, including the following steps:

[0022] S1 Preprocess the input face video data, extract video frames, perform face detection and alignment, and perform data augmentation on each frame.

[0023] First, perform preprocessing operations on the deepfake face video dataset, and use a face extraction and alignment tool library to perform face detection and extraction operations on each video frame. For forged video frames, additional binary masks for the forged areas need to be processed; next, save the authenticity information, binary forged areas, and storage location paths of each video frame in the form of a configuration file for easy reading during model training. Finally, divide the extracted video frame data into a training set, a validation set, and a test set for model training and testing.

[0024] Secondly, in order to improve the generalization ability of the model and avoid overfitting to the original data, common data augmentation strategies are performed on the training data. By rotating, compressing, and performing color transformation on the original video frames, diverse training samples are generated, enabling the model to learn different representations of the same image during training, thereby improving the detection performance of the model.

[0025] S2 extracts the global spatial features and high-frequency features in the frequency domain for each frame, and uses the high-frequency features to guide the learning of spatial features to capture fine-grained forgery traces.

[0026] The spatial domain feature extraction module in the present invention uses the Xception model as the main feature extractor. This model extracts high-level features of the image through deep convolutional and depthwise separable convolutional layers, and maps the input image into a high-dimensional feature space, providing a basis for subsequent analysis and classification tasks. At the same time, the present invention uses the SRM edge filter to extract the high-frequency part of the image to capture robust forgery traces. Finally, the high-frequency features and global features are fused. Through multi-scale convolution, spatial domain aggregation and channel domain enhancement, the spatial domain and frequency domain features are effectively fused, thereby enhancing the final feature expression ability of the model.

[0027] For the input image with a feature dimension of where b represents the batch size, h and w represent the height and width of the input image respectively, c represents the number of channels (usually c = 3 for RGB images), the Xception model receives an image of size (b×c×h×w) and extracts features from it through multiple convolutional operations.

[0028] After passing through the Xception model, the image will be mapped to a high-dimensional feature space F_output through a series of deep convolutional operations, and its dimension is:

[0029] To capture the high-frequency features of the image, the SRM filter is used here to extract the high-frequency components. Combining convolutional operations and activation function constraints, the high-frequency features in the input image are explicitly extracted. This filter includes three types: edge enhancement filter, texture enhancement filter, and directional high-frequency filter, which are used to capture different forms of high-frequency information respectively.

[0030] Edge enhancement filter F 1 is as follows:

[0031]

[0032] Texture enhancement filter F 2 is as follows:

[0033]

[0034] Second-order horizontal enhancement filter F 3 is as follows:

[0035]

[0036] To enhance the expression ability of the model, the extracted high-frequency features and global features are processed by multi-scale feature fusion here.

[0037] First, use convolutional kernels of multiple different scales (3×3, 5×5, 7×7) to process the input features in the spatial domain and capture multi-scale local information. Then, after global pooling of the features, aggregate important spatial information through the spatial attention mechanism.

[0038] By performing convolutional operations of different scales, multi-scale features are extracted, which helps the neural network capture multi-level information from details to the whole in the image and has good robustness. The formula is as follows:

[0039] X = Concat(X s , X f ) (4)

[0040]

[0041] where X s represents the global information in the spatial domain extracted, X f represents the high-frequency information in the frequency domain extracted, and X represents the result after concatenating the two.

[0042] Then, the multi-scale features are fused by element-wise addition, adding the values of the features of different scales at each position to enhance the expression of multi-scale information. The formula is as follows:

[0043]

[0044] Finally, spatial aggregation is performed on the multi-scale fused features, which means that after fusing the features from different scales, these features are spatially integrated in some way, and then more representative features or higher-level semantic information are extracted. Spatial aggregation helps compress the spatial dimension of multi-scale features. The formula is as follows:

[0045] F sp = σ(Conv 7×7 (Pool(F s )))·F s (9)

[0046] where Pool(·) is the global pooling operation; Conv 7×7 (·) is a 7×7 convolution; σ(·) is the Sigmoid activation function.

[0047] Secondly, importance modeling is performed on the feature channels through global average pooling and per-channel non-linear activation. The main purpose is to dynamically adjust the weight of each channel by modeling the importance of the feature channels, thereby enhancing the model's attention and discrimination ability for useful information. The formula is as follows:

[0048] F c = ReLU(Conv1×1 (AvgPool(x))) (10)

[0049] Then, the enhanced channel weights are broadcasted and fused with the spatial features. This process applies the channel attention information to the entire feature map, thereby weighting and adjusting the features at each spatial position to strengthen the performance of the key regions. The formula is as follows:

[0050] F enh = F sp ·F c (11)

[0051] Finally, by combining the spatial domain and frequency domain features, the final fusion output is obtained through per-channel weighted summation.

[0052] Y = F enh + X (12)

[0053] S3 reconstructs by combining the high-dimensional features after spatial and frequency domain fusion, strengthens the learning of forged features through the residual attention mechanism, and conducts deepfake detection.

[0054] First, in order to recover the original mask information in the input image, we designed and trained a reconstruction network F based on the encoder-decoder structure. The input mask is a binary image mask representing the forged area and the real area. To enable the model to learn robust representations that are different between real and forged faces, we added some white noise during the training process to obtain the value X enc '. The standard deviation of the noise is dynamically adjusted according to the training process, aiming to simulate the noise interference in the real world and prompt the model to learn the anti-interference ability against interference signals. The image reconstruction formula is as follows:

[0055] X enc ' = X enc + Z noise (13)

[0056] where X enc is the representation of the high-dimensional feature vector combined with the spatial domain and frequency domain. Z noise is the added white noise.

[0057] The output of the network is the reconstructed result F(X enc '). We map the decoder output to the interval [0, 1] through the Sigmoid function to obtain the final binary mask M'. That is:

[0058] M' = Sigmoid(F(X enc ')) (14)

[0059] During the reconstruction process, we calculated the reconstruction loss between the input original mask and the reconstructed mask as

[0060]

[0061] where B is the batch size, M i and M i ′ are the ground truth mask and the reconstructed mask of the i-th sample respectively, and ‖·‖ 1 represents the L1 norm. By minimizing this loss function, the model can learn more robust mask reconstruction, thereby improving the model's ability to identify forged regions.

[0062] Secondly, under the constraint of the reconstruction network, we use the reconstruction difference to indicate the regions that may need to be strengthened for learning. Through reconstruction guidance, the model can pay more attention to the regions that may be forged.

[0063] Given the reconstructed mask M′ and the original mask M, first calculate the pixel difference between them to obtain the interpolation mask as:

[0064] m = |M′ - M| (16)

[0065] where |·| represents the absolute value function.

[0066] Perform differential mask attention on the difference and the high-dimensional feature vector output by the intermediate layer. And apply it to the F enh high-dimensional feature output to obtain F enh ′. The calculation formula is as follows:

[0067]

[0068] F att = F′ enh + F enh (18)

[0069] where f 1 and f 2 represent convolution operations. σ is the sigmoid function. represents element-wise multiplication.

[0070] Finally, send F att into the classification network and perform binary classification to distinguish true and false through the cross-entropy loss function. The cross-entropy loss function is as follows:

[0071]

[0072] where N represents the total number of samples. By averaging the loss values of N samples. i represents the index of the sample, that is, calculate for the i-th sample. y iDenote the true label of the $i$-th sample, usually 0 or 1 in a binary classification problem. Denote the model prediction value of the $i$-th sample, which represents the probability that the sample is predicted as the positive class, ranging from 0 to 1, and outputs the result of true / false classification through the sigmoid activation function.

[0073] Example 2

[0074] Such as Figure 3 , this example relates to a deepfake face detection device based on spatio-temporal frequency domain reconstruction, including a memory and one or more processors. Executable code is stored in the memory. When the one or more processors execute the executable code, it is used to implement the deepfake face detection method based on spatio-temporal frequency domain reconstruction in Example 1.

[0075] Example 3

[0076] This example relates to a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the deepfake face detection method based on spatio-temporal frequency domain reconstruction in Example 1.

[0077] The content described in the embodiments of this specification is only an enumeration of the implementation forms of the inventive concept. The protection scope of the present invention should not be regarded as limited to the specific forms stated in the embodiments. The protection scope of the present invention also extends to equivalent technical means that can be conceived by those skilled in the art according to the inventive concept of the present invention.

Claims

1. A deep fake face detection method based on spatial-frequency domain reconstruction, comprising the following steps: S1 preprocesses the input face video data, extracts video frames, performs face detection and alignment, and performs data enhancement on each frame; S2 extracts the global spatial features and high-frequency features in the frequency domain of each frame, and uses the high-frequency features to guide the learning of spatial features to capture fine-grained forgery traces. S3 reconstructs the high-dimensional features after fusion of spatial domain and frequency domain, and strengthens the learning of forged features through the residual attention mechanism to perform deep forgery detection.

2. A deep fake face detection method based on space-frequency domain reconstruction according to claim 1, characterized in that: In step S1, each input video is divided into frames, face extraction and alignment are performed on each frame, and the diversity of the training set is expanded through data enhancement.

3. The deep fake face detection method based on space-frequency domain reconstruction according to claim 1 is characterized in that: In step S2, the model extracts global information in the spatial domain through an encoder, and extracts high-frequency information in the frequency domain through a filter, and combines a multi-domain and multi-scale feature fusion method to guide spatial domain feature learning to forge details through the high-frequency part of the frequency domain.

4. The deep fake face detection method based on space-frequency domain reconstruction according to claim 1 is characterized in that: In step S3, the reconstruction features of the spatial domain and the frequency domain are combined, and the expression of forgery traces is enhanced through the residual attention module to improve the detection accuracy of deep forged faces. Finally, the classification module is used to detect deep forgery of faces on the input face data.

5. The deep fake face detection method based on space-frequency domain reconstruction according to claim 3 is characterized in that: The spatial domain feature extraction is to perform global forgery feature extraction on the input human face, and the frequency domain feature extraction is to perform high-frequency forgery feature extraction on the input human face.

6. The deep fake face detection method based on space-frequency domain reconstruction according to claim 3 is characterized in that: The multi-scale feature fusion module is used to implement high-frequency features to guide global features for learning.

7. The deep fake face detection method based on space-frequency domain reconstruction according to claim 4 is characterized in that: The reconstruction attention is used to enable the model to learn fine-grained forgery traces and to detect deep fake faces.

8. The deep fake face detection method based on space-frequency domain reconstruction according to claim 4 is characterized in that: The classification module is used to achieve the final determination of the image, that is, to determine whether the input image is a forged image or a real image.

9. A deep fake face detection device based on spatial-frequency domain reconstruction, characterized in that: It comprises a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, it is used to implement a deep fake face detection method based on space-frequency domain reconstruction as described in any one of claims 1-8.

10. A computer-readable storage medium, characterized in that: A program is stored thereon, and when the program is executed by a processor, a deep fake face detection method based on space-frequency domain reconstruction as described in any one of claims 1-8 is implemented.

Citation Information

Patent Citations

  • Deeply-forged face image detection method and related equipment

    CN116188956A

  • Deep fake face detection method based on generative adversarial network

    CN117831137A

Cited By

  • Data security sharing method and system

    CN122087881A