Deep fake detection method based on space-frequency feature integration and dynamic edge optimization
By adopting the methods of space-frequency feature integration and dynamic edge optimization in deep forgery detection, the problem of limited generalization capability and detection accuracy in the prior art dealing with complex or unknown forgery modes is solved, and higher generalization capability and detection accuracy are achieved, especially in the case of category imbalance.
Patent Information
- Application Number
- CN202510282109.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2045-03-11
AI Technical Summary
Existing deep forgery detection methods have limited generalization capabilities and detection accuracy when dealing with complex or unknown forgery modes.
The deep forgery detection method based on space-frequency feature integration and dynamic edge optimization is adopted, and the spatial domain and frequency domain features of the image are fused through the space-frequency feature integration module, and the model is optimized through the reality-perceptual boundary loss function, and the boundaries are dynamically adjusted to deal with the problem of category imbalance.
It improves the generalization ability and detection accuracy of deep forgery detection in unknown environments, can more effectively capture complex features of forged images, and significantly improve classification performance under category imbalance.
Smart Images

Figure CN119785193B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of deepfake detection, and particularly relates to the design of a deepfake detection method based on the integration of spatial-frequency features and dynamic edge optimization. Background Art
[0002] The breakthrough of deep learning technology has significantly expanded the boundaries of digital content creation, and deep generative models have become the core of this field. However, the double-edged sword effect of technology has also given rise to the proliferation of fake face images and videos. These highly realistic contents quickly spread in the cyberspace, which may not only mislead the public, but also threaten the social trust system in terms of synthesizing fake news. With the progress of deepfake technology, its applications in aspects such as privacy infringement and speech tampering pose a severe challenge to the authenticity of information. To address this issue, deepfake detection technology has emerged as the times require.
[0003] Deepfake detection technology is constantly innovating with the rapid development of generative models. The significant breakthrough of generative adversarial networks (GANs) in the field of image generation has provided new ideas for deep learning detection algorithms, enabling them to capture the unique features of GAN-generated images, thereby improving the detection accuracy. However, new generative technologies such as diffusion models, with higher fidelity, have posed new challenges to existing detection technologies. These advancements have made forged images more delicate, further weakening the effectiveness of traditional detection algorithms.
[0004] In the field of deepfake detection, early methods mainly relied on convolutional neural networks (CNNs) to perform binary classification on facial images through existing network architectures. These methods could relatively effectively identify local features, but only focused on the local information of the images and failed to deeply capture the complexity of forged images, resulting in limited detection effects. To solve this problem, recent research methods have begun to introduce new detection ideas, focusing on identifying specific patterns in forged images, such as noise features, local textures, and frequency information. Although these improvements have increased the detection accuracy in some cases, there are still limitations, especially when dealing with unknown forgery techniques not included in the training set, and the effect drops significantly.
[0005] Currently, there are numerous deepfake detection methods emerging in an endless stream. LSDA adopts a teacher-student model architecture. The teacher module learns features in a specific domain and is trained through a domain loss. The student module extracts features from the teacher to enhance the generalization ability. However, LSDA has relatively high requirements for the diversity of training data, and the generalization effect will be limited when the data is insufficient. UCF decouples image features through a multi-task learning strategy and a conditional decoder, and extracts general forgery features for detection. However, UCF increases the model complexity during the feature decoupling process, and the decision-making process lacks transparency, making it difficult to intuitively explain the detection results. Methods such as SRM and F3Net distinguish between forged and real images by analyzing the frequency components of images. Although they perform well in dealing with known forgery patterns, their detection ability for unknown forgery patterns is still limited. In addition, CORE and RECCE improve the detection algorithm by designing specific loss functions and reconstruction learning. However, these methods are often vulnerable to interference from irrelevant information (such as race, gender, or identity features) and it is difficult to completely eliminate the influence of these factors. On the other hand, models such as Meso4 and Xception, as pure convolutional network structures, fail to fully utilize data augmentation, feature decoupling, or frequency information, so their performance significantly degrades when dealing with unknown deepfake videos.
[0006] In summary, the existing deepfake detection methods have the following deficiencies: (1) Although the existing deepfake detection methods have made certain progress, when dealing with forgery patterns not seen in the training set, the generalization ability and detection effect of the model still need to be improved. (2) The local details of forged images are not captured sufficiently, and the spectral information is not utilized fully, resulting in limited generalization ability and detection accuracy when dealing with complex or unknown forgery patterns. Summary of the Invention
[0007] The purpose of the present invention is to solve the problem that the generalization ability and detection accuracy of existing deepfake detection methods are limited when dealing with complex or unknown forgery patterns, and a deepfake detection method based on spatio-frequency feature integration and dynamic edge optimization is proposed.
[0008] The technical solution of the present invention is as follows: A deepfake detection method based on spatio-frequency feature integration and dynamic edge optimization includes the following steps:
[0009] S1. Obtain the dataset required for deepfake detection and divide the dataset into a training set and a test set.
[0010] S2. Perform preprocessing and data augmentation on the training set and the test set respectively to obtain the processed training set and test set.
[0011] S3. Construct and initialize a deepfake detection network based on spatio-frequency feature integration and dynamic edge optimization.
[0012] S4. Input the processed training set into the deepfake detection network, train the deepfake detection network, and obtain the output probability of the deepfake detection network.
[0013] S5. Construct a authenticity-aware boundary loss function based on the output probability of the deepfake detection network, and optimize the deepfake detection network according to the authenticity-aware boundary loss function.
[0014] S6. Repeat steps S4 - S5 iteratively until the preset number of iterations is reached. At the end of each iteration, input the processed test set into the trained deepfake detection network for testing, and calculate the test metrics of the current deepfake detection network.
[0015] S7. Input the processed test set into the deepfake detection network with the highest test metric value, and output the deepfake detection result.
[0016] Furthermore, the dataset in step S1 includes FaceForensics++, CelebDF-v1, CelebDF-v2, and DFDCP, where FaceForensics++ is used as the training set, and CelebDF-v1, CelebDF-v2, and DFDCP are used as the test sets.
[0017] Furthermore, the preprocessing in step S2 is specifically as follows: perform frame extraction on the training set and the test set respectively, use the Dlib face detection algorithm to detect the faces in each video frame, use the Dlib shape predictor model to align and crop the faces according to the detected facial landmarks, and save the processed face images in separate folders.
[0018] Furthermore, the deepfake detection network based on spatial frequency feature integration and dynamic edge optimization in step S3 includes a spatially - frequency feature integration module SFFI, a pre-trained Xception network, and a fully connected layer connected in sequence.
[0019] Furthermore, step S4 includes the following sub-steps:
[0020] S41. Input the processed training set into the spatially - frequency feature integration module SFFI, fuse the spatial domain and frequency domain features of the image, and output a spatio-frequency fused feature map.
[0021] S42. Input the spatio-frequency fused feature map into the pre-trained Xception network for in-depth feature extraction, and output a refined feature map. 。
[0022] S43. Input the refined feature map into the fully connected layer for processing, and obtain the output probability of the deepfake detection network.
[0023] Furthermore, step S41 includes the following sub-steps:
[0024] S411. Input the processed training set into the spatial-frequency feature integration module SFFI, calculate the per-pixel mean of the three channels of the input image, and obtain a single-channel grayscale image.
[0025] S412. Perform a fast Fourier transform on the grayscale image to obtain a spectrogram.
[0026] S413. Perform frequency-domain offset and amplitude calculation on the spectrogram, and use logarithmic transformation to compress the dynamic range of the spectrogram to obtain a compressed spectrogram.
[0027] S414. Normalize the compressed spectrogram to obtain a normalized spectrogram. .
[0028] S415. Concatenate the normalized spectrogram and the input image along the channel dimension to obtain a spatio-frequency fusion feature map.
[0029] Furthermore, the authenticity-aware boundary loss function constructed in step S5 is:
[0030]
[0031] where represents the authenticity-aware boundary loss function, represents the batch size, represents the number of classes, represents the output probability of the i-th sample on the j-th class, represents the true class of the i-th sample, represents the output probability of the i-th sample on the true class ., represents the true class . represents the scaling factor.
[0032] Furthermore, the test metrics in step S6 include accuracy ACC, area under the curve AUC, and average precision AP.
[0033] The beneficial effects of the present invention are:
[0034] (1) The present invention provides a deepfake detection method based on spatial-frequency features and dynamic edge optimization to achieve deepfake detection in an unknown environment, with high generalization ability and detection accuracy.
[0035] (2) The present invention proposes a Spatial-Frequency Feature Integration Module (SFFI) that combines the original color image with frequency-domain information, enabling the model to capture richer and more useful feature information during the feature extraction stage.
[0036] (3) The present invention proposes a Realness-Aware Boundary Loss Function that successfully addresses the class imbalance problem by dynamically adjusting the boundary, especially improving the classification performance significantly when real image samples are scarce. Description of the Drawings
[0037] Figure 1 The following shows a flowchart of a deepfake detection method based on spatio-frequency feature integration and dynamic edge optimization provided by an embodiment of the present invention.
[0038] Figure 2 The following shows a schematic structural diagram of a deepfake detection network based on spatio-frequency feature integration and dynamic edge optimization provided by an embodiment of the present invention. Detailed Embodiment
[0039] Now, exemplary embodiments of the present invention will be described in detail with reference to the accompanying drawings. It should be understood that the embodiments shown and described in the drawings are merely exemplary, intended to illustrate the principles and spirit of the present invention, rather than limiting the scope of the present invention.
[0040] An embodiment of the present invention provides a deepfake detection method based on spatio-frequency feature integration and dynamic edge optimization, as Figure 1 shown, including the following steps S1 to S7:
[0041] S1. Obtain the dataset required for deepfake detection and divide the dataset into a training set and a test set.
[0042] In an embodiment of the present invention, the dataset required for deepfake detection includes FaceForensics++, CelebDF-v1, CelebDF-v2, and DFDCP, where FaceForensics++ is used as the training set, and CelebDF-v1, CelebDF-v2, and DFDCP are used as the test sets.
[0043] In an embodiment of the present invention, the divided datasets are downloaded from the papers of FaceForensics++ and the official website of Celeb-DF, and the source datasets are all videos ranging from 10 seconds to 30 seconds.
[0044] S2. Perform preprocessing and data augmentation on the training set and the test set respectively to obtain the processed training set and test set.
[0045] In the embodiments of the present invention, the specific method for preprocessing is as follows: perform frame extraction on the training set and the test set respectively, use the Dlib face detection algorithm to detect the faces in each video frame, use the Dlib shape predictor model to align and crop the faces according to the detected facial landmarks, and save the processed face images in a separate folder.
[0046] In the embodiments of the present invention, the data augmentation processing includes methods such as flip_prob, rotate_prob, rotate_limit, blur_prob, blur_limit, brightness_prob, brightness_limit, contrast_limit, quality_lower, and quality_upper.
[0047] S3. Construct and initialize a deepfake detection network based on the integration of spatial frequency features and dynamic edge optimization.
[0048] In the embodiments of the present invention, as Figure 2 shown, the deepfake detection network based on the integration of spatial frequency features and dynamic edge optimization includes a sequentially connected spatial-frequency feature integration module SFFI, a pre-trained Xception network, and a fully connected layer. The spatial-frequency feature integration module SFFI is used to fuse the spatial domain and frequency domain features of the image, enabling the model to more comprehensively identify the subtle differences in forged images. The pre-trained Xception network is used to perform the feature extraction task, which can capture rich feature representations from the input image. These features are then passed to the fully connected layer. In the fully connected layer, through a series of complex calculations, a logits (probability) is finally output to predict the authenticity of the input image.
[0049] S4. Input the processed training set into the deepfake detection network, train the deepfake detection network, and obtain the output probability of the deepfake detection network.
[0050] As Figure 2 shown, step S4 includes the following sub-steps S41 to S43:
[0051] S41. Input the processed training set into the spatial-frequency feature integration module SFFI, fuse the spatial domain and frequency domain features of the image, and output a spatio-frequency fusion feature map.
[0052] As Figure 2 shown, step S41 includes the following sub-steps S411 to S415:
[0053] S411. Input the processed training set into the Spatial-Frequency Feature Integration Module (SFFI), calculate the per-pixel mean for the three channels of the input image to obtain a single-channel grayscale image.
[0054] S412. Perform a fast Fourier transform on the grayscale image to obtain a frequency spectrum image, which can extract texture information and forgery traces under different frequency components.
[0055] S413. Perform frequency domain offset (through centering processing) and amplitude calculation on the frequency spectrum image, and use logarithmic transformation to compress the dynamic range of the frequency spectrum image to obtain a compressed frequency spectrum image, enhancing the visibility of tiny features.
[0056] S414. Normalize the compressed frequency spectrum image to obtain a normalized frequency spectrum image , ensuring that the numerical range of the frequency spectrum data meets the requirements of subsequent calculations.
[0057] S415. Concatenate the normalized frequency spectrum image and the input image along the channel dimension to obtain a spatial-frequency fusion feature map.
[0058] S42. Input the spatial-frequency fusion feature map into the pre-trained Xception network for in-depth feature extraction, and output a refined feature map .
[0059] As Figure 2 shown, the fusion feature map is first fed into the encoder part of Xception. Inside the encoder, the fusion feature map undergoes a series of convolutional and pooling operations, gradually abstracting and compressing to extract the deep features of the image. Subsequently, these features are fed into the middle layer of the network (consisting of several RELU layers and depthwise separable convolutional layers), and the expressive power of the features is further refined and enhanced through an iterative process. Finally, through the reverse operation of the decoder, these deep features are mapped back to the spatial dimension of the original feature map, thus obtaining a refined feature map .
[0060] S43. Input the refined feature map into the fully connected layer for processing to obtain the output probability of the deepfake detection network.
[0061] The output probability logits represent the likelihood that the input image belongs to the real category or the forged category. Finally, this probability value is used as the basis for judging the authenticity of the image, achieving efficient identification of the authenticity of the image.
[0062] S5. Construct a authenticity-aware boundary loss function based on the output probability of the deepfake detection network, and optimize the deepfake detection network according to the authenticity-aware boundary loss function.
[0063] In the task of image authenticity detection, class imbalance is a common challenge, especially when fake images often constitute a large number of samples while real images are relatively scarce. By dynamically adjusting the margin to enhance the model's discriminative ability for real images, this significantly improves the classification performance, especially in the case of severe data imbalance.
[0064] Based on this, the authenticity-aware boundary loss function constructed in the embodiments of the present invention is:
[0065]
[0066] where represents the authenticity-aware boundary loss function, represents the batch size, represents the number of classes, represents the output probability of the i-th sample on the j-th class, represents the true class of the i-th sample, represents the output probability of the i-th sample on the true class , represents the true class 's boundary, and there is , where represents the boundary of the j-th class, represents the true class 's number of samples, represents the maximum boundary value, represents the scaling factor.
[0067] In the embodiments of the present invention, the dynamic boundary of each class is calculated according to the number of samples in each class, and the weight is adjusted through the reciprocal of the square root of the class frequency to generate the corresponding dynamic boundary to solve the class imbalance problem. At the same time, a scaling factor s is introduced to amplify the output of the loss function, so that during the training process, the model pays more attention to those samples that are difficult to classify. Combining the output after dynamic weight adjustment with the cross-entropy loss function improves the classification performance of the model for imbalanced datasets.
[0068] S6. Repeatedly iterate and execute steps S4 - S5 until the preset number of iterations (10 times in the embodiments of the present invention) is reached. At the end of each iteration, the processed test set is input into the trained deepfake detection network for testing, and the test metrics of the current deepfake detection network are calculated.
[0069] In the embodiments of the present invention, the test metrics include accuracy ACC, area under the curve AUC, and average precision AP.
[0070] S7. Input the processed test set into the deepfake detection network with the highest test metric value, and output the deepfake detection result.
[0071] To further illustrate the effectiveness of the method of the present invention, the method of the present invention is compared with other existing methods. For a fair comparison, the official released codes of other methods are used and their experimental settings are followed, where all methods are implemented in the same computing environment and quantitative and qualitative analyses are carried out simultaneously. The 8 methods are specifically as follows: Method 1 is the Meso4 method, which is a deep learning-based image analysis technology; Method 2 is Xception. Like Meso4, Xception is also a pure convolutional network structure focusing on the extraction of image features; Method 3 is F3Net, which uses the frequency components of images to distinguish forged images from real images; Method 4 is SRM, which also adopts image frequency analysis technology to enhance the accuracy of forgery detection; Method 5 is CORE, which focuses on designing specific loss functions and reconstruction learning strategies to improve the detection performance; Method 6 is Recce, which is an innovative detection algorithm that improves the accuracy of forgery detection through different strategies; Method 7 is UCF, which decouples image features through a multi-task learning strategy and a conditional decoder to extract general forgery features for deepfake detection; Method 8 is LSDA, which adopts a teacher-student model architecture, where the teacher module learns domain-specific features and is trained through domain losses; the student module refines features from the teacher module to improve the generalization ability.
[0072] Since the method of the present invention mainly aims at deepfake detection in an unknown environment, the forgery techniques of forged videos in the training set and the test set are from different sources. Tables 1, 2, and 3 give the quantitative comparison results of the precision ACC, area under the curve AUC, and average precision AP of 9 different network structures on 3 datasets.
[0073] Table 1 Comparison results of different methods on CelebDF-v1
[0074]
[0075] Table 2 Comparison results of different methods on CelebDF-v2
[0076]
[0077] Table 3 Comparison results of different methods on DFDCP
[0078]
[0079] In Tables 1, 2, and 3, the percentage counting method was adopted to expand the measurement index by 100 times. From the comparison of the results in Tables 1, 2, and 3, it can be seen that the method of the present invention has higher detection accuracy and better detection coverage for the target to be detected compared to all other methods.
[0080] The effectiveness of each proposed module was verified through ablation experiments. The training set for the ablation experiments was FaceForensics++, and the test sets were CelebDF-v1, CelebDF-v2, and DFDCP datasets.
[0081] Table 4 Comparison results of ablation experiments on CelebDF-v1
[0082]
[0083] Table 5 Comparison results of ablation experiments on CelebDF-v2
[0084]
[0085] Table 6 Comparison results of ablation experiments on DFDCP
[0086]
[0087] The quantitative results are shown in Tables 4 to 6. The method of the present invention selected Xception as the baseline model. When adding modules one by one to the baseline, it was verified that the proposed method had obvious improvements. In addition, the influence of adding a single module to the baseline was tested. The experimental results showed that adding any module improved the detection performance, thus proving the effectiveness of the proposed modules.
[0088] Those of ordinary skill in the art will realize that the embodiments described herein are for helping readers understand the principles of the present invention and should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations that do not depart from the essence of the present invention based on the technical revelations disclosed in the present invention, and these deformations and combinations are still within the protection scope of the present invention.
Claims
1. A deep fake detection method based on space-frequency feature integration and dynamic edge optimization, characterized in that: The following steps are involved: S1. Obtain the data set required for deep fake detection and divide the data set into a training set and a test set; S2, preprocessing and data enhancement processing are performed on the training set and the test set respectively to obtain the processed training set and the test set; S3. Build and initialize a deep fake detection network based on spatial frequency feature integration and dynamic edge optimization; S4. Input the processed training set into the deep fake detection network, train the deep fake detection network, and obtain the output probability of the deep fake detection network; S5. Constructing an authenticity-aware boundary loss function based on the output probability of the deep fake detection network, and optimizing the deep fake detection network according to the authenticity-aware boundary loss function; S6. Repeat steps S4 to S5 until a preset number of iterations is reached. At the end of each iteration, the processed test set is input into the trained deep fake detection network for testing, and the test index of the current deep fake detection network is calculated. S7, inputting the processed test set into the deep fake detection network with the highest test index value, and outputting the deep fake detection result; The step S4 comprises the following sub-steps: S41, inputting the processed training set into the space-frequency feature integration module SFFI, fusing the space domain and frequency domain features of the image, and outputting a space-frequency fusion feature map; S42, input the space-frequency fusion feature map into the pre-trained Xception network for in-depth feature extraction, and output a fine feature map ; S43, fine feature map The input is processed in the fully connected layer to obtain the output probability of the deep fake detection network; The step S41 includes the following sub-steps: S411, input the processed training set into the space-frequency feature integration module SFFI, calculate the pixel-by-pixel mean of the three channels of the input image, and obtain a single-channel grayscale image; S412, performing fast Fourier transform on the grayscale image to obtain a spectrum image; S413, performing frequency domain offset and amplitude calculation on the spectrum graph, and compressing the dynamic range of the spectrum graph by logarithmic transformation to obtain a compressed spectrum graph; S414, normalize the compressed spectrum to obtain a normalized spectrum ; S415, normalize the spectrum It is concatenated with the input image according to the channel dimension to obtain a space-frequency fusion feature map.
2. The deep fake detection method based on space-frequency feature integration and dynamic edge optimization according to claim 1 is characterized in that: The data sets in step S1 include FaceForensics++, CelebDF-v1, CelebDF-v2 and DFDCP, wherein FaceForensics++ is used as a training set, and CelebDF-v1, CelebDF-v2 and DFDCP are used as test sets.
3. The deep fake detection method based on space-frequency feature integration and dynamic edge optimization according to claim 1 is characterized in that: The preprocessing in step S2 is specifically as follows: extracting frames for the training set and the test set respectively, detecting faces in each video frame using the Dlib face detection algorithm, aligning and cropping faces according to detected facial landmarks using the Dlib shape predictor model, and saving the processed face images in a separate folder.
4. The deep fake detection method based on space-frequency feature integration and dynamic edge optimization according to claim 1 is characterized in that: The deep fake detection network based on spatial-frequency feature integration and dynamic edge optimization in step S3 includes a spatial-frequency feature integration module SFFI, a pre-trained Xception network and a fully connected layer connected in sequence.
5. The deep fake detection method based on space-frequency feature integration and dynamic edge optimization according to claim 1 is characterized in that: The authenticity perception boundary loss function constructed in step S5 is: ; in represents the authenticity-aware boundary loss function, represents the batch size, represents the number of categories, represents the output probability of the i-th sample in the j-th category, represents the true category of the i-th sample, Indicates that the i-th sample is in the true category The output probability on Represents the true category The borders of Represents the scaling factor.
6. The deep fake detection method based on space-frequency feature integration and dynamic edge optimization according to claim 1 is characterized in that: The test indicators in step S6 include accuracy ACC, area under the curve AUC and average precision AP.
Citation Information
Patent Citations
Video face forgery detection method and system based on space-frequency time sequence characteristics
CN118072400A
Face forgery detection method and system based on multi-modal collaborative learning
CN118397681A
Cited By
Depth counterfeit image identification method based on differential feature search
CN121583008A