Deep fake face detection method based on adaptive convolution and bidirectional adapter
By using adaptive convolution and bidirectional adapter methods in deep fake face detection, the forged features of the interactive space domain and frequency domain are extracted, and the problem of poor generalization ability in the prior art is solved, and the accurate authenticity detection of unknown data is achieved.
Patent Information
- Application Number
- CN202510258950.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-24
AI Technical Summary
When detecting the authenticity of unknown data, the prior art has poor generalization ability and it is difficult to accurately identify the authenticity of deep fake face pictures and videos.
The deep fake face detection method based on adaptive convolution and bidirectional adapter is adopted. The fake features of the spatial domain and the frequency domain are extracted through the adaptive area dynamic convolution module and the adaptive frequency dynamic filter module, and the feature interaction and fusion is carried out through the bidirectional adapter module to improve the accuracy of detection.
The detection and discrimination of false content of pictures from different sources is realized, the authenticity detection ability of unknown data is improved, and the robustness and generalization ability of the model are improved.
Smart Images

Figure CN120198945A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the detection technology of forged face pictures and videos in digital content, and specifically relates to a deepfake face detection method based on adaptive convolution and bidirectional adapter. Background Art
[0002] The spread of deepfake content poses a severe challenge to social trust and information security. The key lies in accurately analyzing and mining the core features of deepfake content generation, identifying the subtle differences between real images and synthetic images, and achieving effective detection of large-scale spread of forged content. Therefore, designing an efficient detection model to accurately extract deepfake features has become one of the important means to address the threat to the authenticity of digital content. According to the different features used in existing detection models, the detection models can be divided into two categories: those based on spatial domain features and those based on frequency domain features. Most deepfake detection methods based on spatial features detect forged images by capturing low-level visual cues from the spatial domain, and deepfake detection methods based on frequency domain features identify forged traces by analyzing the frequency components of images and capturing subtle frequency changes in forged content.
[0003] The literature [Shiohara K, Yamasaki T. Detecting Deepfakes With Self-Blended Images[C], 2022 IEEE / CVF Conference on Computer Vision and Pattern Recognition, 2022: 18720-18729.] proposed a new type of synthetic training data SBI to detect deepfakes, encouraging the classifier to learn more general forgery features by generating more difficult-to-recognize forged samples. The literature [Guo X, Liu X, Ren Z, et al. Hierarchical Fine-Grained Image Forgery Detection And Localization[C], 2023 IEEE / CVF Conference on Computer Vision and Pattern Recognition, 2023: 3155-3165.] proposed a hierarchical fine-grained network to identify multiple forgery attributes and detect and localize deepfakes. The literature [Jiahe Tian, Cai Yu, Peng Chen, et al. Discriminative Feature Mining Based on Frequency Information and Metric Learning for Face Forgery Detection[J], IEEE Transactions on Knowledge and Data Engineering, 2023, 35(12): 12167-12180.] proposed a frequency-aware discriminative feature learning framework, developing a frequency feature adaptive generation module to adaptively mine frequency features in a completely data-driven manner, getting rid of incomplete prior knowledge. The above methods achieve high accuracy in public datasets. However, when faced with unknown data, it is difficult to accurately identify its authenticity and the generalization ability is poor. Summary of the Invention
[0004] Aiming at the problems in the prior art, the present invention provides a deepfake face detection method based on adaptive convolution and bidirectional adapter, aiming to detect and discriminate false content for pictures from different sources.
[0005] The deepfake face detection method based on adaptive convolution and bidirectional adapter includes the following steps:
[0006] Step 1: Construct a deepfake face detection network, which includes an adaptive region dynamic convolution module, an adaptive frequency dynamic filtering module, and several dynamic interaction modules;
[0007] Step 2: Input the forged images in the training set into the deepfake face detection network for training, including the following steps:
[0008] Step 2.1: Dynamically extract the spatial domain forgery features of different input images through the adaptive region dynamic convolution module;
[0009] Step 2.2: Dynamically extract the frequency domain forgery features of different input images through the adaptive frequency dynamic filtering module;
[0010] Step 2.3: Interact the spatial domain forgery features and the frequency domain forgery features through several dynamic interaction modules in sequence; among them, the previous dynamic interaction module performs downsampling on the feature data after interaction before sending it into the next dynamic interaction module;
[0011] Step 2.4: Concatenate the output of the last dynamic interaction module and send it into the fully connected layer for classification to obtain the detection result of the forged image;
[0012] Step 2.5: Update the model parameters of the deepfake face detection network according to the cross-entropy loss and the SGD optimization strategy to obtain the final deepfake face detection network;
[0013] Step 3: Perform deepfake face detection through the final deepfake face detection network.
[0014] Furthermore: The dynamic interaction module includes two bidirectional adapters and two backbone networks. The backbone network is a swin transformer model. Insert one bidirectional adapter between the W-MSA modules of the two swin transformer models, and insert the other bidirectional adapter between the SW-MSA modules of the two swin transformer models; the two backbone networks are used to respectively input two types of feature data input by the dynamic interaction module, and these two types of feature data are respectively the spatial domain forgery features and the frequency domain forgery features, or respectively the two feature data output by the previous dynamic interaction module;
[0015] Among them, the bidirectional adapter is used to fuse the inputs of the normalization layers before the two W-MSA modules or SW-MSA modules, and add the two outputs after fusion by the bidirectional adapter to the outputs of these two W-MSA modules or SW-MSA modules respectively.
[0016] Specifically, the Adaptive Region Dynamic Convolution Module (ARDConv) includes a convolution module and an attention module. Step 2.1 specifically includes the following steps:
[0017] Step 2.1.1: Generate a guiding feature map Y of the input image through the convolution module h,w,c ;
[0018]
[0019] where h, w, and c are the height, width, and number of channels of the input image respectively; x i is the input image, W c represents the standard convolution of the c-th channel;
[0020] Step 2.1.2: Process the guiding feature map through the attention module to generate an attention weight map A k ;
[0021] Step 2.1.3: Normalize the attention weight map and make the sum of all region weights at each position equal to 1 to obtain the normalized weight map
[0022]
[0023] where, A k (i, j) represents the attention weight at position (i, j), m represents the number of divided regions, represents the degree of belonging of position (i, j) to the k-th region;
[0024] Step 2.1.4: Divide the image into different regions in the spatial dimension according to the weight map and then perform convolution operations on the feature maps of each region respectively to obtain feature representations of different regions;
[0025] Step 2.1.5: Use the attention weight map to multiply the feature representations of different regions element by element to obtain the weighted spatial domain feature representation F s ;
[0026]
[0027] where, Y k is the k-th region feature map, w k is the convolution kernel of the k-th region, Y k *w k represents the operation of region division and convolution, and ⊙ represents element-by-element multiplication.
[0028] Specifically, Step 2.2 is as follows:
[0029] Step 2.2.1: Decompose the input image into different frequency components by downsampling in the learnable latent space;
[0030] Step 2.2.2: Generate weights α for different frequency components through the attention module k , dynamically combining a set of basic filters g k Construct a frequency dynamic filter W and convolve it with the feature representation containing all frequency components to obtain the image frequency domain feature F f ;
[0031]
[0032] F f =W*F (13)
[0033] Among them, g k is the kth basic filter, α k is the attention weight, W is the constructed dynamic filter, k∈{l,m,h}, and F is the un-downsampled feature of the input image.
[0034] Further: different frequencies include low frequency components, medium frequency components and high frequency components;
[0035] f l =Conv↓2(Conv↓2(I)) (14)
[0036] f m =Conv↓2(I)-Conv↓2(Conv↓2(I))↑2 (15)
[0037] f h =Conv(I)-Conv↓2(I)↑2 (16)
[0038] Among them, f l 、f m 、f h They represent low-frequency component, medium-frequency component and high-frequency component respectively, and I represents the input image.
[0039] The beneficial effects of the present invention are as follows: specific features are extracted from different input images through an adaptive regional dynamic convolution module and an adaptive frequency dynamic filtering module to mine specific forgery clues; then, based on the complementarity of spatial domain features and frequency domain features, a bidirectional adapter module is introduced to interactively fuse spatial features and frequency features to complement each other, and extract image features containing complete forgery clues, thereby accurately judging the authenticity of images and videos in digital content. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 is a flow chart of the present invention;
[0041] Figure 2 It is a schematic structural diagram of the deepfake face detection network in the present invention;
[0042] Figure 3 It is a schematic structural diagram of the dynamic interaction module in the present invention;
[0043] Figure 4 It is a schematic structural diagram of the adaptive region dynamic convolution module in the present invention;
[0044] Figure 5 It is a schematic structural diagram of the adaptive frequency dynamic filtering module in the present invention;
[0045] Figure 6 It is a schematic structural diagram of the bidirectional adapter in the present invention. Detailed implementation manners
[0046] The present invention will be described in detail below with reference to the accompanying drawings. The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, in which the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions from beginning to end. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention. The orientation terms such as left, middle, right, up, and down in the embodiments of the present invention are only relative concepts to each other or are referenced based on the normal use state of the product, and should not be considered as restrictive.
[0047] Data processing: In the data processing stage, for the data during training and testing, including videos from the FaceForensics++ (ff++), Celeb-DF, and DFDC datasets, 32 frames are uniformly sampled and face extraction and alignment are performed using MTCNN, and the frame size is adjusted to 299×299 as the input image; the maximum number of training epochs T and the batch size B are set.
[0048] The deepfake face detection method based on adaptive convolution and bidirectional adapter, as Figure 1 and Figure 2 shown, includes the following steps:
[0049] Step 1: Construct a deepfake face detection network, which includes an adaptive region dynamic convolution module, an adaptive frequency dynamic filtering module, and several dynamic interaction modules;
[0050] Step 2: Input the forged images in the training set into the deepfake face detection network for training, that is, input a batch of image samples (x i , y i ) in the training set into the deepfake face detection network; where x i is the input image, and yi is a label of an image, including the following steps:
[0051] Step 2.1: Dynamically extract the spatial domain forgery features of different input images through the Adaptive Region Dynamic Convolution Module;
[0052] The Adaptive Region Dynamic Convolution Module (ARDConv) includes a convolution module and an attention module. As Figure 4 shown, dynamically extract the spatial domain forgery features of different input images through the Adaptive Region Dynamic Convolution Module; Step 2.1 specifically includes the following steps:
[0053] Step 2.1.1: Generate the guiding feature map Y of the input image through the convolution module h,w,c ;
[0054]
[0055] where h, w, and c are the height, width, and number of channels of the input image respectively; x i is the input image, and W c represents the standard convolution of the c-th channel;
[0056] Step 2.1.2: Process the guiding feature map through the attention module to generate the attention weight map A k ;
[0057] Step 2.1.3: Normalize the attention weight map and make the sum of the weights of all regions at each position equal to 1 to obtain the normalized weight map
[0058]
[0059] where, A k (i,j) represents the attention weight at position (i,j), m represents the number of divided regions, represents the degree of belonging of position (i,j) to the k-th region;
[0060] Step 2.1.4: Divide the image into different regions in the spatial dimension according to the weight map and then perform convolution operations on the feature maps of each region respectively to obtain the feature representations of different regions;
[0061] Step 2.1.5: Use the attention weight map to multiply the feature representations of different regions element by element to obtain the weighted spatial domain feature representation F s ;
[0062]
[0063] where, Y kis the k-th regional feature map, w k is the convolution kernel of the k-th region, Y k *w k represents the operation of region division and convolution, and ⊙ represents element-wise multiplication;
[0064] Compared with other traditional methods for extracting spatial domain features, this method dynamically adjusts feature weights to enhance the features of important regions; especially in the scenario of deepfake face detection, the model can focus more on the key regions related to face forgery, reducing the interference of irrelevant background information; at the same time, it adaptively divides regions to capture local detailed information; in the face image, the forgery traces are relatively rich in areas such as facial feature contours, and the adaptive regions can better extract local features and improve the robustness of the model;
[0065] Step 2.2: Dynamically extract the frequency domain forgery features of different input images through the Adaptive Frequency Dynamic Filtering Module (AFDF), specifically:
[0066] Step 2.2.1: Decompose the input image by learnable latent space downsampling to obtain different frequency components;
[0067] The different frequencies include low-frequency components, medium-frequency components, and high-frequency components;
[0068] f l = Conv↓2(Conv↓2(I)) (20)
[0069] f m = Conv↓2(I)-Conv↓2(Conv↓2(I))↑2 (21)
[0070] f h = Conv(I)-Conv↓2(I)↑2 (22)
[0071] where f l 、f m 、f h represent the low-frequency component, medium-frequency component, and high-frequency component respectively, and I represents the input image; specifically, use a convolutional layer with a stride of 4 to downsample the feature space to obtain the corresponding low-frequency component f l ,then remove the corresponding original feature f l ,and perform downsampling with a stride of 2 to obtain the corresponding medium-frequency component f m ,similarly, to obtain the high-frequency component f h ,remove the downsampled feature with a stride of 2 from the non-downsampled feature; ultimately helping the model to more accurately capture global and local information during multi-level feature extraction, while improving flexibility and generalization ability;
[0072] Step 2.2.2: Generate the weights α of different frequency components through the attention module k , dynamically combine a set of basic filters g k Construct the frequency dynamic filter W, convolve it with the feature representation containing all frequency components to obtain the image frequency domain feature F f ;
[0073]
[0074] F f = W * F (24)
[0075] where g k is the k-th basic filter, α k is the attention weight, W is the constructed dynamic filter, k ∈ {l, m, h}, and F is the unsampled feature of the input image;
[0076] In deepfake face detection, forgery clues are implicitly embedded in the frequency information. However, existing frequency feature extraction methods are difficult to fully capture the subtle artifacts in the frequency domain. This method uses different frequency components and assigns weights according to the importance of different frequency components, which can more comprehensively extract forgery information in the frequency domain and more accurately extract the key frequency components of forged images. By introducing frequency decomposition, multi-frequency hierarchical filtering, dynamic frequency weight adjustment, and adaptive filter design, this method greatly improves the accuracy and robustness of frequency domain feature extraction in image forgery detection. Compared with traditional fixed-frequency filtering methods, the adaptive frequency filter in this paper can more accurately extract the key frequency components of forged images and avoid the under-generalization of traditional methods for different forgery types.
[0077] Step 2.3: Interact the spatial domain forgery feature and the frequency domain forgery feature through a number of dynamic interaction modules in sequence; among them, the previous dynamic interaction module downsamples the feature data after interaction before sending it into the next dynamic interaction module.
[0078] Among them, as Figure 3 shown, the dynamic interaction module includes two bidirectional adapters and two backbone networks. The backbone network is a swintransformer model. Insert one bidirectional adapter between the W-MSA modules of the two swintransformer models, and insert the other bidirectional adapter between the SW-MSA modules of the two swintransformer models. The two backbone networks are used to input the two types of feature data input to the dynamic interaction module respectively. These two types of feature data are the spatial domain forgery feature and the frequency domain forgery feature respectively, or the two feature data output by the previous dynamic interaction module respectively.
[0079] Among them, the bidirectional adapter is used to fuse the inputs of the normalization layers in the front positions of two W-MSA modules or SW-MSA modules, and add the two outputs after fusion by the bidirectional adapter to the respective outputs of these two W-MSA modules or SW-MSA modules; as Figure 3 shown, Fs and Ff are the spatial domain forgery feature and the frequency domain forgery feature respectively. After being processed by the bidirectional adapter, Ts and Tf are the spatial domain features affected by the frequency domain and the frequency domain features affected by the spatial domain respectively; these features not only contain the original spatial and frequency information, but also fuse the correlation between the two types of features, and can better reflect the global and local characteristics of the image; the essence of feature interaction is to enable features from different sources (such as the spatial domain and the frequency domain) to exchange and fuse information in a certain shared representation space; therefore, the bidirectional adapter is designed to make the two types of features affect each other in the low-dimensional space through dimensionality reduction, non-linear transformation and reconstruction.
[0080] The swintransformer model has powerful feature extraction capabilities. With its self-attention mechanism and hierarchical design, it can more effectively capture image context information and fine-grained features; the input image spatial domain features and frequency domain features pass through the swintransformer model and the bidirectional adapter, obtaining features that contain both the information of each branch and the forgery information of the other party, realizing dynamic interaction, significantly improving the model's feature representation ability in complex scenarios, and improving the accuracy and robustness of detection;
[0081] Step 2.4: Concatenate the outputs of the last dynamic interaction module and send them into the fully connected layer for classification to obtain the detection result of the forged image;
[0082] Step 2.5: Update the model parameters of the deep fake face detection network according to the cross-entropy loss and the SGD optimization strategy to obtain the final deep fake face detection network;
[0083] Step 3: Perform deep fake face detection through the final deep fake face detection network.
[0084] Experimental test
[0085] This paper mainly uses the representative public datasets FaceForensics++, Celeb-DF, and DFDC in deepfake detection. The FaceForensics++ dataset contains 1000 real videos and 4000 fake videos, which are generated by 4 face forgery algorithms: Deepfakes (DF), Face2Face (F2F), FaceSwap (FS), and NeuralTextures (NT). At the same time, FF++ provides data with 3 different video compression qualities: the original version (raw), the lightly compressed version (HQ), and the heavily compressed version (LQ). The Celeb-DF dataset contains 590 real videos and 5639 fake videos, covering tasks of different ages, races, and genders. The quality, diversity, and detection difficulty of the fake videos have been significantly improved. The DFDC dataset contains more than 100,000 video clips, and fake videos are generated by various methods.
[0086]
[0087] Select Xception, SPSL+ID3, F3-Net, ATSC, SFIC, and F2Trans-B methods with excellent detection effects in deepfake detection as the baseline methods to verify the advantages of the deepfake detection method proposed in this paper in terms of detection effect and generalization. The comparison results with each baseline method in the dataset are shown in the above table. It can be seen that most of the indicators of this invention are higher than those of the prior art.
[0088] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. A deep fake face detection method based on adaptive convolution and bidirectional adapter, characterized by: The following steps are involved: Step 1: Build a deep fake face detection network, which includes an adaptive region dynamic convolution module, an adaptive frequency dynamic filtering module, and several dynamic interaction modules; Step 2: Input the fake images in the training set into the deep fake face detection network for training, which includes the following steps: Step 2.1: Dynamically extract spatial domain forgery features of different input images through the adaptive region dynamic convolution module; Step 2.2: Dynamically extract frequency domain forgery features of different input images through the adaptive frequency dynamic filtering module; Step 2.3: The spatial domain forged features and the frequency domain forged features are sequentially interacted through a number of dynamic interaction modules; wherein the front dynamic interaction module downsamples the feature data after interaction before sending it to the next dynamic interaction module; Step 2.4: Concatenate the output of the last dynamic interaction module and send it to the fully connected layer for classification to obtain the detection result of the forged image; Step 2.5: Update the model parameters of the deep fake face detection network according to the cross entropy loss and SGD optimization strategy to obtain the final deep fake face detection network; Step 3: Perform deepfake face detection through the final deepfake face detection network.
2. The deep fake face detection method based on adaptive convolution and bidirectional adapter according to claim 1 is characterized in that: The dynamic interaction module includes two bidirectional adapters and two branch networks. The branch network is a swintransformer model. One bidirectional adapter is inserted between the W-MSA modules of the two swintransformer models, and the other bidirectional adapter is inserted between the SW-MSA modules of the two swintransformer models. The two branch networks are used to respectively input two feature data input by the dynamic interaction module, and the two feature data are respectively spatial domain forged features and frequency domain forged features, or respectively two feature data output by the front-end dynamic interaction module; wherein the bidirectional adapter is used to fuse the inputs of the normalization layers of the front-ends of the two W-MSA modules or SW-MSA modules, and the two outputs after the fusion of the bidirectional adapter are respectively added to the respective outputs of the two W-MSA modules or SW-MSA modules.
3. The deep fake face detection method based on adaptive convolution and bidirectional adapter according to claim 1, characterized in that: The adaptive region dynamic convolution module includes a convolution module and an attention module. Step 2.1 specifically includes the following steps: Step 2.1.1: Generate the guided feature map Y of the input image through the convolution module h,w,c ; Where h, w, c are the height, width and number of channels of the input image respectively; x i is the input image, W c represents the standard convolution of the cth channel; Step 2.1.2: Process the guided feature map through the attention module to generate the attention weight map A k ; Step 2.1.3: Normalize the attention weight map and make the sum of all region weights at each position 1 to obtain the normalized weight map Among them, A k (i,j) represents the attention weight of position (i,j), m represents the number of divided regions, Indicates the degree to which the position (i, j) belongs to the kth region; Step 2.1.4: According to the weight map The image is divided into different regions in the spatial dimension, and then the feature map of each region is convolved to obtain the feature representation of different regions; Step 2.1.5: Use the attention weight map Multiply the feature representations of different regions element by element to obtain the weighted spatial domain feature representation F s ; Among them, Y k is the kth region feature map, w k is the convolution kernel of the kth region, Y k *w k It represents the operation of region division and convolution, and ⊙ represents element-by-element multiplication.
4. The deep fake face detection method based on adaptive convolution and bidirectional adapter according to claim 1, characterized in that: Step 2.2 is as follows: Step 2.2.1: Decompose the input image into different frequency components by downsampling in the learnable latent space; Step 2.2.2: Generate weights α for different frequency components through the attention module k , dynamically combining a set of basic filters g k Construct a frequency dynamic filter W and convolve it with the feature representation containing all frequency components to obtain the image frequency domain feature F f ; F f =W*F (5) Among them, g k is the kth basic filter, α k is the attention weight, W is the constructed dynamic filter, k∈{l,m,h}, and F is the un-downsampled feature of the input image.
5. The deep fake face detection method based on adaptive convolution and bidirectional adapter according to claim 4 is characterized in that: Different frequencies include low-frequency components, mid-frequency components, and high-frequency components; f l =Conv↓2(Conv↓2(I)) (6) f m =Conv↓2(I)-Conv↓2(Conv↓2(I))↑2 (7) f h =Conv(I)-Conv↓2(I)↑2 (8) Among them, f l 、f m 、f h They represent low-frequency component, medium-frequency component and high-frequency component respectively, and I represents the input image.