Self-adaptive deep fake face detection method and system based on space-frequency domain graph learning

Through the combination of the dynamic adaptive wavelet module and the normalized residual isomerographic neural network module, the problem of insufficient adaptability and generalization in face forgery detection is solved, and more efficient forgery face detection is achieved.

CN120356254AActive Publication Date: 2025-07-22EAST CHINA JIAOTONG UNIVERSITY

Patent Information

Application Number
CN202510845971.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-07-22
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

The existing deep forgery detection model has problems in the face forgery detection that cannot adapt to data changes, large amount of calculation, insufficient distinction and poor cross-domain generalization.

Method used

The dynamic adaptive wavelet module, the normalized residual isomerographic neural network module and the adaptive feature fusion module based on gating convolution are adopted, combined with the dynamic asymmetric triple loss function, and the adaptive fusion of frequency domain and airspace features is achieved through multi-scale analysis and improved graph structure construction.

Benefits of technology

It significantly improves the detection accuracy and generalization ability of the model, can better adapt to data changes, reduce the amount of calculation, and improve the stability and robustness of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356254A_ABST
    Figure CN120356254A_ABST
Patent Text Reader

Abstract

The invention discloses a space-frequency domain graph learning-based adaptive deep fake face detection method and system, and the method comprises the steps: randomly extracting an image frame from a video, intercepting a face image, adjusting the feature dimension of the face image, and transmitting the face image to a depth adaptive wavelet module and a normalized residual homomorphic composition neural network module; a depth adaptive wavelet module extracts frequency features of the face image; a normalized residual homograph neural network module extracts spatial domain features of the face image; performing weighted fusion on the frequency domain features and the spatial domain features by using a self-adaptive feature fusion module based on gated convolution, realizing class attention guidance by using the gated convolution, dynamically adjusting the channel of a feature map and the weight of a spatial dimension, and finally obtaining fusion features; and performing classification according to the fusion features by using a classifier. According to the method, the extraction capability of forged detail clues is enhanced, and the detection precision and stability of the model are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of face fraud detection, and in particular to an adaptive deepfake face detection method and system based on spatio-frequency domain graph learning. Background Technique

[0002] Deepfake technology is a means of generating or replacing real picture and video content by means of artificial intelligence. While it opens up new ideas for artistic creation, it may also be used to spread false media information, disrupt social order, and even be applied to war, and its potential hazards cannot be underestimated. Deepfake detection technology helps people distinguish the authenticity of digital content and maintain the credibility of digital information by identifying and analyzing the subtle features and patterns in the tampered or forged digital content.

[0003] The existing deepfake detection technologies mainly rely on deep learning models, such as convolutional neural networks (CNNs) and graph neural networks (GNNs). These models learn a large amount of real and forged image, video or audio data, and extract and identify the subtle texture, color, light and shadow changes and other features in the forged content, as well as the information such as the actions and facial expressions of the people in the video to perform forgery detection. The existing face forgery detection models usually adopt a method combining the spatial domain and the frequency domain to detect forged faces, and identify the generation distortion or mixing inconsistency phenomenon through the spatial domain features, and identify the GAN fingerprints and other phenomena through frequency analysis. However, the existing face forgery detection models have the following problems: First, the discrete cosine transform (DCT) or fixed filter bank is used to extract frequency features, which cannot provide signals in both the spatial domain and the frequency domain at the same time and cannot adapt to data changes; second, the computational cost is large when constructing the graph structure; third, the discrimination degree is insufficient and the cross-domain generalization ability is poor. Summary of the Invention

[0004] The present invention aims to provide an adaptive deepfake face detection method and system based on spatio-frequency domain graph learning, and solve the problems existing in the prior art through a dynamic adaptive wavelet module, a normalized residual homogeneous graph neural network module, an adaptive feature fusion module, and a comprehensive loss optimization function, so as to improve the detection accuracy and generalization ability.

[0005] The present invention adopts the following technology, an adaptive deepfake face detection method based on spatio-frequency domain graph learning, and the steps are as follows: Step 1: Randomly extract image frames from the video, intercept the face images, and after adjusting the feature dimensions of the face images, send them into the deep adaptive wavelet module and the normalized residual homogeneous graph neural network module respectively; Step 2: The deep adaptive wavelet module extracts frequency features of face images at different scales; the normalized residual homogeneous graph neural network module uses the approximate nearest neighbor algorithm to construct the graph topology and updates node features using multi-layer normalized residual homogeneous graph convolution to extract the spatial domain features of face images; Step 3: Use the adaptive feature fusion module based on gated convolution to perform weighted fusion on the frequency domain features and spatial domain features, use gated convolution to achieve class attention guidance, dynamically adjust the weights of the channels and spatial dimensions of the feature map, and finally obtain the fused features, and use a classifier to classify according to the fused features.

[0006] Further preferably, perform image enhancement on the collected face images to expand the face image samples, construct a face image dataset, divide the training set and the test set, and train the adaptive deep fake face detection model composed of the deep adaptive wavelet module, the normalized residual homogeneous graph neural network module, the adaptive feature fusion module and the classifier.

[0007] Further preferably, the deep adaptive wavelet module includes several layers of 2D adaptive wavelet transform and a channel attention module. The 2D adaptive wavelet transform includes a horizontal lifting step and two independent vertical lifting steps, generating four wavelet transform sub-band signals: the wavelet transform sub-band signal LL decomposed by low frequency in the horizontal direction and low frequency in the vertical direction, the wavelet transform sub-band signal LH decomposed by low frequency in the horizontal direction and high frequency in the vertical direction, the wavelet transform sub-band signal HL decomposed by high frequency in the horizontal direction and low frequency in the vertical direction, and the wavelet transform sub-band signal HH decomposed by high frequency in the horizontal direction and high frequency in the vertical direction; the input of each layer of 2D adaptive wavelet transform is the wavelet transform sub-band signal LL decomposed by low frequency in the horizontal direction and low frequency in the vertical direction of the previous layer of 2D adaptive wavelet transform; all the final wavelet transform sub-band signals obtained by decomposing each layer of 2D adaptive wavelet transform are concatenated and then sent to the channel attention module for weighted fusion.

[0008] Further preferably, the channel attention module consists of two multi-layer perceptrons and two activation functions, and its formula is expressed as follows: ; In the formula, respectively represent the learnable weights of the two multi-layer perceptrons, is the RELU activation function, is the Sigmoid activation function, is the attention weight, is the feature sequence output by the 2D adaptive wavelet transform, is the fused frequency domain feature sequence output by the channel attention module.

[0009] Further preferably, the target loss function for training the deep adaptive wavelet module is: ; where M is the number of wavelet decomposition levels, is the concatenation of the wavelet transform sub-band signals (HH, HL, LH) of the th layer, is the Huber norm, is the mean of the input signal of the th layer, is the mean of the wavelet transform sub-band signal LL of the low-frequency decomposition in the horizontal and vertical directions output by the th layer.

[0010] Further preferably, the approximate nearest neighbor algorithm is improved as follows: when finding the neighbors of a node, the average pooling method is used to downsample the face image samples, and the number of nodes in the topological graph is reduced from to , where is the downsampling ratio; the nodes after downsampling are the representative nodes of each region. When querying, only the representative point closest to the query node needs to be found among the nodes; the distance metric method used replaces the Euclidean distance with the cosine similarity.

[0011] Further preferably, the normalized residual homogeneous graph convolution is expressed as: ; where is the feature representation of node in the th layer, is the feature representation of node in the th layer, is the multi-layer perceptron of the th layer, is the learnable parameter, is the neighbor set of node , is the feature representation of neighbor node in the th layer; is the learnable weight of the neighbor node; , where is the learnable weight matrix, is the learnable weight vector, is the transpose, is the feature representation of neighbor node in the th layer, is the leaky rectified linear unit.

[0012] Further preferably, , where BN represents batch normalization, represents the weight matrix of the first fully connected layer of the k-th layer, represents the weight matrix of the second fully connected layer of the k-th layer.

[0013] Further preferably, the adaptive feature fusion module based on gated convolution has two inputs: spatial domain feature and frequency domain feature ; First, the features of both inputs go through a 3×3 convolution to convert the feature dimension to , and one branch of the spatial domain feature is added pixel by pixel to the frequency domain feature after going through a 3×3 convolution, so as to obtain the first fusion feature map , goes through gated convolution and an activation function to obtain the adaptive feature ; After another branch of the spatial domain feature goes through gated convolution transformation, it is multiplied pixel by pixel with the adaptive feature to obtain the hybrid feature map with a class attention guiding mechanism . Finally, the hybrid feature map is added pixel by pixel to the spatial domain feature to obtain the fusion feature .

[0014] Further preferably, a dynamic asymmetric triplet loss function is introduced to optimize the adaptive deepfake face detection model. The mathematical expression of the dynamic asymmetric triplet loss function is: ; In the formula, is the triplet loss, is the Gaussian weight of the -th triplet, and are the distances of the negative sample pair and the positive sample pair in the -th triplet respectively, is the boundary value; ; In the formula, is the relative distance difference between the negative sample pair and the positive sample pair , and are the mean and standard deviation of the Gaussian function; ; In the formula, is the initial boundary value, is the change rate is the number of iterations, is the switching threshold, is the linear attenuation rate.

[0015] Further preferably, the total loss function is expressed as follows: ; In the formula, is the binary cross - entropy loss, is the dynamic asymmetric triplet loss, is the wavelet optimization objective loss, is the weight of the dynamic asymmetric triplet loss function, is the weight of the adaptive wavelet constraint function.

[0016] The present invention also provides an adaptive deep - fake face detection system for spatio - frequency domain graph learning, including a memory and a processor, where computer instructions are stored on the processor, and the steps of the adaptive deep - fake face detection method described by the computer instructions.

[0017] The present invention has the following advantages: 1. The present invention uses dynamic adaptive wavelet transform to extract frequency features, has multi - scale analysis ability, and can adaptively adjust the scale and position parameters of the wavelet basis function to adapt to the changes of the input signal, so as to more effectively discover forgery details. The extracted multi - scale sub - band signals are weighted and fused through the channel attention module, further enhancing the ability to extract forgery detail clues and significantly improving the detection accuracy of the model.

[0018] 2. The present invention utilizes the powerful complex relationship representation ability of the graph neural network (GNN) to capture subtle clues in forged faces. Through the improved approximate nearest neighbor algorithm (ANN), the present invention reduces the computational complexity of graph structure construction and improves the graph construction efficiency. At the same time, the introduction of normalized residual homogeneous graph convolution avoids the problem of gradient disappearance, improves the stability of the model, and accelerates the training convergence.

[0019] 3. The present invention introduces a dynamic asymmetric triplet loss function, weights the samples through a Gaussian function, assigns higher weights to difficult - to - learn sample pairs, thereby improving the detection performance of difficult - to - classify samples. In addition, the method of dynamically adjusting the boundary value further enhances the data adaptation ability of the model, prompting the model to train quickly and converge stably. The unilateral domain generalization technology promotes the aggregation of real samples in different domains, effectively eliminates the distribution gap between domains, and significantly improves the domain generalization ability and detection accuracy of the model.

[0020] 4. The present invention adopts an adaptive fusion module based on gated convolution to deeply fuse the frequency-domain features extracted by dynamic adaptive wavelet and the spatial-domain features extracted by graph neural network. Gated convolution learns a dynamic feature selection mechanism for each channel and spatial position to achieve class attention guidance, and can adaptively select task-related features. This fusion method not only enhances the data adaptability of the model, but also improves the generalization and robustness of deepfake face detection. Through the complementarity of frequency-domain and spatial-domain features, the present invention further improves the data representation ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 It is a framework diagram of the adaptive deepfake face detection model of the present invention.

[0022] Figure 2 It is a structural diagram of the adaptive feature fusion module based on gated convolution.

[0023] Figure 3 It is a schematic diagram of dynamic asymmetric triplet optimization. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0024] The present invention will be further described in detail below with reference to the accompanying drawings.

[0025] Refer to Figure 1 , an adaptive deepfake face detection method based on spatio-frequency domain graph learning, the steps are as follows: Step 1: Randomly extract image frames from the video and intercept the face images. After adjusting the feature dimensions of the face images, they are respectively sent into the deep adaptive wavelet module and the normalized residual homogeneous graph neural network module; Step 2: The deep adaptive wavelet module extracts the frequency features of different scales of the face image; the normalized residual homogeneous graph neural network module uses the approximate nearest neighbor algorithm to construct the graph topology structure, and uses multi-layer normalized residual homogeneous graph convolution to update the node features and extract the spatial-domain features of the face image; Step 3: Use the adaptive feature fusion module based on gated convolution to perform weighted fusion on the frequency-domain features and the spatial-domain features, use gated convolution to achieve class attention guidance, dynamically adjust the weights of the channels and spatial dimensions of the feature map, and finally obtain the fusion features, and use the classifier to classify according to the fusion features.

[0026] The computational cost of a single-frame image is relatively low and the detection is fast. Therefore, the present invention selects single-frame images to carry out forged face detection. As Figure 1As shown, the sample video includes data from different data domains (data domain 1 and data domain 2). First, the present invention randomly extracts image frames from the video, uses the MTCNN algorithm to intercept face images, and then independently applies methods such as random rotation, brightness adjustment, and Gaussian noise to each face image for image enhancement. The face image samples after image enhancement pass through 2 groups of convolutional networks (Conv-BN-ReLU) to increase the feature dimension from to , where H is the height, W is the width, and C is the number of channels.

[0027] As Figure 1 shown, the present invention uses a deep adaptive wavelet module to extract the frequency domain features of fraudulent faces. Existing face forgery detection models usually use the discrete cosine transform (DCT) to extract the frequency domain features of forged faces. However, the wavelet transform (WT) can provide local information of the signal in both the spatial and frequency domains, while the discrete cosine transform can only provide frequency domain information and cannot provide local spatial features; moreover, the wavelet transform can provide frequency information at different scales and can better capture the local detail features of the signal, while the discrete cosine transform does not have this multi-scale analysis ability. Therefore, the present invention uses the second-generation wavelet transform to extract the frequency domain features of faces and uses a non-linear function represented by a neural network to replace the updater and predictor to adaptively adjust the scale and position parameters of the wavelet basis function to better adapt to the signal features at different frequencies and spatial positions.

[0028] The second-generation wavelet transform takes the signal as input, and its output is the approximation of the wavelet transform and the detail subband signal 。The second-generation wavelet transform includes three stages: separating the signal, an updater, and a predictor. Drawing on the Deep Adaptive Wavelet Module (DAWN), the present invention uses non-linear functions represented by neural networks to replace the updater and predictor of the second-generation wavelet transform to adapt to the changes in the input signal, thereby obtaining a 2D adaptive wavelet transform. The 2D adaptive wavelet transform includes a horizontal lifting step and two independent vertical lifting steps. These lifting steps are applied sequentially to decompose the signal, generating four wavelet transform sub-band signals: the wavelet transform sub-band signal LL obtained by decomposing the low frequency in the horizontal direction and the low frequency in the vertical direction, the wavelet transform sub-band signal LH obtained by decomposing the low frequency in the horizontal direction and the high frequency in the vertical direction, the wavelet transform sub-band signal HL obtained by decomposing the high frequency in the horizontal direction and the low frequency in the vertical direction, and the wavelet transform sub-band signal HH obtained by decomposing the high frequency in the horizontal direction and the high frequency in the vertical direction. Each lifting step has its own predictor and updater, and the internal structures of the updater and predictor are the same in both the vertical and horizontal directions, including convolutional layers and non-linear activations, etc. The working process of the updater and predictor: First, the Padding module is a replication and extension of the signal, then 1x3 convolution and RELU activation are performed, and then 1x1 convolution and Tanh activation are performed. After updating and prediction, the 2D adaptive wavelet transform also performs spatial pooling, reducing the spatial size of the output by half relative to the input.

[0029] Such a 2D adaptive wavelet decomposition layer can have at most layers, where W is the width of the input image. The input of each layer corresponds to the wavelet transform sub-band signal LL obtained by decomposing the low frequency in the horizontal direction and the low frequency in the vertical direction at the previous level. Each layer outputs four wavelet transform sub-band signals (LL, LH, HL, and HH). Therefore, M decomposition layers generate a total of wavelet transform sub-band signals.

[0030] Finally, after these sub-wavelet transform sub-bands are concatenated, they will be sent to the channel attention module for weighted fusion. The concatenation formula is expressed as follows: ; where represents the wavelet transform sub-band signal obtained by decomposing the low frequency in the horizontal direction and the high frequency in the vertical direction at the th level, represents the wavelet transform sub-band signal obtained by decomposing the high frequency in the horizontal direction and the low frequency in the vertical direction at the th level, represents the wavelet transform sub-band signal obtained by decomposing the high frequency in the horizontal direction and the high frequency in the vertical direction at the th level, i = [1, 2,..., M], represents the wavelet transform sub-band signal obtained by decomposing the low frequency in the horizontal direction and the low frequency in the vertical direction at the th level, X represents the feature sequence output by the 2D adaptive wavelet transform, and M is the number of wavelet decomposition layers.

[0031] In order to extract more useful frequency features, the present invention sends the sub-band signals of each wavelet transform into a channel attention module for weighted fusion. The channel attention module of the present invention consists of two multi-layer perceptrons and two activation functions, and its formula is expressed as follows: ; In the formula, respectively represent the learnable weights of the two multi-layer perceptrons, is the RELU activation function, is the Sigmoid activation function. Through the training of two-layer multi-layer perceptrons, the importance of each wavelet transform sub-band signal can be learned, and then the normalized weight between 0 and 1 can be obtained through the Sigmoid function. Finally, it is weighted by multiplying the channel width to the feature sequence output by the 2D adaptive wavelet transform, and the weighted fusion of features in the channel dimension is completed to obtain the fused frequency domain feature sequence .

[0032] When training the deep adaptive wavelet module, on the one hand, it is required that the input signal of each 2D adaptive wavelet transform has the same average value as its output low-frequency approximation (low-frequency in the horizontal direction and low-frequency in the vertical direction) in order to retain the mean value of the input signal; on the other hand, it is required that the detail coefficients output by each wavelet module are as small as possible. Therefore, the wavelet optimization objective loss function of this deep adaptive wavelet module is: ; In the formula, is the wavelet optimization objective loss, M is the number of wavelet decomposition layers, is the concatenation of the wavelet transform sub-band signals HH, HL, and LH of the th layer (which can be called the detail sub-band signal), is the Huber norm, and the first term in the objective loss function is to minimize the sum of the detail coefficients on all decomposition layers. is the mean value of the input signal of the th layer, is the mean value of the wavelet transform sub-band signal LL of the low-frequency decomposition in the horizontal direction and the vertical direction output by the th layer. The second term in the objective loss function is to minimize the sum of the L2 norms of the difference between the input signal and the output wavelet transform sub-band signal LL.

[0033] As Figure 1 shown, another branch of the present invention is to use a normalized residual homogeneous graph neural network module to extract the spatial domain features of fraudulent faces. This module first passes the face image through a convolutional network (3 groups of Conv-BN-ReLU) to obtain a feature vector , the feature dimension changes from to , where is the number of nodes, is the number of channels, and the feature vector adds learnable positional encoding , , and then is fed into the normalized residual homogeneous graph neural module (NRGIN) to extract spatial domain features. The normalized residual homogeneous graph neural module includes multiple layers of networks. Each layer of the network will dynamically construct a topological graph using an improved approximate nearest neighbor algorithm, update the node features using normalized residual homogeneous graph convolution, and further enrich the node features using a feed-forward neural network (FFN).

[0034] Graph neural networks only accept the input of topological graphs. Therefore, it is necessary to convert the image into a topological graph. The topological graph is defined as , where is the set of each node in the topological graph, is the set of each edge in the topological graph. The topological graph can be regarded as a total set of nodes and edges. During the process of constructing the topological graph, the K-nearest neighbor algorithm (KNN) is usually used to select K nearest neighbor nodes , thereby constructing the edges connecting the nodes.

[0035] When using the KNN algorithm to construct the topological graph, since it is necessary to calculate the distance between all pairs of nodes, there is a relatively high computational complexity. For this reason, the present invention improves the approximate nearest neighbor (ANN) algorithm to quickly find neighbor nodes. ANN sacrifices a certain amount of accuracy in exchange for a faster search speed. There are various implementation schemes for ANN. Among them, the inverted file index (IVF) is an index technology for efficiently retrieving similar vectors. It is a spatial partitioning index that requires dividing the data into multiple regions, and each region is represented by a representative point. During the query, first find the representative point closest to the query point, and then scan these regions.

[0036] The inverted file index needs to apply clustering to divide the index regions, and the computational complexity of the clustering algorithm is already relatively high and does not meet the needs of the present invention. Therefore, the present invention does not use the clustering algorithm to divide the index regions, but uses the principle that adjacent pixels in the image have high correlation, and simply divides the image into blocks, and each block corresponds to an index region. The specific method is that when looking for node neighbors, the average pooling method is used to downsample the face image samples. The number of nodes in the topological graph changes from to , where is the downsampling ratio. The nodes after downsampling are the representative nodes of each region. During the query, only Among the nodes, find the representative point closest to the query node. Therefore, compared with the computational complexity of KNN , the computational complexity of the improved approximate nearest neighbor (ANN) algorithm will be reduced to . In addition, the present invention uses cosine similarity to replace Euclidean distance, further reducing the computational complexity of finding the nearest neighbor.

[0037] The graph neural network updates the features of the current node by using the features of neighbor nodes through graph convolution operations. Among various graph convolution operations, the original homogeneous graph convolution network can effectively capture the topological structure of the graph and achieve a discrimination ability equivalent to that of the Weisfeiler-Lehman (WL) graph isomorphism test. The original homogeneous graph convolution network updates the node features through a multi-layer perceptron and a summation aggregation mechanism. The formula for updating the node features of the -th layer is: In the formula, is the feature representation of node in the -th layer, is the feature representation of node in the -th layer, is the multi-layer perceptron of the -th layer, are learnable parameters used to adjust the weights of self-features and neighbor aggregation features, is the neighbor set of node , is the feature representation of neighbor node in the -th layer. The aggregation method of the traditional graph neural network (GCN) mainly focuses on the information of local neighbor nodes and lacks the capture of the global structure information of the graph. When the local structures of two graphs are similar but the global structures are different, GCN may have difficulty distinguishing these graph structures. The discriminative / representative ability of GIN has been proven to be equivalent to that of the WL test and can distinguish more graphs with different structures.

[0038] However, the number of stacked layers of the original homogeneous graph convolution network is limited. As the network deepens, the gradient gradually disappears or explodes during backpropagation, resulting in difficulty in effectively training the deep network. Therefore, the present invention introduces residual connections and normalization in the homogeneous graph convolution network, and automatically focuses on important neighbor nodes through adaptive attention weights, making it a normalized residual homogeneous graph convolution, and its formula is expressed as: ; In the formula, finally adding is to add the transformed feature to the input feature. This residual connection method can directly retain the input feature and avoid the smoothing of key information. is the learnable weight of neighbor nodes. It utilizes the correlation between neighbor nodes and the central node, assigns different weights to neighbors, enabling the model to focus on more important neighbor nodes. At the same time, normalizes the weights using the Softmax function. , where is the learnable weight matrix. is the learnable weight vector. is the transpose of is the neighbor node at the -th layer's feature representation. is the leaky rectified linear unit function.

[0039] , where BN represents batch normalization. represents the weight matrix of the first fully connected layer at the k-th layer. represents the weight matrix of the second fully connected layer at the k-th layer. Its main function is to activate and normalize the transformed features, further improving the stability of the model. The original GIN did not normalize the node features. The distribution of the input features changes during the training process, resulting in internal covariate shift. Since the input distribution of each layer changes continuously, the network parameters need to be adjusted frequently to adapt to the new distribution, which makes the training process more complex and increases the training difficulty. The normalized residual homogeneous graph convolution stabilizes the distribution of the input to each layer through neighbor weight normalization and batch normalization, accelerates the training convergence, and reduces the sensitivity of the training process to hyperparameters.

[0040] To improve the feature selection ability of the model, the present invention introduces an adaptive feature fusion module based on gated convolution to perform weighted fusion on frequency features and spatial domain features. The module structure is as shown in Figure 2 . Gated convolution enhances the model's ability to represent and flexibility for input data by introducing a gating mechanism. Its core idea is to learn a dynamic feature selection mechanism for each channel and each spatial position, so as to better capture the important features in the input data and suppress the unimportant features. Gated convolution contains two convolutional layers: one for generating feature maps and the other for generating gating signals. Its core formula is: ; where is the input of the gated convolution, Y is the output of the gated convolution. and are the convolutional kernel and the gated convolutional kernel respectively. is the Sigmoid activation function, which is used to ensure that the gating value is between 0 and 1.

[0041] The present invention realizes class attention guidance through gated convolution, that is, dynamically adjusts the channel and spatial dimensions of the feature map through gated convolution to enhance the model's attention to important features. The adaptive feature fusion module based on gated convolution has two inputs: spatial domain features and frequency domain features ; First, the features of both inputs go through 3×3 convolution to convert the feature dimension to , one branch of the spatial domain features goes through a 3×3 convolution and then is added pixel by pixel to the frequency domain features to obtain the first fused feature map , goes through gated convolution and an activation function to obtain the adaptive feature ; Another branch of the spatial domain features is transformed through gated convolution and then multiplied pixel by pixel with the adaptive feature to obtain the hybrid feature map with a class attention guidance mechanism . Finally, the hybrid feature map is added pixel by pixel to the spatial domain features to obtain the fused feature , and the fused feature will be fed into the classifier for end-to-end training.

[0042] In order to improve the generalization ability of the adaptive deepfake face detection model and the discrimination of each category, the present invention introduces a dynamic asymmetric triplet loss function to optimize the adaptive deepfake face detection model. The following will be introduced from two aspects: sample mining strategy and target loss function respectively.

[0043] Referring to Figure 3 , the training samples of the present invention come from different data domains (data domain 1, data domain 2, and data domain 3) and have two categories (real samples and forged samples). In order to improve the data domain generalization ability of the model, it is necessary to minimize the gap between data domains. However, different data domains use different forgery methods, and the distribution of their forged features varies greatly. It is difficult and unnecessary to eliminate the differences in forged features between different data domains; while the real sample feature distributions in different data domains are similar, and the differences in real features between different data domains are relatively easy and necessary to be eliminated. Therefore, the present invention uses single-sided real samples to carry out data domain generalization optimization, and forged samples are only used to promote the discrimination between classes. Since different optimization strategies are adopted for positive and negative samples, this method is called asymmetric triplet optimization. The mining strategy of the dynamic asymmetric triplet is as Figure 3 shown. The specific mining strategy is: (1) When the anchor sample is a real sample, then select real samples from other data domains as positive samples, and select forged samples from the same data domain as negative samples to minimize the gap between real samples in different data domains and promote the separation between classes; (2) When the anchor sample is a forged sample (Figure 3 If the middle border is bold (the anchor sample), then select the forged samples in the same data domain as the positive samples, and select the real samples in the same data domain as the negative samples to promote the compactness within the class and the separation between classes.

[0044] The objective of the triplet loss function is to make the relative difference between the distance between the anchor sample and the negative sample and the distance between the anchor sample and the positive sample as large as possible. On the basis of the triplet loss function, the dynamic asymmetric triplet loss function weights the samples, and assigns higher weights to those difficult-to-learn sample pairs, so that the model pays more attention to these samples during training. The mathematical expression of the dynamic asymmetric triplet loss function is: ; In the formula, is the dynamic asymmetric triplet loss, is the Gaussian weight of the th triplet, and are the distances of the negative sample pair and the positive sample pair in the th triplet respectively, is the boundary value.

[0045] The present invention uses the method of Gaussian function to weight the samples, and the Gaussian weight can be calculated by the following formula: ; In the formula, is the relative distance difference between the negative sample pair and the positive sample pair , , and are the mean and standard deviation of the Gaussian function, which can be dynamically adjusted according to the data. When the distance difference is small, it means that the current triplet is a difficult-to-classify sample, and a larger weight will be assigned.

[0046] In addition, the boundary value in the objective function can control the relative distance interval between the positive and negative sample pairs. The setting of the boundary value is crucial for the performance and training process of the model, and different data sets and tasks may require different boundary values. If the boundary value is too small, it cannot ensure that the model learns a feature representation with sufficient discriminative power; if the boundary value is too large, it means that the model needs to pull the distance between the positive sample and the anchor point closer, and at the same time push the distance between the negative sample and the anchor point farther, which will increase the difficulty of model training and cause the network to be difficult to converge. Therefore, the present invention introduces a dynamic boundary value. At the beginning of training, a larger boundary value is used to prompt the model to learn quickly at the beginning of training, and the boundary value is reduced in the later stage to make more refined adjustments to the model, thereby improving the performance and generalization ability of the model. The calculation formula of this boundary value is: ; In the formula, is the initial boundary value, is the rate of change, which can be set to 0.05 in the experiment. is the number of iterations. As the iteration progresses, the boundary value It decays exponentially with the number of iterations. is the lower limit to prevent the boundary value from being too small. It can be set to 1e-5 in the experiment. In addition, when the number of training steps reaches a certain requirement, the boundary value switches to linear attenuation to prevent the feature distance from converging too much. The calculation formula is: ; In the formula, For the switching threshold, the experiment can be set to 300. is a linear decay rate, and is experimentally set to 0.05. Such a setting can force the model to learn robust feature representations in the early training stage, and allow the model to be fine-tuned in a reasonable feature space in the later stage, so that the model pays more attention to subtle differences in features and improves detection accuracy.

[0047] In order to collaboratively optimize the deep adaptive wavelet module, the normalized residual isomorphic graph neural network module, and the adaptive feature fusion module based on gated convolution, the present invention integrates multi-task loss functions and jointly supervises model training. The total loss function consists of three parts: wavelet optimization target loss function, dynamic asymmetric triple loss function, and binary cross entropy loss function, and its formula is as follows: ; In the formula, is the binary cross entropy loss, It is a dynamic asymmetric triplet loss, whose optimization goal is to eliminate the gap between domains as much as possible while promoting compactness within the class and separation between classes; It is the wavelet optimization target loss, and its optimization goal is to force the input signal and the output signal to have the same mean, while the detail coefficient is as small as possible. and It is a hyperparameter used to balance the impact of each objective loss function on the model. is the weight of the dynamic asymmetric triplet loss function, is the weight of the adaptive wavelet constraint function.

[0048] The present invention also provides an adaptive deep fake face detection system for space-frequency domain graph learning, comprising a memory and a processor, wherein computer instructions are stored on the processor, and the steps of the adaptive deep fake face detection method described in the computer instructions.

[0049] The above description only represents the preferred embodiments of the present invention and does not limit the present invention in other forms. Any person skilled in the relevant art may use the disclosed content above to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention still fall within the protection scope of the technical solution of the present invention.

Claims

1. An adaptive deepfake face detection method based on spatio-temporal frequency domain graph learning, characterized by the steps As follows: Step 1: Randomly extract image frames from the video and intercept face images. After adjusting the feature dimensions of the face images, they are respectively sent into the deep adaptive wavelet module and the normalized residual homogeneous graph neural network module; Step 2: The deep adaptive wavelet module extracts frequency features of different scales of the face image; the normalized residual homogeneous graph neural network module uses the approximate nearest neighbor algorithm to construct the graph topology structure, and uses multi-layer normalized residual homogeneous graph convolution to update the node features and extract the spatial domain features of the face image; Step 3: Use the adaptive feature fusion module based on gated convolution to perform weighted fusion on the frequency domain features and the spatial domain features, use gated convolution to achieve class attention guidance, dynamically adjust the weights of the channels and spatial dimensions of the feature map, and finally obtain the fused features. Use the classifier to classify according to the fused features.

2. The adaptive deepfake face detection method according to claim 1, characterized in that, Image enhancement is performed on the collected face images to expand the face image samples, construct a face image data set, and divide it into a training set and a test set. The adaptive deep fake face detection model composed of the deep adaptive wavelet module, the normalized residual homogeneous graph neural network module, the adaptive feature fusion module and the classifier is trained.

3. The adaptive deepfake face detection method according to claim 1, wherein The deep adaptive wavelet module includes several layers of 2D adaptive wavelet transform and a channel attention module. The 2D adaptive wavelet transform includes a horizontal lifting step and two independent vertical lifting steps, generating four wavelet transform sub-band signals: the wavelet transform sub-band signal LL decomposed by horizontal low-frequency and vertical low-frequency, the wavelet transform sub-band signal LH decomposed by horizontal low-frequency and vertical high-frequency, the wavelet transform sub-band signal HL decomposed by horizontal high-frequency and vertical low-frequency, and the wavelet transform sub-band signal HH decomposed by horizontal high-frequency and vertical high-frequency; the input of each layer of 2D adaptive wavelet transform is the wavelet transform sub-band signal LL decomposed by horizontal low-frequency and vertical low-frequency of the previous layer of 2D adaptive wavelet transform; all the final wavelet transform sub-band signals obtained by the decomposition of each layer of 2D adaptive wavelet transform are concatenated and then sent into the channel attention module for weighted fusion.

4. The adaptive deepfake face detection method according to claim 3, characterized in that, The channel attention module consists of two multi-layer perceptrons and two activation functions, and its formula is expressed as follows: ; wherein, respectively represent the learnable weights of two multi-layer perceptrons, is the RELU activation function, is the Sigmoid activation function, is the attention weight, is the feature sequence output by the 2D adaptive wavelet transform, is the fused frequency domain feature sequence output by the channel attention module.

5. The adaptive deepfake face detection method according to claim 3, characterized in that, The wavelet optimization objective loss function for training the deep adaptive wavelet module is: ; In the formula, is the wavelet optimized objective loss, M is the number of wavelet decomposition layers, is the concatenation of the wavelet transform subband signals HH, HL, and LH of the th layer, is the Huber norm, is the mean of the input signal of the th layer, is the mean of the wavelet transform subband signal LL of the low-frequency decomposition in the horizontal direction and the low-frequency decomposition in the vertical direction output by the th layer.

6. The adaptive deepfake face detection method according to claim 1, characterized in that, The approximate nearest neighbor algorithm is improved as follows: when finding the nearest neighbor of a node, the average pooling method is used to downsample the face image samples, and the number of nodes in the topological graph is reduced from to , where is the downsampling ratio; the nodes after downsampling are the representative nodes of each region. During query, only the representative point closest to the query node needs to be found among the nodes; the cosine similarity is used to replace the Euclidean distance as the distance metric method.

7. The adaptive deepfake face detection method according to claim 1, characterized in that, The normalized residual homogeneous graph convolution is expressed as: ; wherein, is the feature representation of node at the layer, is the feature representation of node at the layer, is the multi-layer perceptron of the layer, are learnable parameters, is the neighbor set of node , is the feature representation of neighbor node at the layer; are the learnable weights of neighbor nodes; , where is a learnable weight matrix, is a learnable weight vector, is the transpose, is the feature representation of neighbor node at the layer, is the leaky rectified linear unit function; , where BN represents batch normalization, represents the weight matrix of the first fully connected layer of the k-th layer, represents the weight matrix of the second fully connected layer of the k-th layer.

8. The adaptive deepfake face detection method according to claim 1, characterized in that The adaptive feature fusion module based on gated convolution has two inputs: spatial domain features and frequency domain features ; First, the features of both inputs pass through a 3×3 convolution to convert the feature dimension to . One branch of the spatial domain features is added pixel by pixel to the frequency domain features after passing through a 3×3 convolution, so as to obtain the first fused feature map , passes through gated convolution and an activation function to obtain the adaptive feature ; Spatial domain features Another branch of , after gated convolutional transformation, is multiplied pixel by pixel with the adaptive features to obtain a hybrid feature map with a class attention guidance mechanism . Finally, the hybrid feature map is added pixel by pixel to the spatial domain features to obtain the fused features .

9. The adaptive deepfake face detection method according to claim 1, characterized in that A dynamic asymmetric triplet loss function is used to optimize the adaptive deep fake face detection model, and the mathematical expression of the dynamic asymmetric triplet loss function is: ; In the formula, is the dynamic asymmetric triple loss, is the Gaussian weight of the -th triple, and are the distances of the negative sample pair and the positive sample pair in the -th triple respectively, is the boundary value; ; In the formula, is the negative sample pair and the positive sample pair is the relative distance difference, and are the mean and standard deviation of the Gaussian function; ; Wherein, is the initial boundary value, is the rate of change, is the number of iterations, is the switching threshold, is the linear attenuation rate; The total loss function is expressed as follows: ; wherein, is the binary cross-entropy loss, is the dynamic asymmetric triplet loss, is the adaptive wavelet constraint function, is the weight of the dynamic asymmetric triplet loss, is the weight of the adaptive wavelet constraint function.

10. An adaptive deepfake face detection system for learning spatial-frequency domain graphs, characterized in that, It includes a memory and a processor, and computer instructions are stored on the processor. The computer instructions are used to execute the steps of the adaptive deep fake face detection method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Forgery detection method based on feature enhancement and spectral analysis

    CN115829909A

  • Deep counterfeit video detection method and device based on wavelet transform and time sequence feature extraction

    CN117496393A

  • Cross-dimensional image fusion adaptive deep counterfeiting discrimination method

    CN119169338A

  • Deep counterfeit compressed face image identification method based on deep learning

    CN119541058A

  • Detecting forged facial images using frequency domain information and local correlation

    US20230081645A1

Cited By

  • Depth forgery detection method and system for priori knowledge guidance and double-domain representation optimization

    CN120748054A

  • False image identification method and system based on multi-modal analysis

    CN120894679A

  • Deep forged video detection method based on multi-dimensional feature collaborative modeling

    CN121121210A

  • A deep fake video detection method based on multi-dimensional feature collaborative modeling

    CN121121210B

  • Deep forgery detection method based on adaptive decomposition and high-frequency modulation

    CN121170909A