A detector training method and a fake face generalization detection device for deep fake face generalization detection

CN122738005APending Publication Date: 2026-09-11JIANGSU COLLEGE OF INFORMATION TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610972478.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-01
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

[0004]尽管上述多特征融合方法在一定程度上提升了伪造人脸检测的泛化性能,但现有技术方案存在以下不足:仅简单拼接或并行处理多种特征,并未有效挖掘不同特征之间的内在关联与互补信息;同时,忽略了对特征通道间相关性的深度建模,导致模型在面对未知伪造算法时特征表达能力受限,泛化检测准确率仍有较大提升空间,因此,现有深度伪造人脸检测方法在面对未知伪造算法及跨数据集场景时,难以同时有效捕获伪造人脸图像中的局部细微伪造痕迹和全局语义不一致特征,并且对特征通道间相关性的建模能力不足,导致检测器的特征表达能力受限,跨数据集泛化性能下降,库外检测准确率较低

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122738005A_ABST
    Figure CN122738005A_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on multi-feature fusion's deep fake face generalization detection method and fake face generalization detection device. It includes the following steps: (1) the network structure of constructing multi-feature fusion block;(2) the network structure of constructing lightweight channel attention network block;(3) the structure of constructing detector network based on multi-feature fusion block and lightweight channel attention network block;(4) set weighted binary cross-entropy loss function;(5) set hyperparameter in training process;(6) training detector network;Advantages are: it can effectively improve the feature expression capability of fake trace in deep fake face image, thereby improve the generalization detection performance of detector to unknown fake algorithm and cross dataset scene, reduce the performance attenuation under cross dataset detection scene, improve the generalization recognition ability of detector to unknown fake algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a deep face spoofing detection method and a face spoofing detection device based on multi-feature fusion, belonging to the field of digital media forensics technology. Background Technology

[0002] With the rapid development of deep learning technology, deepfake technology has become increasingly sophisticated, posing serious challenges to social security, cyberspace security, and information authenticity. Existing face detection methods have achieved high detection accuracy on multiple training datasets; however, when faced with unknown forgery algorithms and cross-dataset scenarios, the detection performance of these methods deteriorates significantly, and their generalization ability still has considerable room for improvement.

[0003] Existing methods for detecting tampered faces can be broadly categorized into two types: single-feature detection and multi-feature fusion detection. Single-feature detection methods typically rely on single-modal information such as spatial, frequency, or texture features for discrimination. Multi-feature fusion detection methods, on the other hand, comprehensively utilize various heterogeneous features to improve the model's overall adaptability to different tampering methods. Early work, such as [“Zhou P, Han X, Morariu VI, et al. Two-stream neuralnetworks for tampered face detection[C] / / Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW).2017: 1831-1839.”], proposed a detection framework based on two-stream networks. By processing RGB spatial domain images and frequency domain representations in parallel, it initially achieved cross-modal feature fusion, effectively capturing the inconsistency between visual content and spectral anomalies in generated images. Subsequently, [“Chen S, Yao T, Chen Y, et al. Local relation learning forface forgery detection[C] / / Proceedings of the AAAI Conference on Artificial Intelligence. 2021: 1081-1088.”] further deepened this idea by explicitly modeling the complementary relationship between RGB features and frequency domain features, and enhancing the response of the forged region through attention mechanisms or feature interaction modules, making multi-feature fusion more discriminative. Faced with increasingly diverse forgery algorithms, [“Yu P, Fei J, Xia Z, et al. Improving generalization by commonality learning in face forgery detection[J]. IEEE Transactions on Information Forensics and Security, 2022, 17: 547-558.”] proposed, from the feature semantic level, to jointly model common forgery features (reflecting general tampering traces) and specific forgery features (characterizing artifacts of specific generative models), constructing a hierarchical multi-feature fusion strategy that exhibits stronger robustness and generalization ability in unknown forgery scenarios.

[0004] Although the aforementioned multi-feature fusion methods have improved the generalization performance of fake face detection to some extent, existing technical solutions have the following shortcomings: they simply concatenate or process multiple features in parallel without effectively mining the intrinsic correlation and complementary information between different features; at the same time, they neglect deep modeling of the correlation between feature channels, resulting in limited feature expression ability of the model when facing unknown fake algorithms, and there is still considerable room for improvement in generalization detection accuracy. Therefore, existing deep fake face detection methods are difficult to simultaneously and effectively capture local subtle fake traces and global semantic inconsistencies in fake face images when facing unknown fake algorithms and cross-dataset scenarios. Furthermore, their insufficient ability to model the correlation between feature channels limits the feature expression ability of the detector, reduces cross-dataset generalization performance, and results in low out-of-database detection accuracy. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a detector training method and a fake face generalization detection device that can effectively capture local subtle forgery traces and global semantic inconsistency features in fake face images, and has strong feature expression ability, high cross-dataset generalization performance, and high out-of-database detection accuracy for deep fake face generalization detection.

[0006] To address the aforementioned technical problems, the present invention provides a detector training method for generalized detection of deepfake faces, comprising the following steps:

[0007] (1) Construct a network structure for a multi-feature fusion (MFF) block;

[0008] The multi-feature fusion (MFF) block includes a local branch, a global branch, and an α-Gate fusion unit. The local branch extracts local forgery trace features from the face image using a Local Feature Enhancement Network (LFAN) block. The global branch extracts global semantic inconsistency features from the face image using an MLP-Mixer structure. The α-Gate fusion unit generates adaptive fusion weights based on the local forgery trace features and the global semantic inconsistency features, and performs weighted fusion of the local forgery trace features and the global semantic inconsistency features to obtain the fused features.

[0009] (2) Construct the network structure of the Lightweight Channel Attention Network (LCAN) block;

[0010] The lightweight channel attention network LCAN block is based on an inverted residual structure and introduces ECA channel attention to enhance the correlation modeling capability between feature channels.

[0011] (3) Construct the detector network structure based on multi-feature fusion blocks and lightweight channel attention network blocks;

[0012] The detector network includes a backbone network formed by stacking LCAN blocks and MFF blocks, and a classification head connected to the backbone network. The classification head is used to generate a single prediction value based on the features output by the backbone network, and obtain the probability that the face image to be detected is a fake face image through the Sigmoid function.

[0013] (4) Set the weighted binary cross-entropy loss function;

[0014] (5) Set hyperparameters during training, including hyperparameters of network structure, hyperparameters of network training, and class weights between real and fake classes;

[0015] (6) Train the detector network;

[0016] This step involves training the detector network using training face images with real labels, and updating the network weights of the detector network according to the weighted binary cross-entropy loss function to obtain the trained detector network.

[0017] Further, after step (6), the detector is tested for performance to verify whether it can effectively identify whether the face image to be detected is a real face image or a fake face image in multiple datasets in both in-database and out-of-database scenarios. During the verification process, the face image to be detected is read, the face image to be detected is input into the trained detector network, and the detection result of whether the face image to be detected is a real face image or a fake face image is output.

[0018] Furthermore, in step (1), the local branch extracts local features through a Local Feature Augmentation Network (LFAN) block. The LFAN block consists of a channel adjustment layer, a depthwise separable convolutional layer, a coordinate convolutional layer, a dynamic convolutional layer, and a convolutional feedforward network layer. The dynamic convolutional layer can dynamically adjust the kernel weights according to the input features, and more flexibly focus on subtle and inconsistent local features in the forgery traces. The dynamic convolutional layer is implemented using conditional convolution, as follows:

[0019]

[0020]

[0021]

[0022]

[0023] Where X represents the input feature map of the dynamic convolutional layer, with dimension 1. GAP represents average pooling, and K represents the predetermined number of expert convolutional kernels, set to 8. This represents the k-th routing function. The normalized attention weights of the k-th expert convolutional kernel are calculated using the softmax function. This represents the k-th learnable convolutional kernel; Y represents the dynamic convolution kernel obtained by weighted fusion of all expert convolution kernels; Y represents the output feature with dimension 1. ;

[0024] Global feature processing uses L stacked Mixer layers, each Mixer layer including Token-mixing MLP and Channel-mixing MLP.

[0025] Token-mixing MLP: Performs information exchange in the spatial dimension, applying a multilayer perceptron independently to each channel; Channel-mixing MLP: Performs feature transformation in the channel dimension, applying an MLP independently to each spatial location, with the following formula:

[0026]

[0027]

[0028]

[0029] in, and All are two-layer fully connected networks with a GELU activation function in between. Z represents the input feature with dimension N×C, U represents the intermediate feature after layer normalization of the input Z, V represents the output feature after token-mixing operation and residual connection, and Y represents the final output feature of the MLP-Mixer layer.

[0030] Furthermore, the fusion of global and local features employs an α-Gate mechanism: weights α are generated by superimposing an adaptive average pooling layer, a 1×1 convolutional layer one, a ReLU activation function, a 1×1 convolutional layer two, and a Sigmoid activation function. These weights are then used to fuse local and global features. Finally, a residual connection is used to add the original input features to the fused features, resulting in an enhanced feature representation that is more complementary to both global and local features. The specific formula is as follows:

[0031]

[0032]

[0033] in, and For learnable matrix weights, For the Sigmoid function, Indicates average pooling. Indicates local features, Represents global features. and These represent the fused features and the input features, respectively.

[0034] Drawing inspiration from residual networks, the fused features are then fused again with the input features through element-wise addition, resulting in an enhanced feature representation that is more globally and locally complementary. The output feature F is then obtained. out for:

[0035]

[0036] Furthermore, in step (2), the lightweight channel attention network block is based on an inverted residual structure. Channel attention is introduced to enhance cross-channel interaction. First, a channel adjustment layer is used, followed by channel expansion by convolutional layers, filtering by deep convolutional layers, ECA (Efficient Channel Attention) channel attention, and channel compression by convolutional layers. At the same time, linear activation is used to reduce information loss. Finally, a skip connection is made with the input to obtain the output of the lightweight channel attention network block. That is to say, since the deep separable convolutional layer performs independent filtering operations on each channel of the input feature map, it is impossible to model the dependency relationship between channels, making it difficult to effectively enhance the features of key channels and suppress the feature interference of noisy channels. In order to enhance the cross-channel feature interaction capability of the module, a channel attention LCAN block is designed. By introducing channel attention, the LCAN block improves the representation capability of channel features with only a small number of parameters added to the one-dimensional convolutional layer, making up for the channel isolation defect of deep convolution, and enabling the model to more accurately focus on the fake region of the face for generalization detection.

[0037] Furthermore, in step (3), the method for constructing the detector network structure based on the MFF block and LCAN block is as follows:

[0038] (3-1) Backbone network: The backbone network adopts an architecture of alternating LCAN blocks and MFF blocks. First, a standard convolutional layer is used to obtain shallow features of the image to be detected, and the spatial dimension of the features is reduced to half of that of the image to be detected. Then, multiple LCAN blocks are stacked to process the shallow features to obtain intermediate features. Furthermore, the intermediate features are processed by stacking MFF blocks and LCAN blocks.

[0039] (3-2) Classification Head: The classification head adopts an architecture of convolutional layer, adaptive average pooling layer, dropout layer and linear layer; firstly, the input feature channels are adjusted by convolutional layer, then the feature map is compressed by adaptive average pooling layer, then the dropout layer is applied to prevent overfitting, and finally the spoofing probability is output by linear layer.

[0040] Furthermore, in step (4), the weighted binary cross-entropy function is used to train the detector's detection capability, and the calculation formula is as follows:

[0041]

[0042] Where B is the batch size, y j Represents the true label of the j-th image, when y j =1 indicates a forged face, y j =0 indicates a real face, p j The output of the detector is the forgery probability, where w0=2.9 and w1=1 are the class weights of the real class and the forgery class, respectively.

[0043] Furthermore, in step (5), the hyperparameters of the network training settings are set as follows:

[0044] (5-1) Setting up training and testing images: Scale or crop the training and testing images to 255×255 pixels;

[0045] (5-2) Hyperparameter settings for network training: The batch size used for training is 64, the training is 100 epochs, gradient accumulation and mixed precision training are used, the learning rate is 0.001, and linear warm-up and cosine are used to schedule the learning rate. The optimizer used for training is Adam.

[0046] (5-3) The weight ratio of the fake class to the real class is 1:2.9;

[0047] (5-4) Setting up training and testing data: The detector network is trained using the first training dataset; the trained detector network is validated within the library using the first test dataset which is from the same source as the first training dataset; the trained detector network is validated outside the library using the second and third test datasets which are from different sources than the first training dataset.

[0048] Furthermore, in step (6), the training steps of the detector network are as follows:

[0049] (6-1) Load training data. In order to enhance the generalization effect of the network, data augmentation is performed on the training data. Data augmentation performs random transformation operations on each batch of loaded training data, including scaling, horizontal flipping, color jittering, Gaussian blur, Gaussian noise, JPEG compression, random occlusion, and affine transformation.

[0050] (6-2) Input the loaded training data into the detector network to obtain the classification result, and use the loss function to calculate the loss for the classification result and the true label of the training data;

[0051] (6-3) Update the network weights based on the loss using the Adam optimizer;

[0052] (6-4) Repeat the above steps, iterating through each batch, until the total number of training rounds reaches the specified value.

[0053] Furthermore, the trained detector network is subjected to performance testing to verify whether it can perform generalized recognition of deepfake faces on multiple datasets. The specific steps are as follows: read the three test set images in step (5) in sequence, input the test sets into the detector network in batches, obtain the detection results, and complete the test.

[0054] Furthermore, the present invention also provides a deepfake face generalization detection device, including a network structure construction module, a loss function setting module, a hyperparameter setting module, a training module, and a detection and testing module;

[0055] The network structure construction module is used to construct the network structure of the multi-feature fusion (MFF) block and the network structure of the lightweight channel attention network (LCAN) block, and to construct the structure of the detector network based on the multi-feature fusion (MFF) block and the lightweight channel attention network (LCAN) block.

[0056] The loss function setting module is used to set the weighted binary cross-entropy loss function for training the detector network;

[0057] The hyperparameter setting module is used to set hyperparameters during the training process. The hyperparameters include hyperparameters of the network structure, hyperparameters of network training, and class weights between the real class and the fake class.

[0058] The training module is used to train the detector network using training data;

[0059] The detection testing module is used to perform performance testing on the trained detector network, read the images to be detected from multiple datasets, input the images to be detected into the trained detector network in sequence, and output the detection result of whether the image to be detected is a real face or a fake face.

[0060] Wherein, the multi-feature fusion (MFF) block, the lightweight channel attention network (LCAN) block, the detector network, the weighted binary cross-entropy loss function, the hyperparameters, the training module, and the detection and testing module are respectively used to perform the corresponding steps in the detector training method for generalized detection of deep fake faces as described in any one of claims 1 to 9.

[0061] The advantages of this invention are:

[0062] By extracting local forgery trace features through local branches and extracting global semantic features through MLP-Mixer-based global branches, and using the α-Gate mechanism to achieve adaptive fusion of local and global features, and by enhancing the correlation modeling ability between feature channels through lightweight channel attention network blocks, the feature representation ability of forgery traces in deepfake face images can be effectively improved. This improves the detector's generalization detection performance for unknown forgery algorithms and cross-dataset scenarios, reduces performance degradation in cross-dataset detection scenarios, and enhances the detector's generalization recognition ability for unknown forgery algorithms. Attached Figure Description

[0063] Figure 1 This is a flowchart of the deep fake face generalization detection method based on multi-feature fusion of the present invention.

[0064] Figure 2 This is a structural diagram of the MFF block in this invention;

[0065] Figure 3 This is a structural diagram of the LFAN block in this invention;

[0066] Figure 4 This is a schematic diagram illustrating the principle of global feature extraction in this invention.

[0067] Figure 5 This is a structural diagram of the LCAN block in this invention;

[0068] Figure 6 This is an algorithm diagram of the deep fake face generalization detection method based on multi-feature fusion of the present invention. Detailed Implementation

[0069] The detector training method and fake face generalization detection device of the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0070] Example 1:

[0071] As shown in the figure, the deepfake face generalization detection method of the present invention mainly involves inputting the face image to be detected into a multi-feature fusion network structure based on channel attention for feature extraction and classification, obtaining a binary classification detection result of real or fake, thereby achieving generalized detection of deepfake faces. The multi-feature fusion deepfake face detection framework based on channel attention includes a multi-feature fusion (MFF) block and a lightweight channel attention network (LCAN) block. The proposed method includes: setting the overall network framework; structural design of the MFF block and LCAN block; loading the pre-trained backbone network; setting the loss function; setting the hyperparameters during training; training the network; and performing performance testing on the model, specifically including the following steps:

[0072] Step 1: Read the FaceForensics++ (FF++) dataset, including the training set and the validation set;

[0073] Step 2: Setting up the multi-feature fusion block structure, its network structure is as follows: Figure 2 As shown;

[0074] Step 3: Lightweight channel attention network block structure setup, its network structure is as follows: Figure 5 As shown;

[0075] Step 4: Configure the overall detector network structure based on the MFF block and LCAN block, as follows: Figure 6 As shown;

[0076] Step 5: Setting the weighted binary cross-entropy loss function for the detector network;

[0077] Step 6: Set the hyperparameters during the training process, including the hyperparameters of the network structure, the hyperparameters of network training, and the class weights between the real class and the fake class;

[0078] Step 7: Training the detector network, i.e., the training process of the overall network;

[0079] Step 8: Perform performance tests on the detector network both inside and outside the database to verify whether it can effectively identify whether the face image to be detected is a real face image or a fake face image in multiple datasets both inside and outside the database. Read the images to be detected from multiple datasets, input the images to be detected into the detector in sequence, and the detector outputs the detection results to verify its effectiveness and generalization.

[0080] Further explanation of the multi-feature fusion block (MFF) described in step 2:

[0081] like Figure 2 As shown, the MFF block as a whole adopts a fusion architecture of local and global features: where, for example... Figure 3As shown, a local feature enhancement network is used to extract local features. This network consists of a channel adjustment layer (using 1×1 convolution if the input and output channels are different), depthwise separable convolution, coordinate convolution, dynamic convolution, and a convolutional feedforward network. Among them, the dynamic convolution layer can dynamically adjust the convolution kernel weights according to the input features, and more flexibly focus on subtle and inconsistent local features in the forgery traces. The dynamic convolution layer is implemented using conditional convolution, and the specific implementation is as follows:

[0082]

[0083]

[0084]

[0085]

[0086] Where X represents the input feature map of the dynamic convolutional layer, with dimension 1. GAP represents average pooling, and K represents the predetermined number of expert convolutional kernels, set to 8. This represents the k-th routing function. This represents the normalized attention weight of the k-th expert convolutional kernel, calculated using the softmax function. This represents the k-th learnable convolutional kernel. Y represents the dynamic convolution kernel obtained by weighted fusion of all expert convolution kernels; Y represents the output feature with dimension 1. .

[0087] like Figure 4 As shown, global features are extracted using an MLP-Mixer structure to discover global semantic inconsistencies between deepfake face images and real face images. Global feature processing uses L stacked Mixer layers, each Mixer layer including a Token-mixing MLP and a Channel-mixing MLP.

[0088] Token-mixing MLP: Information exchange occurs in the spatial dimension (i.e., the token sequence dimension), with a multilayer perceptron applied independently to each channel; Channel-mixing MLP: Feature transformation occurs in the channel dimension, with an MLP applied independently to each spatial location. The formulas are as follows:

[0089]

[0090]

[0091]

[0092] in, and All are two-layer fully connected networks with a GELU activation function in between. Z represents the input feature, with dimensions N×C. U represents the intermediate feature after layer normalization of the input Z; V represents the output feature after token-mixing and residual connections; Y represents the final output feature of the MLP-Mixer layer.

[0093] To better utilize the two features, an adaptive channel weight α is set during fusion to dynamically optimize the weight relationship between them. Local features are multiplied by α, and global features are multiplied by 1-α, then the two are summed. The weight α is generated by stacking an adaptive average pooling layer, a 1×1 convolutional layer, a ReLU activation function, another 1×1 convolutional layer, and a Sigmoid activation function. In other words, global and local features are dynamically fused through an α-Gate mechanism. This mechanism uses adaptive average pooling, 1×1 convolution, ReLU, 1×1 convolution, and Sigmoid activation to generate the weight α, which is then used to weight and fuse local and global features. Finally, a residual connection is used to add the original input features to the fused features, resulting in an enhanced feature representation that complements both global and local features. The specific formula is as follows:

[0094]

[0095]

[0096] in, and For learnable matrix weights, For the Sigmoid function, Indicates average pooling. Indicates local features, Represents global features. and These represent the fused features and the input features, respectively.

[0097] Drawing inspiration from residual networks, the fused features are then fused again with the input features through element-wise addition, resulting in an enhanced feature representation that is more globally and locally complementary. The output feature F is then obtained. out for:

[0098]

[0099] Further explanation of the lightweight channel attention network block in step 3:

[0100] Figure 5As shown, the LCAN block is based on an inverse residual architecture and incorporates an ECA channel attention mechanism to enhance the interaction of cross-channel features. First, a channel adjustment layer (using 1×1 Conv stacked with batch normalization if the input and output channels are different) is used. Then, convolutional layers expand the channels, deep convolutional layers perform filtering, ECA efficiently addresses channel attention, and convolutional layers compress the channels. Linear activation is used to reduce information loss. Finally, a skip connection is established with the input to obtain the LCAN block's output. Specifically, its internal structure is as follows: the input undergoes 1×1 convolution and batch normalization to achieve channel alignment; expanded convolution increases the channel dimension, and deep convolution performs spatial filtering; the ECA module achieves adaptive cross-channel weight calibration through local 1D convolution; compressed convolution restores the number of channels and uses a linear activation function to mitigate information decay; finally, residual connections fuse the input and output features to form an efficient feature representation.

[0101] Further explanation is provided regarding the overall detector network structure configuration built based on the MFF block and LCAN block in step 4:

[0102] like Figure 6 As shown, the detector network as a whole adopts a hybrid architecture of alternating LCAN and MFF blocks:

[0103] In the initial stage, shallow features of the image to be detected are obtained by standard 3×3 convolution, and the spatial dimension of these features is reduced to half that of the image to be detected. Then, four LCAN blocks are processed layer by layer (the shallow features are then processed by stacking four LCAN blocks to obtain intermediate features), and each block outputs 16 channels. Deep processing uses MFF and LCAN blocks to stack alternately to process the intermediate features, in the following order: MFF block (256 channels, MLP expansion factor of 4, MLP-Mixer layer of 2), LCAN block (64 channels), MFF block, LCAN block (64 channels), MFF block, to complete encoding and downsampling. Finally, 512 channels of deep features are output by 3×3 convolution.

[0104] The classification head adjusts the features to 1280 channels through convolution, compresses them using adaptive average pooling, and then maps them to two output classes through a linear layer. Specifically, the classification head uses an architecture of convolutional layers, adaptive average pooling layers, Dropout layers, and linear layers. First, the input feature channels are adjusted to 1280 through convolutional layers. Then, the feature map is compressed to a 1×1 size through adaptive average pooling layers. Next, a Dropout layer is applied to prevent overfitting. Finally, a linear layer outputs the spoofing probability.

[0105] Further explanation is provided regarding the loss function settings in step 5. During training, a weighted binary cross-entropy loss function is used to supervise the detector's detection capability. The calculation formula is as follows:

[0106]

[0107] Where B is the batch size, y j Represents the true label of the j-th image, when y j =1 indicates a forged face, y j =0 indicates a real face, p j The output of the detector is the forgery probability, where w0=2.9 and w1=1 are the class weights of the real class and the forgery class, respectively.

[0108] Further explanation of the hyperparameter settings in step 6 is provided. Training and testing images are uniformly scaled to 255×255 pixels; the batch size is set to 64, the total epochs are 100, gradient accumulation and mixed precision strategies are used to accelerate convergence, the optimizer is Adam, the initial learning rate is 0.001, supplemented by linear warm-up and Cosine annealing scheduling; the weight ratio of the fake class to the real class is 1:2.9 to alleviate class imbalance.

[0109] In this embodiment, the first training dataset is preferably the training set of the FF++ dataset, the first test dataset is preferably the test set of the FF++ dataset, and the test dataset is preferably the Celeb-DF dataset and the DFDC dataset, wherein 20,000 images are randomly selected from each of the Celeb-DF dataset and the DFDC dataset for external validation.

[0110] Further explanation of the training process in step 7: The algorithm flowchart of the method proposed in this invention is as follows. Figure 1 As shown. First, the training data is loaded and augmented to improve the network's generalization ability. Specifically, for each batch of training data, operations such as scaling, horizontal flipping, color dithering, Gaussian blurring, Gaussian noise, JPEG compression, random occlusion, and affine transformation are randomly performed. Then, the augmented training data is input into the detector network to obtain the classification result. The loss value between the result and the true label is calculated using the loss function described in step 5. Subsequently, the Adam optimizer is used to update the network weights based on the loss value to optimize the detector network. The above process is repeated for all batches until the total number of training iterations reaches a preset threshold, or the verification result meets expectations.

[0111] Furthermore, the performance test in step 8 will be explained in more detail. The trained detector network was subjected to performance testing to verify its ability to generalize deepfake face recognition across multiple datasets. The performance test verified two aspects: first, its ability to effectively identify both real and deepfake images; and second, its generalization ability on unknown data and unknown forgery methods. Accuracy (ACC) was used as the evaluation metric. The specific steps were as follows: three test set images were read sequentially, and the test sets were input into the detector network in batches to obtain the detection results, thus completing the test. The comparison methods were the classic algorithm B4ATT [“N. Bonettini, ED Cannas, S. Mandelli, L. Bondi, P. Bestagini, and S. Tubaro, “Video face manipulation detection through ensemble of cnns,”in 2020 25th international conference on pattern recognition (ICPR). IEEE, 2021,pp. 5012–5019.”] and MCX-API [“Y. Xu, K. Raja, L. Verdoliva, and M. Pedersen, “Learning pairwise interaction for generalizable deepfake”]. detection,” inProceedings of the IEEE / CVF Winter Conference on Applications of ComputerVision (WACV) Workshops, January 2023, pp. 672–682.”]

[0112] Firstly, this embodiment tests using three datasets: internal and external. Specifically, for the internal dataset, 22,354 images from the test set of the FF++ dataset are used as test data. For the external dataset, 20,000 images are randomly selected from both the DFDC and Celeb-DF datasets. The testing process is as follows: the datasets are loaded, and the trained detector network is input in batches to obtain classification results. The classification results are then compared with the true labels to obtain the ACC index, as shown in Table 1. The results show that the ACC index of the method presented in this invention is superior to the comparative scheme, achieving optimal results on all three datasets, demonstrating the recognition performance and generalization ability of the method presented in this invention.

[0113] Table 1

[0114]

[0115] This method centers on multi-feature fusion, forming complementary feature representations at both global and local scales. It enhances spatial perception of local artifacts in forged faces by constructing a local feature enhancement network, extracts global semantic features using an MLP-Mixer-based global branch, and employs an α-Gate mechanism for mutual fusion. Simultaneously, a lightweight channel attention network block is designed to improve cross-channel feature interaction. Verification shows that the average ACC of this invention outperforms existing methods on the FF++, DFDC, and Celeb-DF datasets, demonstrating that the proposed method can effectively and with high generalization identify unknown deepfake face images.

[0116] To further illustrate the fusion effect of the MFF block proposed in this invention and the role of ECA in the LCAN block within the overall detection model, an ablation experiment was conducted in this embodiment. The network was evaluated by removing the global branch and local branch in the MFF block and the ECA in the LCAN block, respectively, and the experimental results are shown in Table 2.

[0117] Table 2

[0118]

[0119] As shown in Table 2, removing the global branch, local branch, or ECA resulted in varying degrees of decrease in the model's ACC on the FF++, DFDC, and Celeb-DF datasets, indicating that all three modules effectively improved detection performance. The most significant decrease in average ACC was observed without the local branch, suggesting that subtle local forgery traces have a strong discriminative effect on deep forgery detection, especially in cross-dataset scenarios where the lack of local texture, edge, and artifact information weakens the model's ability to identify unknown forgery samples. Without the global branch, the model can still rely on local features to achieve a certain degree of discrimination, but the decrease is more pronounced on external datasets such as DFDC and Celeb-DF, indicating that global semantic inconsistency features help improve the model's generalization ability to data distribution variations and unknown forgery methods. In contrast, removing ECA resulted in a relatively smaller performance decrease, but it still reduced inter-channel feature interaction and key channel response capabilities, leading to a decrease in ACC on external datasets. Overall, the local branch, global branch, and ECA complement each other: the local branch enhances fine-grained forgery detection, the global branch models overall semantic inconsistencies, and ECA strengthens channel attention representation. Together, they improve the model's in-database detection performance and cross-dataset generalization ability.

[0120] Example 2:

[0121] This embodiment provides a deep fake face generalization detection device with multi-feature fusion, including a network structure construction module, a loss function setting module, a hyperparameter setting module, a training module, and a detection and testing module.

[0122] The network structure construction module is used to construct the network structure of the Multi-Feature Fusion (MFF) block and the Lightweight Channel Attention Network (LCAN) block, and to construct the detector network structure based on the MFF block and the LCAN block. The MFF block extracts local forgery trace features through local branches and extracts global semantic inconsistency features through global branches, and adaptively weights and fuses the local forgery trace features and global semantic inconsistency features using an α-Gate mechanism. The LCAN block enhances the correlation modeling capability between feature channels through channel attention.

[0123] The loss function setting module is used to set the weighted binary cross-entropy loss function for training the detector network.

[0124] The hyperparameter setting module is used to set hyperparameters during the training process. The hyperparameters include hyperparameters of the network structure, hyperparameters of network training, and class weights between the real class and the fake class.

[0125] The training module is used to train the detector network using training data with real labels, and update the network weights of the detector network according to the weighted binary cross-entropy loss function to obtain the trained detector network.

[0126] The detection testing module is used to perform performance testing on the trained detector network, read the face image to be detected, input the face image to be detected into the trained detector network, and output the detection result of whether the face image to be detected is a real face image or a fake face image.

[0127] The network structure construction module, loss function setting module, hyperparameter setting module, training module, and detection and testing module are respectively used to execute the corresponding steps in the deep fake face generalization detection method described in Example 1.

[0128] Example 3:

[0129] This embodiment provides a system including a processor and a memory coupled to the processor; the memory stores computer-executable instructions; the processor is configured to execute the steps of the method described in Embodiment 1 when running the computer-executable instructions, which enhances the local feature extraction capability of the network by introducing coordinate convolution and extracts global features using an MLP-Mixer-based structure; on the other hand, it uses residual connections to deeply fuse the global and local fused features with the input features. Furthermore, a channel attention network block is designed to improve the cross-channel interaction capability of features. Experiments show that the model trained on the FF++ dataset achieves ACC of 80.29% and 85.24% on DFDC and Celeb-DF, respectively, achieving effective generalized detection.

[0130] Example 4:

[0131] This embodiment provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, is used to implement all or part of the steps of the method as described in Embodiment 1.

[0132] Of course, the above description is not intended to limit the present invention, and the present invention is not limited to the examples given above. Any changes, modifications, additions or substitutions made by those skilled in the art within the scope of the present invention should also fall within the protection scope of the present invention.

Claims

1. A detector training method for generalized detection of deepfake faces, characterized in that, Includes the following steps: (1) Construct a network structure with multiple feature fusion blocks; The multi-feature fusion (MFF) block includes a local branch, a global branch, and an α-Gate fusion unit. The local branch extracts local forgery trace features from the face image using a Local Feature Enhancement Network (LFAN) block. The global branch extracts global semantic inconsistency features from the face image using an MLP-Mixer structure. The α-Gate fusion unit generates adaptive fusion weights based on the local forgery trace features and the global semantic inconsistency features, and performs weighted fusion of the local forgery trace features and the global semantic inconsistency features to obtain the fused features. (2) Construct a network structure for lightweight channel attention network blocks; The lightweight channel attention network LCAN block is based on an inverted residual structure and introduces ECA channel attention to enhance the correlation modeling capability between feature channels. (3) Construct the detector network structure based on multi-feature fusion blocks and lightweight channel attention network blocks; The detector network includes a backbone network formed by stacking LCAN blocks and MFF blocks, and a classification head connected to the backbone network. The classification head is used to generate a single prediction value based on the features output by the backbone network, and obtain the probability that the face image to be detected is a fake face image through the Sigmoid function. (4) Set the weighted binary cross-entropy loss function; (5) Set hyperparameters during the training process, including hyperparameters of the network structure, hyperparameters of network training, and class weights between the real class and the fake class; (6) Train the detector network; This step involves training the detector network using training face images with real labels, and updating the network weights of the detector network according to the weighted binary cross-entropy loss function to obtain the trained detector network.

2. The detector training method for generalized detection of deepfake faces according to claim 1, characterized in that: After step (6), the detector is tested for performance to verify whether it can effectively identify whether the face image to be detected is a real face image or a fake face image in multiple datasets in both in-database and out-of-database scenarios. During the verification process, the face image to be detected is read, the face image to be detected is input into the trained detector network, and the detection result of whether the face image to be detected is a real face image or a fake face image is output.

3. The detector training method for generalized detection of deepfake faces according to claim 1, characterized in that: The local feature enhancement network block in step (1) consists of a channel adjustment layer, a depthwise separable convolutional layer, a coordinate convolutional layer, a dynamic convolutional layer, and a convolutional feedforward network layer. The dynamic convolutional layer can dynamically adjust the kernel weights according to the input features, and more flexibly focus on the subtle and inconsistent local features in the forgery traces. The dynamic convolutional layer is implemented using conditional convolution, and the specific implementation is as follows: ; ; ; ; Where X represents the input feature map of the dynamic convolutional layer, with dimension 1. GAP represents average pooling, and K represents the predetermined number of expert convolutional kernels, set to 8. This represents the k-th routing function. The normalized attention weights of the k-th expert convolutional kernel are calculated using the softmax function. This represents the k-th learnable convolutional kernel; Y represents the dynamic convolution kernel obtained by weighted fusion of all expert convolution kernels; Y represents the output feature with dimension 1. ; The global feature processing uses L stacked Mixer layers, each Mixer layer including Token-mixing MLP and Channel-mixing MLP; Token-mixing MLP: Performs information exchange in the spatial dimension, applying a multilayer perceptron independently to each channel; Channel-mixing MLP: Performs feature transformation in the channel dimension, applying an MLP independently to each spatial location, with the following formula: ; ; ; in, and All are two-layer fully connected networks with a GELU activation function in between. Z represents the input feature with dimension N×C, U represents the intermediate feature after layer normalization of the input Z, V represents the output feature after token-mixing operation and residual connection, and Y represents the final output feature of the MLP-Mixer layer.

4. The detector training method for generalized detection of deepfake faces according to claim 1, characterized in that: The fusion of global and local features employs an α-Gate mechanism: a weight α is generated by superimposing an adaptive average pooling layer, a 1×1 convolutional layer one, a ReLU activation function, a 1×1 convolutional layer two, and a Sigmoid activation function. This weighted fusion of local and global features is then performed. Finally, a residual connection is used to add the original input features to the fused features, resulting in an enhanced feature representation that is more complementary to both global and local features. The specific formula is as follows: ; ; in, and For learnable matrix weights, For the Sigmoid function, Indicates average pooling. Indicates local features, Represents global features. and These represent the fused features and the input features, respectively. With reference to the idea of residual network, the fused features are element-wise added to the input features to obtain enhanced feature representation with global and local complementarity, and the output feature F out is: 。 5. The detector training method for generalized detection of deepfake faces according to claim 1, characterized in that: In step (2), the LCAN lightweight channel attention network block is based on the inverted residual structure. It introduces channel attention to enhance cross-channel interaction. First, it passes through a channel adjustment layer, then a convolutional layer to expand the channel, a deep convolutional layer to perform filtering operations, ECA channel attention, and a convolutional layer to compress the channel. At the same time, linear activation is used to reduce information loss. Finally, it is connected to the input to obtain the output of the lightweight channel attention network block.

6. The detector training method for generalized detection of deepfake faces according to claim 1, characterized in that: In step (3), the method for constructing the detector network structure based on the MFF block and LCAN block is as follows: (3-1) Backbone network: The backbone network adopts an architecture of alternating LCAN blocks and MFF blocks. First, a standard convolutional layer is used to obtain shallow features of the image to be detected, and the spatial dimension of the features is reduced to half of that of the image to be detected. Then, multiple LCAN blocks are stacked to process the shallow features to obtain intermediate features. Furthermore, the intermediate features are processed by stacking MFF blocks and LCAN blocks. (3-2) Classification Head: The classification head adopts an architecture of convolutional layer, adaptive average pooling layer, dropout layer and linear layer; firstly, the input feature channels are adjusted by convolutional layer, then the feature map is compressed by adaptive average pooling layer, then the dropout layer is applied to prevent overfitting, and finally the spoofing probability is output by linear layer.

7. The detector training method for generalized detection of deepfake faces according to claim 1, characterized in that: In step (4), the weighted binary cross-entropy function is used to train the detector's detection capability, and the calculation formula is as follows: ; Where B is the batch size, y j Represents the true label of the j-th image, when y j =1 indicates a forged face, y j =0 indicates a real face, p j The output of the detector is the forgery probability, where w0=2.9 and w1=1 are the class weights of the real class and the forgery class, respectively.

8. The detector training method for generalized detection of deepfake faces according to claim 1, characterized in that: In step (5), the hyperparameters for network training are set as follows: (5-1) Setting up training and testing images: Scale or crop the training and testing images to 255×255 pixels; (5-2) Hyperparameter settings for network training: The batch size used for training is 64, the training is 100 epochs, gradient accumulation and mixed precision training are used, the learning rate is 0.001, and linear warm-up and cosine are used to schedule the learning rate. The optimizer used for training is Adam. (5-3) The weight ratio of the fake class to the real class is 1:2.9; (5-4) Setting up training and testing data: The detector network is trained using the first training dataset; the trained detector network is validated within the library using the first test dataset which is from the same source as the first training dataset. The trained detector network is validated externally using a second test dataset and a third test dataset that are from a different source than the first training dataset.

9. The detector training method for generalized detection of deepfake faces according to claim 1, characterized in that: In step (6), the specific training steps of the detector network are as follows: (6-1) Load training data. In order to enhance the generalization effect of the network, data augmentation is performed on the training data. Data augmentation performs random transformation operations on each batch of loaded training data, including scaling, horizontal flipping, color jittering, Gaussian blur, Gaussian noise, JPEG compression, random occlusion, and affine transformation. (6-2) Input the loaded training data into the detector network to obtain the classification result, and use the loss function to calculate the loss on the classification result and the true label of the training data; (6-3) Update the network weights based on the loss using the Adam optimizer; (6-4) Repeat the above steps, iterating through each batch, until the total number of training rounds reaches the specified value.

10. The detector training method for generalized detection of deepfake faces according to claim 2, characterized in that: The trained detector network is subjected to performance testing to verify whether it can perform generalized recognition of deep fake faces on multiple datasets. The specific steps are as follows: Read the three test set images in step (5) in sequence, input the test sets into the detector network in batches, obtain the detection results, and complete the test. A deepfake face generalization detection device, characterized in that it includes: The network structure construction module is used to construct the network structure of the multi-feature fusion (MFF) block and the network structure of the lightweight channel attention network (LCAN) block, and to construct the structure of the detector network based on the multi-feature fusion (MFF) block and the lightweight channel attention network (LCAN) block. The loss function setting module is used to set the weighted binary cross-entropy loss function for training the detector network; The hyperparameter setting module is used to set hyperparameters during the training process. The hyperparameters include hyperparameters of the network structure, hyperparameters of network training, and class weights between the real class and the fake class. The training module is used to train the detector network using training data; The detection testing module is used to perform performance testing on the trained detector network, read the images to be detected from multiple datasets, input the images to be detected into the trained detector network in sequence, and output the detection result of whether the image to be detected is a real face or a fake face. Wherein, the multi-feature fusion (MFF) block, the lightweight channel attention network (LCAN) block, the detector network, the weighted binary cross-entropy loss function, the hyperparameters, the training module, and the detection and testing module are respectively used to perform the corresponding steps in the detector training method for generalized detection of deep fake faces as described in any one of claims 1 to 10.