A Deepfake Detection Algorithm Based on a Domain Generalization Framework for Facial Semantic Content Decomposition

By adopting a two-branch feature extractor based on facial semantic content decomposition and an asymmetric alignment constraint module with maximum mean difference in the deep forgery video detection algorithm, the problem of poor generalization of the detection model in the prior art is solved, and better cross-domain detection performance is achieved.

CN116343279BActive Publication Date: 2025-06-20SICHUAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211224568.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-09
Publication Date
2025-06-20
Estimated Expiration
2042-10-09

AI Technical Summary

Technical Problem

The existing deep-fake video detection algorithms perform poorly in cross-domain detection, mainly due to the model overfitting features unique to a single source domain, resulting in poor generalization.

Method used

The dual-branch feature extractor FD-DBN based on facial semantic content decomposition is adopted to extract global and local features through convolutional neural networks, and the fusion module CAFM fusion features of the coordinate attention mechanism is used, and combined with the asymmetric alignment constraint module MAAC with the maximum mean difference, the network framework is optimized to improve the generalization ability of the detection model.

Benefits of technology

It effectively overcomes the problem of overfitting face semantic content by convolutional neural networks, forcing feature extractors to focus on more subtle, local and shared discriminant features, and improves the detection performance of the model in multiple source domains and target domains.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116343279B_ABST
    Figure CN116343279B_ABST
Patent Text Reader

Abstract

The present invention discloses a deepfake detection algorithm based on a domain generalization framework for facial semantic content decomposition, which relates to the field of image passive forensics. Aiming at the problem that the existing deep learning-based deepfake detection algorithms overfit the facial semantic information in a certain source domain, resulting in poor generalization performance of the learned model, a targeted solution is proposed. Specifically, a two-branch feature extractor based on facial semantic content decomposition is constructed to extract more local, subtle and shared trace features. In addition, a domain generalization framework is introduced, and an asymmetric alignment constraint module based on the maximum mean discrepancy is constructed. By aligning the feature distributions of multiple source domains, a shared feature space and decision boundary can be learned, which can be used for the detection of unknown target domains. The present invention can effectively improve the accuracy of deepfake detection and has practical application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of image processing and forensics, and particularly to a deepfake face detection algorithm based on domain generalization. Background Art

[0002] Deepfake video detection has become an urgent problem in multimedia forensics technology in the field of information security. Existing deepfake passive forensics algorithms are mainly divided into two categories: algorithms based on traditional handcrafted features and data-driven algorithms. Among them, the algorithms based on traditional handcrafted features are realized by extracting statistical features, biometric signal features, etc. that are inconsistent with natural images brought about by tampering operations. However, such detection algorithms only show good detection capabilities on some early datasets. With the continuous maturity of the generation model and the focus on eliminating obvious artifacts in the generation algorithm, the detection capabilities of such detection algorithms have been greatly weakened. In contrast, data-driven algorithms based on deep learning can more effectively detect deepfake videos by adaptively extracting the artifact features existing in images and videos through convolutional neural networks or recurrent neural networks. The specific process is to collect a large number of deepfake datasets and feed them all into a specific neural network (such as ResNet, Xception, Transformer, etc.), encouraging the network to adaptively learn the feature distribution of the training set to achieve accurate prediction of the detection model on the target dataset. However, the high-dimensional features of images extracted by existing data-driven detection algorithms are highly dataset-biased, that is, overfitting to the training set, resulting in poor detection performance of the pre-trained model in unknown datasets, that is, the generalization of the detection model is poor. "Thinking in frequency: Face forgery detection by mining frequency-aware clues" (Qian Y, Yin G, Sheng L, et al. Thinking in frequency: Face forgery detection by mining frequency-aware clues[C] / / European Conference on Computer Vision. Springer, vol. 12357, pp. 86-103, 2020) effectively improves the detection ability of the model in the domain by learning the frequency-domain feature differences between deepfake image data and natural real images. However, this method focuses on the frequency-domain differences between real and forged in a certain data domain, and once the learned model is applied to other data domains, the detection ability will drop significantly.Qian Y, Yin G, Sheng L, et al. Thinking in frequency: Face forgery detection by mining frequency-aware clues[C] / / European conference on computer vision. Springer, Cham, 2020: 86-103. In the paper "Improving the Efficiency and Robustness of Deepfakes Detection through Precise Geometric Features", it is pointed out that in the existing deepfake video generation process, no targeted processing is done on the consistency between frames. Therefore, deepfake detection is achieved by learning the differences in the continuity of facial key points between frames of real and forged videos. Such algorithms also learn the temporal features in a certain data domain and have extremely poor performance in cross-domain detection. In real applications, given a video to be detected, it is impossible to know its specific tampering method, which greatly limits the application of deep learning-based detection algorithms. Therefore, it is crucial to improve the generalization ability of the detection model. Generally speaking, most of the existing deepfake detection algorithms based on deep learning design a feature extractor based on a convolutional neural network to extract specific features of a single source domain. The learned features will inevitably overfit the unique traces of that source domain. According to the specific property that the convolutional neural network will focus on regions with rich textures, the facial regions it focuses on are generally local semantic information such as eyes, nose, mouth, etc. And for different source domains, the regions used to judge authenticity are inconsistent, resulting in poor detection performance of the model trained in a single source domain in other unknown target domains. Summary of the Invention

[0003] The purpose of the present invention is to solve the above limitations and provide a domain generalization network framework based on facial semantic content decomposition to improve the generalization of the deepfake detection model and enable it to overcome the above deficiencies of the prior art.

[0004] The technical solution for achieving the purpose of the present invention is as follows:

[0005] A deepfake detection algorithm based on a domain generalization framework for facial semantic content decomposition, comprising the following steps:

[0006] Step 1: Obtain multiple existing publicly available deepfake datasets. Different datasets are generated by different tampering methods, and each dataset has authenticity labels and domain labels of the videos. Use any dataset in the deepfake dataset as the target domain to be tested, and use the datasets other than the target domain as the source domain for training. Perform preprocessing operations on the data in the target domain and source domain, including face detection, cropping, and alignment.

[0007] Step 2: Construct a dual-branch feature extractor FD-DBN based on facial semantic content decomposition, and feed the source domain datasets in the preprocessed deepfake dataset in Step 1 into FD-DBN, where K represents the number of datasets in the source domain, and the superscript i represents the serial number of a certain dataset in the source domain. The specific processing method is as follows: on the one hand, use a convolutional neural network to encode the entire face to extract global features; on the other hand, decompose the face image through a diagonal-directional cross-shuffling module DCSM (Diagonal-directional Cross-Shuffling Module), and use a convolutional neural network to encode the decomposed face image to extract local features. Finally, fuse the above-extracted global and local features through a coordinate attention mechanism-based fusion module CAFM to obtain the feature sets of each source domain data.

[0008] Step 3:

[0009] 3.1 Through a maximum mean discrepancy-based asymmetric alignment constraint module MAAC (MMD-based Asymmetric Aligning Constraint) with real-feature clustering constraint RFC and intra-inter-domain triplet constraint ICT, perform distribution adaptation on each source domain data feature set FT extracted in Step 2 s and calculate the maximum mean discrepancy-based asymmetric alignment constraint loss

[0010] Among them, MAAC takes into account the data distribution characteristics of multiple source domains, that is, the data distribution differences of forged image data generated by different tampering methods are relatively large, so it is difficult to align the forged images of all source domains. In contrast, since all real images are natural images, they are easier to align. Therefore, MAAC performs distribution adaptation by first using the real-feature clustering constraint RFC (Real-Feature Clustering Constraint) to align the real data distributions of multiple source domains, and then using the intra-cross triplet constraint ICT (Intra-Cross Triplet Constraint) to constrain the separation of real and forged samples in the feature space. The corresponding adaptation loss is:

[0011]

[0012] 3.2 The conventional deepfake authenticity classification module combines the feature sets FT of each source domain data extracted in step 2 s with the authenticity labels to calculate the classification loss

[0013] Step 4: Combine the content described in step 3 and to calculate the total loss Train the domain generalization framework model based on facial semantic content decomposition through backpropagation; obtain the target domain generalization framework model FDDG based on facial semantic content decomposition;

[0014] The final loss function of the network is shown as follows:

[0015]

[0016] Step 5: Use the target domain generalization framework model FDDG obtained by training in step 4 to extract the feature of the target domain data to be measured and calculate the prediction score p. If p > 0.5, it is determined that the data to be measured is true, otherwise it is false, thus completing the detection of the deepfake face video.

[0017] The present invention also proposes a deepfake detection system based on a domain generalization framework for facial semantic content decomposition, including:

[0018] Module 1: Obtain multiple existing publicly available deepfake data sets. Different data sets are generated by different tampering methods, and each data set has authenticity labels and domain affiliation labels of the video; use any data set in the deepfake data set as the target domain to be measured, and use the data sets other than the target domain as the source domain for training; perform preprocessing operations on the data in the target domain and source domain, including face detection, cropping, and alignment.

[0019] Module 2: Construct a two-branch feature extractor FD-DBN based on facial semantic content decomposition, and feed the source domain data sets in the deepfake data set preprocessed by Module 1 into FD-DBN, where K represents the number of data sets in the source domain, and the superscript i represents the serial number of a certain data set in the source domain; the specific processing method is, on the one hand, use a convolutional neural network to encode the entire face to extract global features, and on the other hand, decompose the face image through the diagonal direction face image block scrambling module DCSM, and use a convolutional neural network to encode the decomposed face image to extract local features. Finally, fuse the global and local features extracted above through the coordinate attention mechanism-based fusion module CAFM to obtain the feature sets of each source domain

[0020] Module 3:

[0021] 3.1: The maximum mean discrepancy based asymmetric alignment constraint module MAAC with real feature aggregation constraint RFC and intra-domain - inter-domain triple constraint ICT is used to perform distribution adaptation on each source domain data feature set FT extracted by Module 2 s and calculate the maximum mean discrepancy based asymmetric alignment constraint loss

[0022] Among them, MAAC takes into account the data distribution characteristics of multiple source domains, that is, the data distribution of forged image data generated by different tampering methods varies greatly, so it is difficult to align the forged images of all source domains. In contrast, since all real images are natural images, they are easier to align. Therefore, MAAC performs distribution adaptation by first using RFC to align the real data distributions of multiple source domains and then using ICT to constrain the separation of real and forged samples in the feature space. The corresponding adaptation loss is:

[0023]

[0024] 3.2: The conventional deepfake authenticity classification module is used to calculate the classification loss by combining the authenticity labels with each source domain data feature set FT extracted by Module 2 s combined with the authenticity labels

[0025] Module 4: Used to combine what is described in Module 3 and calculate the total loss Train the domain generalization framework model based on facial semantic content decomposition through backpropagation; obtain the target domain generalization framework model FDDG based on facial semantic content decomposition;

[0026] The final loss function of the network is shown as follows:

[0027]

[0028] Module 5: Used to extract the data features of the target domain to be measured by the target domain generalization framework model FDDG obtained by training with Module 4 and calculate the prediction score p. If p > 0.5, it is determined that the data to be measured is genuine, otherwise it is fake, thus completing the detection of deepfake face videos.

[0029] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0030] (1) The present invention utilizes a feature extractor based on facial semantic content decomposition, which can solve the problem that the existing detection algorithms of convolutional neural networks overfit the facial semantic content and ignore the unique traces of deep fakes, forcing the feature extractor to focus on more subtle, local, and shared discriminative features;

[0031] (2) By adopting a domain generalization framework based on facial semantic content decomposition, it can simultaneously consider the feature distributions of multiple source domains, encourage the network to select a feature space shared by multiple source domains, and this feature space can be applied to other unknown but related target domain data, solving the defect that the original algorithm would learn the dataset offset features;

[0032] (3) Since the tampering types of multiple source domains are different and the data distributions vary greatly, it is difficult to align the forged data. In contrast, the real data are all natural images without tampering and are easier to fit. Therefore, an asymmetric alignment constraint based on the maximum mean discrepancy is proposed, which can better optimize the network framework and improve the detection speed and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 is the feature extraction process of the prior art

[0034] Figure 2 is the flowchart of the embodiment of the present invention.

[0035] Figure 3 is the framework structure diagram of the embodiment of the present invention.

[0036] Figure 4 is the schematic diagram of the dual-branch feature extractor based on facial semantic content decomposition of the embodiment of the present invention.

[0037] Figure 5 is the schematic diagram of the asymmetric alignment constraint feature distribution based on the maximum mean discrepancy of the embodiment of the present invention.

[0038] Figure 6 is the effect diagram of the ablation experiment of the embodiment of the present invention.

[0039] Figure 7 is the list diagram of the comparison test results of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments.

[0041] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings and specific embodiments. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. For those skilled in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0042] As Figure 1 shown, the existing deepfake detection algorithm based on convolutional neural network is data-driven, which causes the detection model to overfit the features with dataset shift, resulting in poor generalization of the model and limiting the practical application of the detection algorithm. We believe that due to the characteristic that the convolutional neural network will pay more attention to regions with complex textures, the existing algorithms based on convolutional neural network will overly focus on a certain semantic region of the face, such as eyes, mouth, nose, etc., and for the same model, the regions it focuses on for genuine and fake images of different datasets and different categories are inconsistent, which leads to difficulty in obtaining good detection performance when the model trained in the source domain is migrated to an unknown target domain. Therefore, we propose a domain generalization framework based on facial semantic content decomposition to solve the problem of poor generalization in deepfake detection.

[0043] As Figures 2 - 3 shown, the embodiment of the present invention provides a domain generalization framework based on facial semantic content decomposition. The detection framework includes the following steps:

[0044] S1: Obtain multiple existing publicly available deepfake datasets. Different datasets are generated by different tampering methods, and each dataset has the authenticity label and domain label of the video. Use any dataset in the deepfake dataset as the target domain to be tested, and use the datasets other than the target domain as the source domain for training. For the embodiment of the present invention, we divide the datasets into K source domains according to different tampering methods, and use the face recognition algorithm MTCNN to perform preprocessing operations including face detection, cropping, and alignment on the data in the target domain and source domain;

[0045] S2: As Figure 4 shown, construct a two-branch feature extractor FD-DBN based on facial semantic content decomposition, and use the source domain dataset in the preprocessed deepfake dataset in step 1 Feed it into the FD-DBN, where K represents the number of datasets in the source domain, and the superscript i represents the serial number of a certain dataset in the source domain; the specific processing method is that, on the one hand, a convolutional neural network is used to encode the entire face to extract global features, and on the other hand, the face image is decomposed by the diagonal direction face image block scrambling module DCSM, and a convolutional neural network is used to encode the decomposed face image to extract local features. Finally, the global and local features extracted above are fused through the coordinate attention mechanism-based fusion module CAFM to obtain the data feature sets of each source domain.

[0046] S2.1: Construct the DCSM, and the specific steps are as follows:

[0047] (1) For a given face image I, it is evenly divided into N×N non-overlapping sub-blocks. We use a matrix A of size N×N to represent the sequence of all sub-blocks, where A(h,w) = (h,w), h,w ∈ {1,…,N};

[0048] (2) Use the cross-scrambling mechanism based on the diagonal direction to scramble the matrix A. The scrambled matrix A′ is expressed as:

[0049]

[0050] Among them, F represents a transformation matrix of size N×N with all sub-diagonal elements being 1 and other elements being 0, which can be expressed as:

[0051]

[0052] (3) Scramble the image block sequence of the original image according to the transformation method in 2). The scrambled image I′ can be expressed as:

[0053]

[0054] S2.2: Use the convolutional neural network XceptionNet to extract the global feature f of the face from the entire face g , and extract the local feature f of the face from the face image preprocessed by the DCSM in S2.1 l ;

[0055] S2.3: Construct a feature fusion module based on the coordinate attention mechanism, and its specific steps are as follows:

[0056] (1) For the input feature f of size (H,W,C) g and f l, where H represents height, W represents width, and C represents the number of feature channels. Pooling operations are performed using pooling kernels of sizes (H,1) and (1,W) respectively. Then, the outputs corresponding to the c-th channel of the global feature in the horizontal and vertical directions are:

[0057]

[0058]

[0059] The outputs corresponding to the c-th channel of the local feature in the horizontal and vertical directions are:

[0060]

[0061]

[0062] where the superscripts h and w represent the vertical and horizontal directions, the subscripts g and l represent the global and local features, and the subscript c represents the c-th channel.

[0063] (2) Fuse the features in the two directions respectively. The global feature can be expressed as:

[0064]

[0065] The local feature can be expressed as:

[0066]

[0067] where [·,·] represents concatenation along the spatial dimension, F1(·) represents a shared 1×1 convolution operation, and δ is a non-linear activation function.

[0068] (3) Divide the global feature F g into two independent tensors along the spatial dimension and and use two 1×1 convolution operations F gh and F gw to map and to tensors with the same number of channels as the initial global feature input f g which are expressed as:

[0069]

[0070]

[0071] Divide the local feature F l into two independent tensors along the spatial dimension and and use two 1×1 convolution operations F lh and F lwMap and to a tensor with the same number of channels as the initial global feature input f l , denoted as:

[0072]

[0073]

[0074] (4) Obtain the enhanced global feature Y g and the local feature Y l , which can be expressed as:

[0075]

[0076]

[0077] (5) Concatenate the enhanced global feature Y g obtained above and the local feature Y l channel-wise:

[0078]

[0079] S3:

[0080] S3.1: Construct an asymmetric alignment constraint module based on maximum mean discrepancy (MMD-based Asymmetric Aligning Constraint, MAAC), which includes a real-feature clustering constraint (Real-Feature Clustering Constraint, RFC) and an intra-cross triplet constraint (Intra-Cross Triplet Constraint, ICT), perform distribution adaptation on the feature of each source domain data extracted by S2, and calculate the asymmetric alignment constraint loss based on maximum mean discrepancy

[0081] Among them, MAAC takes into account the data distribution characteristics of multiple source domains, that is, the data distribution differences of forged image data generated by different tampering methods are relatively large, so it is difficult to align the forged images of all source domains. In contrast, since all real images are natural images, they are easier to align. Therefore, MAAC performs distribution adaptation by first using RFC to align the real data distributions of multiple source domains, and then using ICT to constrain the separation of real and forged samples in the feature space. The specific steps are as follows:

[0082] (1) Use RFC to align the feature distributions of real samples in multiple source domains in the learned feature space. For two given source domains S l and S t (whose feature distributions are different and are respectively and )'s characteristics and where n l and n t respectively represent the number of samples in two source domains. First, map them to the Reproducing Kernel Hilbert Space (RKHS) respectively, and calculate the feature distance between the two source domains in the RKHS:

[0083]

[0084] where, represents mapping the source domain data to the RKHS, and k(·,·) represents the kernel function determined by φ(·). In the example of the present invention, we adopt the Gaussian kernel function, that is where σ is the bandwidth parameter.

[0085] By minimizing the distance between the feature distributions of the real samples in multiple source domains, to align the feature distributions of the real samples in multiple source domains, adopt R m represents the real feature from the m-th source domain, m ∈ {1,…,M}, and adopt represents all samples in the m-th source domain, then the RFC loss function can be described as:

[0086]

[0087] (2) To solve the problem that the feature distance between the same categories within and between domains is greater than the feature distance between different categories, the example of the present invention designs ICT, that is, in the low-dimensional feature space, constrain the feature distribution distance between real and forged samples to be as large as possible, then the ICT loss function can be described as:

[0088]

[0089] where, r and f are used to refer to real and forged samples respectively, and the subscript (m,n,q) refers to the serial numbers of different source domains.

[0090] Therefore, MAAC is composed of two sub-constraints, RFC and ICT, to achieve the adaptation constraint of multi-source domain features. The corresponding adaptation loss is:

[0091]

[0092] S3.2: Through the deepfake authenticity classification module, input the feature of each source domain data extracted by S2 into the deepfake authenticity classification module to obtain the model prediction result, and calculate the classification loss in combination with the authenticity label

[0093]

[0094] S4: Combine as described in S3 with calculate the total loss Train the domain generalization framework model based on facial semantic content decomposition through backpropagation; obtain the target domain generalization framework model based on facial semantic content decomposition;

[0095] The final loss function of the network is shown as follows:

[0096]

[0097] S5: Use the trained domain generalization framework model based on facial semantic content decomposition to extract the feature of the data of the target domain to be tested and calculate the prediction score p. If p > 0.5, it is determined that the data to be tested is true, otherwise it is false, so as to complete the detection of deepfake face videos.

[0098] In this embodiment, we select 5 datasets, namely FaceForensics++ (FF++), Celeb-DF-v2 (CDF-v2), DeepFake Detection Challenge (DFDC), Deep Fake Detection (DFD), and DeeperForensics-1.0 (DFo). Among them, λ, m1, and m2 used for constructing the intra-domain - inter-domain triplet constraints are set to 0.25, 0.1, and 0.5 respectively. The hyperparameters α, β, and γ used for assigning weights to the real feature aggregation constraint, the intra-domain - inter-domain triplet constraint, and the classification loss are set to 0.1, 0.2, and 1.0 respectively.

[0099] This embodiment uses the accuracy rate (ACC) and the area under the ROC curve (AUC) as evaluation indicators.

[0100] Figure 6 For the ablation experiment of each component of the present invention, we will successively select 4 of the datasets as the source domain for training and test in the remaining one target domain. For the convenience of display, we use A, B, C, D, and E in the table to represent the datasets FF++, CDF-v2, DFDC, DFD, and DFo respectively. The results are as Figure 6As shown, the line "Backbone(Xception)" represents the performance of the backbone network before adding any components. The line "w / DCSM" represents the detection performance of the dual-branch detection model formed after introducing the diagonal face image patch scrambling module. The line "w / DCSM&CAFM" represents the detection performance after fusing the features of the two branches using CAFM. The line "w / DCSM&CAFM&DG" represents the detection performance after introducing the domain generalization framework and using MAAC for constraint. The results show that each component proposed in the present invention can effectively improve the model detection ability in multiple cross-domain experiments compared with the backbone network. Moreover, when all components are fused, the detection performance of the model of the present invention reaches the optimal level.

[0101] Figure 7 This is the comparative experiment of the present invention. To compare with previous algorithms, we also adopt the experimental setting of training on the FF++ dataset and testing on the CDF-v2, DFDC, and DFo datasets respectively. The experimental results are as Figure 7 shown. Compared with previous algorithms, the algorithm of the present invention achieves better detection performance in cross-domain experiments.

Claims

1. A deepfake detection algorithm based on a domain generalization framework for facial semantic content decomposition constructs a domain generalization framework for facial semantic content decomposition in deepfake detection to solve the problem of overfitting of facial semantic information in a single data domain, which leads to poor generalization performance of the learned model; and extracts more local and shared trace features through a dual-branch feature extractor based on facial semantic content decomposition. The asymmetric alignment constraint module based on the maximum mean discrepancy aligns the feature distributions of multiple source domains, combines the asymmetric alignment constraint loss based on the maximum mean discrepancy with the classification loss to calculate the total loss, and trains and obtains the target domain generalization framework model FDDG based on facial semantic content decomposition, including the following main steps: Step 1: Obtain multiple existing publicly available deepfake datasets, which are generated by different tampering methods, and each dataset has the authenticity label and domain belonging label of the video; use any dataset in the deepfake dataset as the target domain to be tested, and use the datasets other than the target domain as the source domains for training; perform preprocessing operations on the data in the target domain and source domains, including face detection, cropping, and alignment; Step 2: Construct a dual-branch feature extractor FD-DBN based on facial semantic content decomposition, and feed the source domain dataset in the preprocessed deepfake dataset in Step 1 into FD-DBN, where K represents the number of datasets in the source domain, and the superscript i represents the serial number of a certain dataset in the source domain; specifically, on the one hand, a convolutional neural network is used to encode the entire face to extract global features, and on the other hand, the face image is decomposed by the diagonal direction face image patch scrambling module DCSM, and a convolutional neural network is used to encode the decomposed face image to extract local features. Finally, the global and local features extracted above are fused through the coordinate attention mechanism-based fusion module CAFM to obtain the feature sets of each source domain data Step 3: 3.1 Through the Maximum Mean Discrepancy-based Asymmetric Alignment Constraint module MAAC with Real Feature Clustering Constraint RFC and Intra-domain-Inter-domain Triplet Constraint ICT, perform distribution adaptation on each source domain data feature set FT extracted in step 2, and calculate the Maximum Mean Discrepancy-based asymmetric alignment constraint loss s ​ 3.2 The conventional deepfake authenticity classification module calculates the classification loss by combining the authenticity labels with each source domain data feature set FT extracted in step 2 s ​ Step 4: Combine what is described in Step 3 with calculate the total loss Train the domain generalization framework model based on facial semantic content decomposition through backpropagation; obtain the target domain generalization framework model FDDG based on facial semantic content decomposition; Step 5: Use the target domain generalization framework model FDDG based on facial semantic content decomposition obtained by training in Step 4 to extract the feature of the data in the target domain to be tested and calculate the prediction score p. If p>0.5, it is determined that the data to be tested is real, otherwise it is fake, thus completing the detection of deepfake face videos.

2. The deepfake detection algorithm according to claim 1, characterized in that In Step 2, the diagonal direction face image block scrambling module DCSM adopts the following working mechanism: (1) For a given face image I, it is evenly divided into N×N non-overlapping sub-blocks. We use a matrix A of size N×N to represent the sequence of all sub-blocks, where A(h,w)=(h,w), h,w∈{1,…,N}; (2) Scramble the matrix A using the cross-scrambling mechanism in the diagonal direction. The scrambled matrix A′ is expressed as: where, F represents a transformation matrix of size N×N with all elements on the secondary diagonal being 1 and other elements being 0, which can be expressed as: (3) Scramble the sequence of image blocks of the original image according to the transformation method in (2) above. The scrambled image I′ can be expressed as:

3. The deepfake detection algorithm according to claim 1, characterized in that, In Step 2, the specific method of the coordinate attention mechanism-based fusion module CAFM is: (1) For the input feature f of size (H, W, C) g and f l , where H represents the height, W represents the width, and C represents the number of feature channels. Pooling operations are performed using pooling kernels of sizes (H, 1) and (1, W) respectively. Then the outputs corresponding to the c-th channel of the global feature in the horizontal and vertical directions are as follows: The outputs corresponding to the c-th channel of the local features in the horizontal and vertical directions are: where, the superscripts h and w represent the vertical and horizontal directions, the subscripts g and l represent the global and local features, and the subscript c represents the c-th channel; (2) Fuse the features in the two directions respectively. The global feature can be expressed as: The local feature can be expressed as: where, [·,·] represents concatenation according to the spatial dimension, F1(·) represents a shared 1×1 convolution operation, and δ is a non-linear activation function; (3) Divide the global feature F g into two independent tensors according to the spatial dimension and and adopt two 1×1 convolution operations on F gh and F gw to and map them into tensors with the same number of channels as the initial global feature input f g which is expressed as: Divide the local feature F l into two independent tensors according to the spatial dimension and and adopt two 1×1 convolution operations on F lh and F lw Map and to tensors with the same number of channels as the initial global feature input f l which is expressed as: (4) Obtain the enhanced global feature Y g and the local feature Y l , which can be expressed as: (5) Cascade the enhanced global feature Y obtained above at the channel level g with the local feature Y l :

4. The deepfake detection algorithm based on the domain generalization framework for facial semantic content decomposition according to claim 1, wherein , In Step 3.1, the specific method of constructing the asymmetric alignment constraint module MAAC based on the maximum mean discrepancy is: First, use RFC to align the real data distributions of multiple source domains, and then use ICT to constrain the distribution adaptation of the feature of each source domain data extracted in Step 2 by separating the features of real and forged samples. The specific steps are as follows: (1) Align the feature distributions of real samples in multiple source domains in the learned feature space using RFC; for two source domains S l and S t with different distributions, their features can be represented as and where n l and n t represent the number of samples in the two source domains respectively. First, map them to the Reproducing Kernel Hilbert Space (RKHS) respectively, and calculate the feature distance between the two source domains in the RKHS: Among them, denotes mapping the source domain data to the RKHS, and k(·,·) denotes the kernel function determined by φ(·). We adopt the Gaussian kernel function, that is where σ is the bandwidth parameter; By minimizing the distances between the feature distributions of real samples in multiple source domains to align the feature distributions of real samples in multiple source domains, using R m denotes the real features from the n-th source domain, m ∈ {1, …, M}, and using to denote all samples in the m-th source domain, then the RFC loss function can be described as: (2) To solve the problem that the feature distance between the same categories within and between domains is greater than the feature distance between different categories, ICT is designed. Here, r and f are used to represent real and forged samples respectively, and the subscript (m, n, q) represents the serial numbers of each source domain. Then, the ICT loss function is described as: Therefore, MAAC consists of two sub-constraints, RFC and ICT, to implement the adaptation constraint of multi-source domain features, and the corresponding adaptation loss is as follows: