A spatial-frequency joint deepfake detection method based on dual-domain attention collaborative deformable convolution
Through the joint space-frequency deep fake detection method, dual-domain attention is used to coordinate deformable convolution to enhance the feature extraction capability, solve the problem of insufficient feature extraction in traditional methods, and achieve more efficient fake image recognition.
Patent Information
- Application Number
- CN202510910826.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-07-02
AI Technical Summary
Traditional deep fake detection methods lack feature extraction capabilities, lack attention to key areas and consideration of frequency domain information, making it difficult to effectively identify fake images.
A space-frequency joint deep fake detection method based on dual-domain attention collaborative deformable convolution is adopted. The feature extraction capability is enhanced through the space-frequency feature extraction structure, including spatial domain feature extraction, frequency domain feature extraction and bidirectional cross attention structure.
Improved forgery detection performance, enabling more effective differentiation between real and forged images, and enhanced security.
Smart Images

Figure CN120449934B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information security, and in particular to a space-frequency joint deep fake detection method based on dual-domain attention collaborative deformable convolution. Background Art
[0002] Deep learning technology, with its superior data processing capabilities and adaptive feature extraction capabilities, has become a core technology that represents the most significant aspect of artificial intelligence. Deep learning-based facial forgery technology can not only generate highly realistic virtual faces but also convey diverse emotions and semantic meanings by manipulating key features such as facial expressions and lip shape in videos or images. Furthermore, this technology can swap the facial features of a specific person with those of a target person, visually replacing one person's identity and making one person say or do something else. As deepfake technology continues to evolve, the authenticity of the generated content has significantly improved, even reaching a level that is difficult for the human eye to discern.
[0003] While facial forgery technology plays a positive and important role in promoting cultural and entertainment activities such as social media, animated films, and games, greatly enriching people's leisure time, its malicious use and lowered accessibility also pose challenges and threats to information and social security.
[0004] Since traditional deep fake detection methods mostly rely on convolutional neural networks for feature extraction, their feature extraction capabilities are insufficient and they are not good at capturing the forgery traces of forged images in the spatial domain. In addition, commonly used forgery detection methods also have the defects of lacking focus on key areas and insufficient consideration of frequency domain information.
[0005] The above content is only used to assist in understanding the technical solution of the present invention and does not constitute an admission that the above content is prior art. Summary of the Invention
[0006] The main purpose of this invention is to provide a space-frequency joint deep fake detection method based on dual-domain attention collaborative deformable convolution, aiming to solve the problems of insufficient feature extraction capabilities of traditional deep detection networks and insufficient attention to key areas.
[0007] To achieve the above objectives, the present invention provides a spatial-frequency joint deepfake detection method based on dual-domain attention collaborative deformable convolution, which comprises the following steps:
[0008] Obtaining a training sample set and a test sample set, and forming a deep learning network model based on a space-frequency feature extraction structure, wherein the training sample set and the test sample set are composed of real faces and forged faces, and the space-frequency feature extraction structure is composed of a spatial domain feature extraction structure, a frequency domain feature extraction structure, and a bidirectional cross attention structure;
[0009] Passing the training sample set to the deep learning network model and performing model training actions to obtain a deep fake detection model;
[0010] The test sample set is passed to the deep fake detection model, and the face authenticity classification action is performed to obtain the authenticity face detection result.
[0011] Optionally, the step of forming a deep learning network model based on the space-frequency feature extraction structure includes:
[0012] Based on deformable convolution and spatial-channel attention mechanism, the spatial domain feature extraction structure is formed;
[0013] Based on the high-frequency processing algorithm and the frequency domain learning algorithm of the feature map, the frequency domain feature extraction structure is formed;
[0014] Based on the spatial domain feature extraction structure, the frequency domain feature extraction structure and the bidirectional cross attention structure, forming the space-frequency feature extraction structure;
[0015] The deep learning network model is obtained according to the space-frequency feature extraction structure and network layer.
[0016] Optionally, the step of forming the spatial domain feature extraction structure based on deformable convolution and spatial-channel attention mechanism includes:
[0017] When a feature map is received, the feature map is passed to the spatial-channel attention mechanism to obtain a first spatial domain parameter; and
[0018] Perform a channel-wise splicing operation on the features in the feature map to obtain a second spatial domain parameter;
[0019] A spatial domain feature map corresponding to the feature map is obtained based on the first spatial domain parameter and the second spatial domain parameter.
[0020] Optionally, the step of forming the frequency domain feature extraction structure by using the high frequency processing algorithm and the frequency domain learning algorithm based on the feature map includes:
[0021] When a feature map is received, a high-frequency processing action is performed on the feature map based on the high-frequency processing algorithm to obtain a high-frequency processing result corresponding to the feature map;
[0022] The high-frequency processing result is transmitted to the frequency domain learning algorithm, a frequency domain learning action is performed, and a frequency domain feature map corresponding to the feature map is obtained.
[0023] Optionally, the step of forming the space-frequency feature extraction structure based on the space-domain feature extraction structure, the frequency-domain feature extraction structure and the bidirectional cross attention structure includes:
[0024] Transferring the spatial domain feature map and the frequency domain feature map to the bidirectional cross attention structure to obtain the cross attention results from the spatial domain to the frequency domain and the cross attention results from the frequency domain to the spatial domain corresponding to the feature map;
[0025] The cross-attention results from the spatial domain to the frequency domain and the cross-attention results from the frequency domain to the spatial domain are fused to obtain a target feature map.
[0026] Optionally, the step of transferring the training sample set to the deep learning network model and performing model training to obtain a deep fake detection model includes:
[0027] After the deep learning network model receives the training sample set, it performs a supervised training operation based on the training sample set to obtain the deep fake detection model.
[0028] Optionally, the step of transferring the test sample set to the deep fake detection model and performing a face authenticity classification operation to obtain a true or false face detection result includes:
[0029] After receiving the test sample set, the deep fake detection model performs a detection action on the test sample in the test sample set to determine whether the test sample is a real human face;
[0030] If the test sample is determined to be a real face, classify the test sample into a set of true and false face detection results;
[0031] If it is determined that the test sample is not a real face, the test sample is classified into a set of true and false face detection results of which are false.
[0032] In addition, to achieve the above-mentioned purpose, the present invention also provides an authenticity detection device, which includes a memory, a processor, and a space-frequency joint deep fake detection program based on dual-domain attention collaborative deformable convolution stored on the memory and run on the processor. When the space-frequency joint deep fake detection program based on dual-domain attention collaborative deformable convolution is executed by the processor, the steps of the space-frequency joint deep fake detection method based on dual-domain attention collaborative deformable convolution as described above are implemented.
[0033] In addition, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, on which is stored a space-frequency joint deep fake detection program based on dual-domain attention collaborative deformable convolution. When the space-frequency joint deep fake detection program based on dual-domain attention collaborative deformable convolution is executed by a processor, the steps of the space-frequency joint deep fake detection method based on dual-domain attention collaborative deformable convolution as described above are implemented.
[0034] The embodiment of the present invention provides a spatial-frequency joint deep fake detection method based on dual-domain attention and collaborative deformable convolution. By performing frame extraction, face extraction and other operations on the video data set, a training sample set and a test sample set consisting of real faces and fake faces are constructed; the spatial domain features of the training sample set are extracted based on the attention mechanism and collaborative deformable convolution, the frequency domain features of the training sample set are extracted based on high-frequency extraction and frequency domain learning strategies, and the spatial domain and frequency domain features are fused based on bidirectional cross attention to obtain the spatial-frequency features of the training sample set; a network is constructed to train the spatial-frequency features to obtain a deep fake detection model; and the test sample set is detected based on the deep fake detection model. By efficiently fusing and enhancing spatial domain features and frequency domain features, the network can extract richer features, thereby improving detection performance, effectively distinguishing real images from fake images, and improving security. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] The accompanying drawings are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present invention, and together with the specification, serve to explain the principles of the present invention. In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following is a brief introduction to the drawings required for describing the embodiments. Obviously, it is possible for a person of ordinary skill in the art to derive other drawings based on these drawings without inventive effort.
[0036] Figure 1 Schematic diagram of the architecture of the hardware operating environment of the authenticity detection device involved in an embodiment of the present invention;
[0037] Figure 2 This is a flow chart of an embodiment of the present invention's method for joint spatial-frequency deepfake detection based on dual-domain attention collaborative deformable convolution;
[0038] Figure 3 This is a schematic diagram of the flow structure of an embodiment of the present invention's space-frequency joint deepfake detection method based on dual-domain attention collaborative deformable convolution;
[0039] Figure 4 This is a flow chart of an embodiment of the spatial domain feature extraction structure in the space-frequency joint deep fake detection method based on dual-domain attention collaborative deformable convolution of the present invention;
[0040] Figure 5 This is a flow chart of an embodiment of the frequency domain feature extraction structure in the space-frequency joint deep fake detection method based on dual-domain attention collaborative deformable convolution of the present invention;
[0041] Figure 6 This is a flow chart of an embodiment of the bidirectional cross-attention structure in the space-frequency joint deep fake detection method based on dual-domain attention coordinated deformable convolution of the present invention;
[0042] Figure 7 This is a flow chart of the structure of an embodiment of a deep fake detection model in the space-frequency joint deep fake detection method based on dual-domain attention collaborative deformable convolution of the present invention.
[0043] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION
[0044] The present application discloses a space-frequency joint deepfake detection method based on dual-domain attention collaborative deformable convolution. The method obtains a training sample set and a test sample set, and forms a deep learning network model based on a space-frequency feature extraction structure. The training sample set and the test sample set are composed of real faces and forged faces. The space-frequency feature extraction structure is composed of a spatial domain feature extraction structure, a frequency domain feature extraction structure, and a bidirectional cross-attention structure. The method then transfers the training sample set to the deep learning network model, executes the model training action, and obtains a deepfake detection model. The method then transfers the test sample set to the deepfake detection model, executes the face authenticity classification action, and obtains the authenticity and fake face detection results. This method improves detection performance, effectively distinguishes real images from forged images, and enhances security.
[0045] To better understand the above technical solutions, exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments described herein. Rather, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.
[0046] As an implementation solution, Figure 1 This is a schematic diagram of the architecture of the hardware operating environment of the authenticity detection device involved in the embodiment of the present invention.
[0047] like Figure 1As shown, the authenticity detection device may include a processor 101, such as a central processing unit (CPU), a memory 102, and a communication bus 103. Memory 102 may be a high-speed random access memory (RAM) or a stable non-volatile memory (NVM), such as a disk drive. Memory 102 may also be a storage device independent of processor 101. Communication bus 103 facilitates communication between these components.
[0048] Those skilled in the art will understand that Figure 1 The structure shown in the figure does not constitute a limitation on the authenticity detection device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0049] like Figure 1 As shown, the memory 102, which is a computer-readable storage medium, may include an operating system, a data storage module, a network communication module, a user interface module, and a spatial-frequency joint deep fake detection method program based on dual-domain attention collaborative deformable convolution.
[0050] exist Figure 1 In the authenticity detection device shown, the processor 101 and the memory 102 can be set in the authenticity detection device. The authenticity detection device calls the spatial-frequency joint deep fake detection program based on dual-domain attention collaborative deformable convolution stored in the memory 102 through the processor 101 and performs the following operations:
[0051] Obtaining a training sample set and a test sample set, and forming a deep learning network model based on a space-frequency feature extraction structure, wherein the training sample set and the test sample set are composed of real faces and forged faces, and the space-frequency feature extraction structure is composed of a spatial domain feature extraction structure, a frequency domain feature extraction structure, and a bidirectional cross attention structure;
[0052] Passing the training sample set to the deep learning network model and performing model training actions to obtain a deep fake detection model;
[0053] The test sample set is passed to the deep fake detection model, and the face authenticity classification action is performed to obtain the authenticity face detection result.
[0054] In one embodiment, the processor 101 may be configured to call a spatial-frequency joint deepfake detection program based on dual-domain attention collaborative deformable convolution stored in the memory 102 and perform the following operations:
[0055] Based on deformable convolution and spatial-channel attention mechanism, the spatial domain feature extraction structure is formed;
[0056] Based on the high-frequency processing algorithm and the frequency domain learning algorithm of the feature map, the frequency domain feature extraction structure is formed;
[0057] Based on the spatial domain feature extraction structure, the frequency domain feature extraction structure and the bidirectional cross attention structure, forming the space-frequency feature extraction structure;
[0058] The deep learning network model is obtained according to the space-frequency feature extraction structure and network layer.
[0059] In one embodiment, the processor 101 may be configured to call a spatial-frequency joint deepfake detection program based on dual-domain attention collaborative deformable convolution stored in the memory 102 and perform the following operations:
[0060] When a feature map is received, the feature map is passed to the spatial-channel attention mechanism to obtain a first spatial domain parameter; and
[0061] Perform a channel-wise splicing operation on the features in the feature map to obtain a second spatial domain parameter;
[0062] A spatial domain feature map corresponding to the feature map is obtained based on the first spatial domain parameter and the second spatial domain parameter.
[0063] In one embodiment, the processor 101 may be configured to call a spatial-frequency joint deepfake detection program based on dual-domain attention collaborative deformable convolution stored in the memory 102 and perform the following operations:
[0064] When a feature map is received, a high-frequency processing action is performed on the feature map based on the high-frequency processing algorithm to obtain a high-frequency processing result corresponding to the feature map;
[0065] The high-frequency processing result is transmitted to the frequency domain learning algorithm, a frequency domain learning action is performed, and a frequency domain feature map corresponding to the feature map is obtained.
[0066] In one embodiment, the processor 101 may be configured to call a spatial-frequency joint deepfake detection program based on dual-domain attention collaborative deformable convolution stored in the memory 102 and perform the following operations:
[0067] Transferring the spatial domain feature map and the frequency domain feature map to the bidirectional cross attention structure to obtain the cross attention results from the spatial domain to the frequency domain and the cross attention results from the frequency domain to the spatial domain corresponding to the feature map;
[0068] The cross-attention results from the spatial domain to the frequency domain and the cross-attention results from the frequency domain to the spatial domain are fused to obtain a target feature map.
[0069] In one embodiment, the processor 101 may be configured to call a spatial-frequency joint deepfake detection program based on dual-domain attention collaborative deformable convolution stored in the memory 102 and perform the following operations:
[0070] After the deep learning network model receives the training sample set, it performs a supervised training operation based on the training sample set to obtain the deep fake detection model.
[0071] In one embodiment, the processor 101 may be configured to call a spatial-frequency joint deepfake detection program based on dual-domain attention collaborative deformable convolution stored in the memory 102 and perform the following operations:
[0072] After receiving the test sample set, the deep fake detection model performs a detection action on the test sample in the test sample set to determine whether the test sample is a real human face;
[0073] If the test sample is determined to be a real face, classify the test sample into a set of true and false face detection results;
[0074] If it is determined that the test sample is not a real face, the test sample is classified into a set of true and false face detection results of which are false.
[0075] In one embodiment, the processor 101 may be configured to call a spatial-frequency joint deepfake detection program based on dual-domain attention collaborative deformable convolution stored in the memory 102 and perform the following operations:
[0076] Obtaining a video data set, and dividing the video data set according to a preset ratio to obtain training video data and test video data;
[0077] Performing a frame extraction operation on the training video data and the test video data to obtain training image data and test image data;
[0078] A face extraction operation is performed based on the training image data and the test image data to obtain a training sample set and a test sample set including the real faces and the forged faces.
[0079] Based on the hardware architecture of the above-mentioned authenticity detection device, an embodiment of the present invention's space-frequency joint deep fake detection method based on dual-domain attention collaborative deformable convolution is proposed.
[0080] Reference Figure 2 In a first embodiment, the spatial-frequency joint deepfake detection method based on dual-domain attention collaborative deformable convolution includes the following steps:
[0081] Step S100: Obtain a training sample set and a test sample set, and form a deep learning network model based on a space-frequency feature extraction structure, wherein the training sample set and the test sample set are composed of real faces and fake faces, and the space-frequency feature extraction structure is composed of a spatial domain feature extraction structure, a frequency domain feature extraction structure and a bidirectional cross-attention structure.
[0082] In this embodiment, the space-frequency feature extraction structure is a deep learning network structure, which obtains richer and more detailed discriminant features by extracting and fusing space domain and frequency domain features.
[0083] Optionally, the spatial domain feature extraction structure is formed based on deformable convolution and spatial-channel attention mechanism; the frequency domain feature extraction structure is formed based on the high-frequency processing algorithm of the feature map and the frequency domain learning algorithm; and the space-frequency feature extraction structure is formed based on the spatial domain feature extraction structure, the frequency domain feature extraction structure and the bidirectional cross attention structure; then, the deep learning network model is obtained according to the space-frequency feature extraction structure and the network layer.
[0084] In this embodiment, a spatial domain feature extraction structure is obtained by introducing a deformable convolution kernel space-channel attention mechanism and performing a rational structural design.
[0085] As an optional implementation, when a feature map is received, the feature map is passed to the spatial-channel attention mechanism to obtain a first spatial domain parameter; and a channel-wise splicing action is performed on the features in the feature map to obtain a second spatial domain parameter; then, based on the first spatial domain parameter and the second spatial domain parameter, a spatial domain feature map corresponding to the feature map is obtained.
[0086] For example, let the input feature map be X, then the spatial domain feature map corresponding to the feature map X is The calculation formula can be expressed as:
[0087]
[0088] Among them, the first spatial domain parameter is The second spatial domain parameter is , represents the spatial-channel attention mechanism, Represents feature disruption, Representative features are spliced along the channel, Represents the result of the convolution operation on X, Represents the result of the deformable convolution operation on X.
[0089] It can be understood that in the process of spatial domain feature extraction, deformable convolution is used to capture local distortion traces of the face, such as unnatural edge excess, abnormal facial muscle movement, etc.; the spatial-channel attention mechanism is used to focus on areas that are easy to expose forgeries, such as dynamic areas such as the eyes and lips.
[0090] The spatial attention mechanism can enhance forgery-sensitive areas, such as abnormal eyebrow shaking; the channel attention mechanism can dynamically adjust the feature channel weights to suppress irrelevant background interference.
[0091] In this embodiment, a frequency domain feature extraction structure is obtained by fusing a high frequency processing algorithm of a feature map with a frequency domain learning algorithm, and is used to obtain frequency domain features carrying key frequency domain information.
[0092] As an optional implementation method, when a feature map is received, a high-frequency processing action is performed on the feature map based on the high-frequency processing algorithm to obtain a high-frequency processing result corresponding to the feature map; then, the high-frequency processing result is passed to the frequency domain learning algorithm, and a frequency domain learning action is performed to obtain a frequency domain feature map corresponding to the feature map.
[0093] For example, let the input feature map be X, then the frequency domain feature map corresponding to the feature map X is The calculation formula can be expressed as:
[0094]
[0095] in, stands for frequency domain learning, Represents high-frequency processing of feature map X.
[0096] It can be understood that the implementation process of the frequency domain feature extraction structure achieves accurate capture of deep fake frequency domain features through the cascade processing of high-frequency signal analysis and frequency domain table learning.
[0097] In this embodiment, a bidirectional cross-attention mechanism is used to cross-enhance and fuse spatial and frequency domain features to obtain more comprehensive and detailed spatial-frequency features.
[0098] As an optional implementation, the spatial domain feature map and the frequency domain feature map are transmitted to the bidirectional cross-attention structure to obtain the spatial domain to frequency domain cross-attention results and the frequency domain to spatial domain cross-attention results corresponding to the feature maps; then, the spatial domain to frequency domain cross-attention results and the frequency domain to spatial domain cross-attention results are fused to obtain the target feature map. It can be understood that the spatial domain feature map includes spatial domain features, and the frequency domain feature map includes frequency domain features.
[0099] For example, convolution operations are performed on the spatial domain features and the frequency domain features to obtain their respective key, value, and query matrices. 、 、 、 、 、 , respectively, in the spatial domain and frequency domain and Perform Attention calculation and convert the frequency domain With spatial domain and Perform attention calculation to obtain enhanced frequency domain features guided by important spatial features, and enhanced spatial domain features guided by important frequency domain features. It can be expressed as:
[0100]
[0101] in, Represents the cross-attention result from spatial domain to frequency domain, Represents the cross-attention result from the frequency domain to the spatial domain. The two will be multiplied by a weight matrix of the same dimension as them and then fused to obtain the final feature map, which is the target feature map mentioned above.
[0102] It can be understood that the bidirectional cross-attention mechanism is a bidirectional cross-modal attention mechanism. The process of fusing the spatial domain and the frequency domain realizes the deep synergy of the spatial domain and the frequency domain features through the bidirectional cross-modal attention mechanism.
[0103] In this embodiment, network layers such as downsampling and classification are added on the basis of the space-frequency feature extraction structure to realize the construction of a deep learning network model.
[0104] Furthermore, the step of obtaining a training sample set and a test sample set includes obtaining a training sample set and a test sample set consisting of real faces and forged faces by processing a public video dataset.
[0105] In this embodiment, a video data set is first obtained, and the video data set is divided according to a preset ratio to obtain training video data and test video data; then, a frame extraction operation is performed on the training video data and the test video data to obtain training image data and test image data; then, a face extraction operation is performed based on the training image data and the test image data to obtain a training sample set and a test sample set including the real face and the fake face.
[0106] It can be understood that by performing frame extraction processing on the training video data, the training image data can be obtained; and by performing frame extraction processing on the test video data, the test image data can be obtained.
[0107] The face extraction operation refers to extracting facial parts from training image data and test image data.
[0108] As an optional implementation, it is assumed that the preset ratio is 8:2, that is, the video data set is divided into training video data and test video data according to the ratio of 8:2.
[0109] Step S200: Transfer the training sample set to the deep learning network model, and perform model training actions to obtain a deep fake detection model.
[0110] Optionally, after receiving the training sample set, the deep learning network model performs a supervised training operation based on the training sample set to obtain the deepfake detection model. In other words, the deepfake detection network model is supervised trained on the training sample set to obtain a trained deepfake detection model capable of detecting forged images.
[0111] In this embodiment, through supervision signals, such as binary classification labels, the model is forced to establish joint discrimination rules in the spatial domain and frequency domain, thereby realizing forced learning of coordinated spatial-frequency dual-domain features; the labeled data is used to optimize the offset prediction network of the deformable convolution, so that it can accurately focus on the forged area, such as unnatural nose wing deformation, to realize deformable convolution parameter calibration.
[0112] Step S300: Pass the test sample set to the deep fake detection model, and perform face authenticity classification to obtain authenticity face detection results.
[0113] Optionally, after the deep fake detection model receives the test sample set, it performs a detection action on the test sample in the test sample set to determine whether the test sample is a real face; if the test sample is determined to be a real face, the test sample is classified into a set whose true and false face detection result is true; if the test sample is determined not to be a real face, the test sample is classified into a set whose true and false face detection result is false.
[0114] In the technical solution provided in this embodiment, a training sample set and a test sample set consisting of real and fake faces are constructed by performing frame extraction and face extraction on a video dataset. The spatial domain features of the training sample set are extracted using an attention mechanism in collaboration with deformable convolution. The frequency domain features of the training sample set are extracted using a high-frequency extraction and frequency domain learning strategy. The spatial and frequency domain features are fused using bidirectional cross-attention to obtain the space-frequency features of the training sample set. A network is constructed to train the space-frequency features to obtain a deepfake detection model. The test sample set is then detected based on the deepfake detection model. By efficiently fusing and enhancing spatial and frequency domain features, the network can extract richer features, thereby improving detection performance, effectively distinguishing real images from fake images, and enhancing security.
[0115] like Figure 3 As shown, in one embodiment, the public video data set is first trained, tested, and divided proportionally, and frame extraction and face extraction are performed on the training video data and the test video data respectively.
[0116] Then, the training face images and the test face images are downsampled.
[0117] The downsampled feature map is further subjected to spatial feature extraction and frequency domain feature extraction.
[0118] Specifically, see Figure 4 . Spatial domain feature extraction mainly depends on the reasonable design of deformable convolution and attention mechanism. In the spatial domain feature extraction module, in the process of obtaining output from input, a feature map with dimensions C (channel), H (height), and W (width) is calculated in parallel through 1×1 convolution, deformable convolution, and spatial-channel attention. Among them, the output dimensions of 1×1 convolution and deformable convolution are both (C / / 2, H, W). They are spliced along the channel to obtain a feature map with dimensions (C, H, W), and then shuffled to obtain the convolution result feature map. The output dimension of spatial-channel attention is (C, H, W), which is consistent with the dimension of the convolution result feature map, so it can be fused by addition to obtain the output of the spatial domain feature extraction structure.
[0119] See also Figure 5Frequency domain feature extraction mainly depends on the reasonable design of high frequency retention and frequency domain learning. In the frequency domain feature extraction structure, the process of obtaining frequency domain feature output from spatial domain feature input includes taking a spatial domain feature map with data dimensions of C (channel), H (height), and W (width) as input into the frequency domain feature extraction structure, and performing fast Fourier transform on it in the spatial dimension (H×W) and channel dimension C respectively, converting it from the spatial domain to the frequency domain, then performing high frequency retention operation on the input feature map and then performing inverse fast Fourier transform to convert it back to the spatial domain to obtain the spatial representation of the high frequency information of the input feature map, then performing fast Fourier transform on this spatial representation to the frequency domain, performing convolution operation on its imaginary and real parts respectively to learn the frequency domain features of the feature map, then converting the obtained frequency domain features to the spatial domain using inverse fast Fourier transform, and taking the real part to obtain the spatial representation of the frequency domain features for subsequent operations.
[0120] After obtaining the spatial domain features and frequency domain features, the spatial domain features and frequency domain features are fused and enhanced to obtain a richer and more comprehensive feature representation.
[0121] Specifically, see Figure 6 . The spatial domain features and frequency domain features are fused through the bidirectional cross-attention structure, so that the space-frequency features can complement and promote each other. In the bidirectional cross-attention structure, the process of fusing spatial domain features with frequency domain features includes performing convolution and other processing on the spatial domain features and frequency domain features respectively to obtain their query (Q), key (K) and value (V), and then performing attention calculations on the Q of the spatial domain features with the K and V of the frequency domain features, as well as the Q of the frequency domain features with the K and V of the spatial domain features, to obtain the spatial domain features and frequency domain features after cross-enhancement. After that, a weight is assigned to each value of the two, and then the weighted spatial domain features are fused with the frequency domain features to obtain the space-frequency fusion features.
[0122] After obtaining the space-frequency fusion features, classification is performed according to the extracted features to obtain the detection results.
[0123] Then, the structure is extracted based on the space-frequency features to form a deep learning network model.
[0124] Specifically, see Figure 7 The process of distinguishing the authenticity of the input image includes first downsampling the preprocessed input image and performing preliminary feature extraction. Thereafter, feature extraction is continuously performed through the space-frequency feature extraction structure to optimize the feature expression capability of the network. Finally, the features extracted by the network are downsampled and mapped into the category space through the classification head to perform authenticity judgment.
[0125] By training a deep learning network model on a training sample set, a deepfake detection model capable of detecting deepfake images is obtained. The performance of the deepfake detection model is then evaluated on a test sample set. It should be noted that the performance evaluation metrics for the deepfake detection model include, but are not limited to, forgery method, accuracy, precision, recall, and AUC. Forgery methods include, but are not limited to, DeepFakes, Face2Face, FaceShifter, FaceSwap, NeuralTextures, Average, and Celeb-DF.
[0126] The present invention also provides a computer-readable storage medium, which stores a space-frequency joint deep fake detection program based on dual-domain attention collaborative deformable convolution. When the space-frequency joint deep fake detection program based on dual-domain attention collaborative deformable convolution is executed by a processor, it implements the various steps of the space-frequency joint deep fake detection method based on dual-domain attention collaborative deformable convolution as described in the above embodiment.
[0127] The computer-readable storage medium may be any computer-readable storage medium that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disk.
[0128] It should be noted that since the storage medium provided in the embodiments of this application is the storage medium used to implement the method of the embodiments of this application, based on the method described in the embodiments of this application, those skilled in the art will be able to understand the specific structure and deformation of the storage medium, and therefore will not be described in detail here. All storage media used in the method of the embodiments of this application fall within the scope of protection to be provided by this application.
[0129] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROMs, optical storage, etc.) containing computer-usable program code.
[0130] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0131] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0132] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.
[0133] It should be noted that, in the claims, any reference signs placed between parentheses shall not be construed as limiting the claims. The word "comprising" does not exclude the presence of components or steps not listed in the claim. The word "a" or "an" preceding a component does not exclude the presence of a plurality of such components. The invention can be implemented by means of hardware comprising several distinct components and by means of a suitably programmed computer. In a unit claim enumerating several means, several of these means may be embodied by one and the same item of hardware. The use of the words first, second and third etc. does not indicate any order. These words may be interpreted as names.
[0134] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0135] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.
Claims
1. A spatial-frequency joint deepfake detection method based on dual-domain attention collaborative deformable convolution, characterized by: The spatial-frequency joint deepfake detection method based on dual-domain attention collaborative deformable convolution includes the following steps: Obtaining a training sample set and a test sample set, and forming a deep learning network model based on a space-frequency feature extraction structure, wherein the training sample set and the test sample set are composed of real faces and forged faces, and the space-frequency feature extraction structure is composed of a spatial domain feature extraction structure, a frequency domain feature extraction structure, and a bidirectional cross attention structure; Passing the training sample set to the deep learning network model and performing model training actions to obtain a deep fake detection model; Passing the test sample set to the deep fake detection model and performing a face authenticity classification action to obtain a true and false face detection result; The step of forming a deep learning network model based on the space-frequency feature extraction structure includes: Based on deformable convolution and spatial-channel attention mechanism, the spatial domain feature extraction structure is formed; Based on the high-frequency processing algorithm and the frequency domain learning algorithm of the feature map, the frequency domain feature extraction structure is formed; Based on the spatial domain feature extraction structure, the frequency domain feature extraction structure and the bidirectional cross attention structure, forming the space-frequency feature extraction structure; Obtaining the deep learning network model according to the space-frequency feature extraction structure and network layer; The step of forming the spatial domain feature extraction structure based on deformable convolution and space-channel attention mechanism includes: When a feature map is received, the feature map is passed to the spatial-channel attention mechanism to obtain a first spatial domain parameter; and Perform a channel-wise splicing operation on the features in the feature map to obtain a second spatial domain parameter; A spatial domain feature map corresponding to the feature map is obtained based on the first spatial domain parameter and the second spatial domain parameter.
2. The spatial-frequency joint deepfake detection method based on dual-domain attention collaborative deformable convolution as claimed in claim 1 is characterized in that The steps of forming the frequency domain feature extraction structure based on the high frequency processing algorithm and the frequency domain learning algorithm of the feature map include: When a feature map is received, a high-frequency processing action is performed on the feature map based on the high-frequency processing algorithm to obtain a high-frequency processing result corresponding to the feature map; The high-frequency processing result is transmitted to the frequency domain learning algorithm, a frequency domain learning action is performed, and a frequency domain feature map corresponding to the feature map is obtained.
3. The spatial-frequency joint deepfake detection method based on dual-domain attention collaborative deformable convolution as described in claim 2 is characterized in that The step of forming the space-frequency feature extraction structure based on the space-domain feature extraction structure, the frequency-domain feature extraction structure and the bidirectional cross attention structure includes: Transferring the spatial domain feature map and the frequency domain feature map to the bidirectional cross attention structure to obtain the cross attention results from the spatial domain to the frequency domain and the cross attention results from the frequency domain to the spatial domain corresponding to the feature map; The cross-attention results from the spatial domain to the frequency domain and the cross-attention results from the frequency domain to the spatial domain are fused to obtain a target feature map.
4. The spatial-frequency joint deepfake detection method based on dual-domain attention collaborative deformable convolution as claimed in claim 1 is characterized in that The step of transferring the training sample set to the deep learning network model and performing model training to obtain a deep fake detection model includes: After the deep learning network model receives the training sample set, it performs a supervised training operation based on the training sample set to obtain the deep fake detection model.
5. The spatial-frequency joint deepfake detection method based on dual-domain attention collaborative deformable convolution as claimed in claim 1 is characterized in that The step of transferring the test sample set to the deep fake detection model and performing a face authenticity classification operation to obtain a true and false face detection result includes: After receiving the test sample set, the deep fake detection model performs a detection action on the test sample in the test sample set to determine whether the test sample is a real human face; If the test sample is determined to be a real face, classify the test sample into a set of true and false face detection results; If it is determined that the test sample is not a real face, the test sample is classified into a set of true and false face detection results of which are false.
6. The spatial-frequency joint deepfake detection method based on dual-domain attention collaborative deformable convolution as claimed in claim 1 is characterized in that The steps of obtaining a training sample set and a test sample set include: Obtaining a video data set, and dividing the video data set according to a preset ratio to obtain training video data and test video data; Performing a frame extraction operation on the training video data and the test video data to obtain training image data and test image data; A face extraction operation is performed based on the training image data and the test image data to obtain a training sample set and a test sample set including the real faces and the forged faces.
7. An authenticity detection device, characterized in that: The authenticity detection device includes: a memory, a processor, and a space-frequency joint deep fake detection program based on dual-domain attention collaborative deformable convolution stored on the memory and runnable on the processor. The space-frequency joint deep fake detection program based on dual-domain attention collaborative deformable convolution is configured to implement the steps of the space-frequency joint deep fake detection method based on dual-domain attention collaborative deformable convolution as described in any one of claims 1 to 6.
8. A readable storage medium, characterized in that: The readable storage medium stores a space-frequency joint deep fake detection program based on dual-domain attention collaborative deformable convolution. When the space-frequency joint deep fake detection program based on dual-domain attention collaborative deformable convolution is executed by the processor, it implements the steps of the space-frequency joint deep fake detection method based on dual-domain attention collaborative deformable convolution as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Face deep false detection method based on multi-modal feature fusion
CN115880749A
Multi-modal fusion detection method for deeply-forged audio and video
CN116797896A