A general deepfake detection method and device based on space-frequency cooperative learning
By extracting features from both the spatial and frequency domains simultaneously using a spatial-frequency collaborative learning framework, and combining it with a hierarchical cross-modal fusion mechanism, the problem of insufficient accuracy and generalization ability in high-quality forgery detection in existing technologies is solved, achieving efficient and accurate deep forgery detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NORTH CHINA UNIVERSITY OF TECHNOLOGY
- Filing Date
- 2026-02-01
- Publication Date
- 2026-06-02
AI Technical Summary
Existing deepfake detection methods suffer from performance degradation when faced with high-quality forged content or post-processing interference, insufficient frequency feature mining, neglect of dependencies between frequency features, and inefficient fusion of spatial and frequency domain features.
A spatial-frequency collaborative learning approach is adopted, which performs joint feature extraction through a spatial-frequency collaborative learning framework and combines a hierarchical cross-modal fusion mechanism to extract features from both the spatial and frequency domains. Deep forgery detection is achieved through multi-stage feature interaction.
It achieves efficient and universal deepfake detection, accurately captures frequency domain artifacts, preserves local spectral features, reduces high-frequency degradation caused by resampling, and improves detection accuracy and generalization ability.
Smart Images

Figure CN122135072A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of information security technology, and in particular to a general-purpose deepfake detection method and device based on spatial-frequency collaborative learning. Background Technology
[0002] With the rapid development of deep generative models, deepfake technology continues to evolve, enabling the creation of highly realistic deepfake content. While these technologies have potential applications in entertainment and creative industries, their misuse can lead to serious social risks such as the spread of misinformation and identity theft. Therefore, deepfake detection has become a key challenge in multimedia security and computer vision.
[0003] Existing deepfake detection methods mostly transform the task into a binary classification problem, relying primarily on backbone networks to extract discriminative features, with analysis and detection focusing on the spatial domain. They distinguish between real and fake content by identifying spatial artifacts such as inconsistencies, unnatural textures, and fused boundaries in forged images. Some methods also incorporate dedicated functional modules or combine noise features with RGB features to improve the learning effect of local forgery patterns.
[0004] As deepfake technology continues to advance, spatial details are constantly optimized, making the pixel-level differences between real and fake images extremely subtle and difficult to detect. This introduces an inherent flaw into methods that rely on spatial domain analysis; their detection performance significantly degrades and their generalization ability is insufficient when faced with high-quality forged content or content that has undergone post-processing interference.
[0005] Some existing methods recognize the complementary value of frequency domain analysis and extract frequency features through techniques such as Discrete Cosine Transform (DCT) and Discrete Fourier Transform (DFT), or combine frequency domain and RGB features for detection. These methods typically enhance spatial anomalies in the spectral representation through intermediate filtering operations, and finally complete feature extraction in the spatial domain. Some methods also employ block-level frequency transformation.
[0006] Most existing frequency-aware methods only use frequency analysis as an auxiliary means of spatial domain detection, lacking in-depth exploration of the inherent abnormal patterns in frequency subbands, especially the frequency features of local forgery regions. Most methods have not fully explored the interdependencies between frequency features, nor have they properly considered the correlation between image semantics and forgery artifacts, resulting in limited generalization ability on different datasets and different forgery types.
[0007] In the existing technology, there is a lack of an efficient and accurate general deepfake detection method based on spatial frequency collaborative learning. Summary of the Invention
[0008] To address the performance degradation of existing technologies when faced with high-quality forged content or post-processing interference, insufficient frequency feature mining, neglect of dependencies between frequency features, and inefficient fusion of spatial and frequency domain features, this invention provides a general-purpose deepfake detection method and apparatus based on spatial-frequency collaborative learning. The technical solution is as follows:
[0009] On the one hand, a general-purpose deepfake detection method based on space-frequency collaborative learning is provided. This method is implemented by a general-purpose deepfake detection device and includes: Acquire the image to be detected; perform image recognition using a preset target detection method based on the image to be detected to obtain the first target image information; perform scale unification processing on the first target image information to obtain the second target image information; Based on the first target image information and the second target image information, a joint feature extraction is performed using a space-frequency collaborative learning framework to obtain space-frequency collaborative features. Based on the spatial-frequency co-operational features, a classifier is used to perform deep forgery detection and obtain the image forgery judgment result.
[0010] On the other hand, a general-purpose deepfake detection device based on spatio-frequency collaborative learning is provided. This device is applied to a general-purpose deepfake detection method based on spatio-frequency collaborative learning, and the device includes: The image information acquisition module is used to acquire the image to be detected; based on the image to be detected, it performs image recognition using a preset target detection method to obtain first target image information; and performs scale unification processing on the first target image information to obtain second target image information. The space-frequency collaborative learning module is used to perform joint feature extraction based on the first target image information and the second target image information using the space-frequency collaborative learning framework to obtain space-frequency collaborative features. The image forgery identification module is used to perform deep forgery identification using a classifier based on spatial-frequency co-location features, and obtain the image forgery judgment result.
[0011] On the other hand, a general-purpose deepfake detection device is provided, comprising: a processor; and a memory storing computer-readable instructions, wherein when the computer-readable instructions are executed by the processor, any one of the general-purpose deepfake detection methods based on space-frequency collaborative learning described above is implemented.
[0012] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction is stored therein, the at least one instruction being loaded and executed by a processor to implement any of the above-described general deepfake detection methods based on space-frequency cooperative learning.
[0013] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: This invention proposes a general deepfake detection method based on spatial-frequency collaborative learning. By simultaneously extracting features from the spatial and frequency domains and leveraging a hierarchical cross-modal fusion mechanism to achieve multi-stage feature interaction, it enables efficient and universal deepfake detection. The proposed space-frequency collaborative learning framework can accurately capture frequency domain artifacts, analyze and preserve local spectral features, and avoid the loss of details in global transformation; it effectively preserves scale-invariant spectral features and reduces high-frequency degradation caused by resampling. The hierarchical cross-modal fusion mechanism fully leverages the complementarity of spatial and frequency domain features. The combination of shallow attention enhancement and deep dynamic modulation effectively models the spatial-frequency interaction relationship, resolving the issue of feature heterogeneity between modalities. This invention significantly outperforms existing technologies in detection accuracy and generalization ability, adapting to different types and qualities of deepfake content. This invention is a highly efficient and accurate general-purpose deepfake detection method based on spatial-frequency collaborative learning. Attached Figure Description
[0014] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0015] Figure 1 This is a flowchart of a general deepfake detection method based on space-frequency collaborative learning provided by an embodiment of the present invention; Figure 2 This is a block diagram of a general-purpose deepfake detection device based on space-frequency collaborative learning provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of a general-purpose deepfake detection device provided in an embodiment of the present invention. Detailed Implementation
[0016] The technical solution of the present invention will now be described with reference to the accompanying drawings.
[0017] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.
[0018] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning.
[0019] In this embodiment of the invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.
[0020] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.
[0021] This invention provides a general-purpose deepfake detection method based on space-frequency collaborative learning. This method can be implemented using a general-purpose deepfake detection device, which can be a terminal or a server. Figure 1 The flowchart shown is for a general deepfake detection method based on space-frequency collaborative learning. The processing flow of this method may include the following steps:
[0022] S1. Obtain the image to be detected; based on the image to be detected, perform image recognition using a preset target detection method to obtain first target image information; perform scale unification processing on the first target image information to obtain second target image information; In one feasible implementation, the present invention is applied to the fields of multimedia security and computer vision, taking the spoofing of facial images as an example.
[0023] The input RGB image is converted to the YCbCr color space, and the corresponding preset face tracking algorithm is used to locate the face region in the image to obtain the face boundary coordinates. According to the size requirements of subsequent modules, the face region is initially adapted to the size. After the detected face region is cropped, it is uniformly scaled to 380×380 resolution. The original resolution of the face region is preserved and cropped.
[0024] S2. Based on the first target image information and the second target image information, use the space-frequency collaborative learning framework to perform joint feature extraction and obtain space-frequency collaborative features; The space-frequency collaborative learning framework includes spatial pipeline branches, local frequency pipeline branches, and global frequency pipeline branches. The spatial pipeline branches include an efficient convolutional backbone network, a frequency-aware attention enhancement module, and a hybrid cross-modal attention fusion module; The local branches of the frequency pipeline include a discrete cosine transform module, a spectral band convolution module, and a frequency backbone network; the frequency backbone network is built based on the improved Xception framework. The global branch of the frequency pipeline includes a derivative feature extraction module, a feature statistics module, and a feature fusion module.
[0025] In one feasible implementation, the present invention proposes a general deep forgery detection framework based on Spatial-Frequency Collaborative Learning (SFCL); the Spatial-Frequency Collaborative Learning framework includes spatial pipeline branches, frequency pipeline local branches, and frequency pipeline global branches.
[0026] Optionally, based on the first target image information and the second target image information, a space-frequency collaborative learning framework is used to perform joint feature extraction to obtain space-frequency collaborative features, including: Based on the first target image information, global frequency statistical coding is performed through the global branch of the frequency pipeline to obtain global frequency features; Based on the second target image information, cross-band feature refinement is performed through local branches of the frequency pipeline to obtain local frequency features; Local frequency features are injected into the spatial branch; based on the hierarchical cross-modal fusion mechanism, the spatial features are refined by the spatial branch according to the information of the second target image and the local frequency features to obtain the spatial features. Based on multi-head attention mechanism and dynamic gating mechanism, collaborative modulation fusion is performed according to spatial features, global frequency features and local frequency features to obtain spatial-frequency collaborative features.
[0027] In one feasible implementation, the general-purpose deepfake detection framework of the present invention extracts features from both the spatial and frequency domains simultaneously and achieves multi-stage feature interaction through a hierarchical cross-modal fusion mechanism, thereby achieving efficient and universal deepfake detection.
[0028] Optionally, based on the first target image information, global frequency statistical coding is performed through a global branch of the frequency pipeline to obtain global frequency features, including: Based on the inter-block row derivative, inter-block column derivative, and intra-block derivative, scale-invariant derivative features are calculated using the information of the first target image. Based on the mean, standard deviation, skewness, and kurtosis, global statistical characteristics are obtained by performing global statistical calculations according to derivative characteristics. Flattened fusion coding is performed based on global statistical features to obtain global frequency features.
[0029] In one feasible implementation, this invention introduces Scale-Invariant Derivative Analysis (SIDA) into the global branch of the frequency pipeline, preserving the block-based DCT transform while maintaining the original resolution, and simultaneously maintaining the shape consistency of spectral features between images with different spatial dimensions. For inter-block analysis, local tampering often leads to abrupt changes in frequency distribution between adjacent blocks.
[0030] To capture transition inconsistencies, this invention calculates horizontal and vertical derivatives to eliminate semantic information while amplifying tamper-specific anomalies. Inter-block row derivatives are calculated from the reconstructed DCT coefficients X˜. and inter-block derivatives , in, Let C represent the set of real numbers; C represents the number of channels. The input RGB image is first converted to the YCbCr color space, so C is usually 3(Y, Cb, Cr); H represents the image height; W represents the image width. The above calculation process is shown in formulas (1) and (2):
[0031] (1); (2); Where m: represents the block index in the row direction, with a value range of [1, (H / 8)-1][1, (H / 8)-1]; n: represents the block index in the column direction, with a value range of [1, (W / 8)-1][1, (W / 8)-1]; : the colon in the tensor index indicates that all elements in that dimension are selected.
[0032] Intra-block analysis reveals that forged regions often exhibit anomalous spectral energy distributions. Therefore, intra-block derivatives are calculated, followed by zero-padding along the intra-block dimensions to maintain dimensional consistency, as shown in Equation 3:
[0033] (3); Where l: represents the index of the DCT coefficients within the block, and its value range is [1,63][1,63].
[0034] Four statistical indicators are calculated from the absolute values of the three derivative characteristic plots: mean, standard deviation, skewness, and kurtosis. The calculation formulas are shown in equations (4), (5), (6), and (7) below: (4); (5); (6); (7); Where k is the index, representing three different types of derivatives: row: Corresponding inline derivative (inter-block row derivatives) are used to capture horizontal transition inconsistencies. col: Corresponding column derivative (inter-block column derivatives) are used to capture transition inconsistencies in the vertical direction; intra: corresponding to the block derivative (intra-block derivatives) are used to capture anomalies in the spectral energy distribution within a block; The index :,:,i,j: indicates the selection of all channels (first colon), all frequency components (second colon), and the image patch located in row i and column j.
[0035] This aggregation process transforms local difference intensities into global statistical descriptors, thereby capturing the overall distribution pattern of forgery traces rather than relying on artifacts at specific locations. Four types of derivative statistical features are flattened and concatenated to obtain the frequency global feature D∈R2304, which effectively encodes horizontal, vertical, and intra-block anomaly patterns, calculated as shown in Equation (8):
[0036] (8); Here, `cat` represents the concatenation operation, which connects multiple vectors or tensors along a specified dimension; `flat` represents the flatten operation, which converts a multidimensional tensor into a one-dimensional vector.
[0037] Optionally, based on the second target image information, cross-band feature refinement is performed through local branches of the frequency pipeline to obtain local frequency features, including: The second target image information is transformed into the spatial-frequency domain to obtain the DCT coefficient matrix; Cross-spectral band correlation is captured based on the DCT coefficient matrix to obtain spectral-semantic features; High-dimensional frequency features are refined based on spectral-semantic features to obtain local frequency features.
[0038] In one feasible implementation, this invention proposes a Spectral Band Convolution Module (SBCM), which employs stacked 3D convolutional blocks with kernel sizes of (x, 1, 1) to capture inter-channel correlations between spectral coefficients. By sliding the 3D convolutional kernel along the spectral band dimension, the SBCM progressively learns local interaction patterns between adjacent frequency bands. The initial layer utilizes a larger kernel to establish long-range dependencies between low-frequency, mid-frequency, and high-frequency components for cross-band feature aggregation; the intermediate layers gradually decrease x to focus on abrupt spectral transitions caused by artifacts, thereby enhancing the anomalous response in the mid-frequency region; the final layer uses the smallest x to isolate high-frequency noise patterns and amplify distribution irregularities within key spectral bands. In the SBCM module, the size of the 3D convolutional kernel can be adjusted according to the actual application scenario and does not necessarily have to strictly follow the 7, 5, 3 order, as long as cross-band feature interaction and artifact capture can be achieved.
[0039] The processed spectral-semantic features are represented as Xflocal∈RC×3×(H / 8)×(W / 8), where C = 64, thus encapsulating a hierarchically refined frequency-specific representation. The first two dimensions of the feature tensor are flattened, which differs from the original direct flattening method, in which the 192-dimensional features at each spatial location now contain high-level semantic features with cross-band correlations. The calculation process is shown in Equation (9):
[0040] (9); The frequency backbone network is built upon an improved Xception framework. The first two convolutional layers and one depthwise separable convolutional block of the classic Xception framework are removed and replaced with SBCM. To maintain architectural compatibility, the input channels of subsequent depthwise separable convolutional blocks are adjusted to match the channel dimensions of the output spectral features Xflocal. This hierarchical processing framework can model intra-block and inter-block correlations between DCT coefficients, thereby generating frequency local features F∈R2048.
[0041] Optionally, based on a hierarchical cross-modal fusion mechanism, according to the second target image information and frequency local features, frequency-aware spatial features are refined through spatial branching to obtain spatial features, including: Multi-scale feature extraction is performed based on the information from the second target image to obtain shallow texture features and deep semantic features; Based on the attention mechanism, frequency-aware attention enhancement processing is performed according to shallow texture features and local frequency features to obtain frequency-spatial joint enhancement features. Spatial features are obtained by cross-layer feature fusion based on frequency-spatial joint enhancement features and deep semantic features.
[0042] In one feasible implementation, this invention constructs a Backbone-S based on the efficient convolutional neural network EfficientNet. It employs compound scaling to coordinate and optimize depth, width, and resolution for efficient feature extraction. Since forgery traces exhibit greater prominence in shallow features, the shallow fusion stage is deployed in early backbone layers. A Frequency-Aware Attention Enhancement Module (FAAE) enhances forgery trace features in shallow features through frequency-localized attention maps. A Hierarchical Cross-Modal Fusion (HCMF) mechanism is used to fuse spatial and frequency domain features. This amplifies localized forgery traces in the shallow layers of the spatial domain model, thereby enhancing the saliency of subtle tampering clues and promoting the learning of discriminative feature representations.
[0043] The shallow texture features XS and the frequency local features Xflocal output by SBCM are used. Cross-modal query-key pairs are generated by parallel convolutional projection, while aligning the channel dimensions to ensure compatibility, as shown in equations (10) and (11):
[0044] (10); (11); in, () indicates the operation of convolution projection on frequency domain features; () indicates the operation of convolution projection on spatial domain features; () indicates the operation of convolution projection on frequency domain features; () represents the operation of convolution projection on spatial domain features; T represents the transpose operation, which is to interchange the rows and columns of the matrix.
[0045] The cross-modal attention weight matrix α is calculated by scaling the dot product between the connected key pairs. It encapsulates the spectral spatial interaction dynamics by associating local frequency modes with spatial anomalies, as shown in Equation (12): (12); Frequency context constraints are injected through attention gating. ConvF represents the projection of spectral domain features to align with the spatial feature dimension; γS represents a learnable parameter that dynamically adjusts the contribution strength of spectral features to the spatial branch; σ represents the Sigmoid activation function; and ConvS represents the dimensional adjustment of the frequency-aware attention matrix to ensure dimensional compatibility during cross-modal feature fusion, as shown in Equation (13). (13); in, This is a batch normalization operation. Element-wise multiplication, also known as the Hadamard product, refers to multiplying the corresponding elements of two matrices, vectors, or tensors of the same shape to obtain a new matrix, vector, or tensor with the same shape as the original operands.
[0046] Optionally, based on a multi-head attention mechanism and a dynamic gating mechanism, collaborative modulation fusion is performed according to spatial features, global frequency features, and local frequency features to obtain spatial-frequency collaborative features, including: Embedded spatial mapping is performed based on spatial features and local frequency features to obtain mapped spatial features and mapped global frequency features of a unified dimension. Based on the multi-head attention mechanism, residual connections are introduced, and weighted fusion is performed according to the mapped spatial features and the mapped frequency global features to obtain cross-modal fusion features. Based on a dynamic gating mechanism, cross-modal fusion features are adaptively modulated according to global frequency characteristics to obtain space-frequency cooperative features.
[0047] In one feasible implementation, to establish in-depth interactive reasoning between the spatial and spectral domains, this invention proposes a Hybrid Cross-Modal Attention Fusion Module (HCMA) for efficient fusion of spatial-frequency features at a deep level. This module first projects the spatial depth features S∈R1792 and the intra-block / inter-block frequency-related features F∈R2048 onto a unified embedding space S′, F′∈R1792, through a linear transformation layer, thereby eliminating modality-specific dimensional differences.
[0048] Spatial features are used as the query vector Q = S′WQ, and frequency features are used as key-value pairs K = F′WK,V = F′WV, where WQ, WK, WV ∈ R1024×1024 is a learnable parameter matrix. Multi-head attention is used to capture cross-modal global dependencies to generate attention-weighted representations, as shown in Equation (14):
[0049] (14); Where h is the number of attention heads; the multi-head attention mechanism divides the input Query, Key, and Value into multiple subspaces for independent attention calculation and then concatenates the results.
[0050] To improve fusion performance, residual connections are introduced. Gradient stability is maintained by adding the attention output to the spatial features processed by 1D convolution and performing batch normalization, as shown in Equation (15): (15); in, It is a 1×1 convolutional layer; in this invention, the convolutional layer is used to adjust the number of channels of the feature map without changing its spatial dimension; in this formula (15), it is used to process the spatial feature S'.
[0051] Considering the sensitivity of derivative statistical features to tampered regions, this invention employs a dynamic gating mechanism. By using Sigmoid activation, the derivative frequency statistical representation D∈R2304 is projected onto the gating coefficients, thereby adjusting the features to amplify the abnormal response, as shown in formulas (16) and (17): (16); (17); in, It is a learnable weight matrix that is responsible for linearly transforming the high-dimensional input features to the dimension required for the gating coefficients; The global frequency features are obtained by calculating the row, column, and block derivatives of the DCT-transformed image and aggregating its statistical characteristics (mean, standard deviation, skewness, kurtosis). These are the gating coefficients; this vector is used to modulate (adjust) other features to amplify the anomalous response.
[0052] The spatial domain backbone network can be replaced with other high-performance convolutional neural networks, such as ResNet and variants of VisionTransformer. As long as it can effectively extract local artifact features in the spatial domain, it can be used in conjunction with the frequency domain processing and fusion mechanism of this invention to achieve the detection objective.
[0053] In the hierarchical cross-modal fusion mechanism, the number of heads in the multi-head attention can be adjusted according to the computational resources and detection performance requirements. The activation function of the dynamic gating mechanism can also be replaced with other suitable activation functions such as ReLU without affecting the core feature fusion effect.
[0054] S3. Based on the spatial-frequency co-operational features, a classifier is used to perform deep forgery discrimination to obtain the image forgery judgment result.
[0055] In one feasible implementation, the present invention proposes an architecture that achieves hierarchical cross-modal feature collaboration, combining global dependency modeling with local anomaly enhancement capabilities. The fused multimodal semantic representation vectors are input into a classifier to perform discriminative decision-making for highly realistic forged media content detection tasks.
[0056] This invention proposes a general deepfake detection method based on spatial-frequency collaborative learning. By simultaneously extracting features from the spatial and frequency domains and leveraging a hierarchical cross-modal fusion mechanism to achieve multi-stage feature interaction, it enables efficient and universal deepfake detection. The proposed space-frequency collaborative learning framework can accurately capture frequency domain artifacts, analyze and preserve local spectral features, and avoid the loss of details in global transformation; it effectively preserves scale-invariant spectral features and reduces high-frequency degradation caused by resampling. The hierarchical cross-modal fusion mechanism fully leverages the complementarity of spatial and frequency domain features. The combination of shallow attention enhancement and deep dynamic modulation effectively models the spatial-frequency interaction relationship, resolving the issue of feature heterogeneity between modalities. This invention significantly outperforms existing technologies in detection accuracy and generalization ability, adapting to different types and qualities of deepfake content. This invention is a highly efficient and accurate general-purpose deepfake detection method based on spatial-frequency collaborative learning.
[0057] Figure 2 This is a block diagram of a general-purpose deepfake detection device based on spatio-frequency collaborative learning, provided by an embodiment of the present invention. This device is used for a general-purpose deepfake detection method based on spatio-frequency collaborative learning. (Refer to...) Figure 2 The device includes an image information acquisition module 210, a space-frequency collaborative learning module 220, and an image forgery identification module 230. Among them:
[0058] The image information acquisition module 210 is used to acquire the image to be detected; perform image recognition using a preset target detection method based on the image to be detected to obtain first target image information; and perform scale unification processing on the first target image information to obtain second target image information. The space-frequency collaborative learning module 220 is used to perform joint feature extraction based on the first target image information and the second target image information using the space-frequency collaborative learning framework to obtain space-frequency collaborative features; The image forgery identification module 230 is used to perform deep forgery identification using a classifier based on the spatial-frequency co-location features, and obtain the image forgery identification result.
[0059] The space-frequency collaborative learning framework includes spatial pipeline branches, local frequency pipeline branches, and global frequency pipeline branches. The spatial pipeline branches include an efficient convolutional backbone network, a frequency-aware attention enhancement module, and a hybrid cross-modal attention fusion module; The local branches of the frequency pipeline include a discrete cosine transform module, a spectral band convolution module, and a frequency backbone network; the frequency backbone network is built based on the improved Xception framework. The global branch of the frequency pipeline includes a derivative feature extraction module, a feature statistics module, and a feature fusion module.
[0060] Optionally, the space-frequency collaborative learning module 220 is further used for: Based on the first target image information, global frequency statistical coding is performed through the global branch of the frequency pipeline to obtain global frequency features; Based on the second target image information, cross-band feature refinement is performed through local branches of the frequency pipeline to obtain local frequency features; Local frequency features are injected into the spatial branch; based on the hierarchical cross-modal fusion mechanism, the spatial features are refined by the spatial branch according to the information of the second target image and the local frequency features to obtain the spatial features. Based on multi-head attention mechanism and dynamic gating mechanism, collaborative modulation fusion is performed according to spatial features, global frequency features and local frequency features to obtain spatial-frequency collaborative features.
[0061] Optionally, the space-frequency collaborative learning module 220 is further used for: Based on the inter-block row derivative, inter-block column derivative, and intra-block derivative, scale-invariant derivative features are calculated using the information of the first target image. Based on the mean, standard deviation, skewness, and kurtosis, global statistical characteristics are obtained by performing global statistical calculations according to derivative characteristics. Flattened fusion coding is performed based on global statistical features to obtain global frequency features.
[0062] Optionally, the space-frequency collaborative learning module 220 is further used for: The second target image information is transformed into the spatial-frequency domain to obtain the DCT coefficient matrix; Cross-spectral band correlation is captured based on the DCT coefficient matrix to obtain spectral-semantic features; High-dimensional frequency features are refined based on spectral-semantic features to obtain local frequency features.
[0063] Optionally, the space-frequency collaborative learning module 220 is further used for: Multi-scale feature extraction is performed based on the information from the second target image to obtain shallow texture features and deep semantic features; Based on the attention mechanism, frequency-aware attention enhancement processing is performed according to shallow texture features and local frequency features to obtain frequency-spatial joint enhancement features. Spatial features are obtained by cross-layer feature fusion based on frequency-spatial joint enhancement features and deep semantic features.
[0064] Optionally, the space-frequency collaborative learning module 220 is further used for: Embedded spatial mapping is performed based on spatial features and local frequency features to obtain mapped spatial features and mapped global frequency features of a unified dimension. Based on the multi-head attention mechanism, residual connections are introduced, and weighted fusion is performed according to the mapped spatial features and the mapped frequency global features to obtain cross-modal fusion features. Based on a dynamic gating mechanism, cross-modal fusion features are adaptively modulated according to global frequency characteristics to obtain space-frequency cooperative features.
[0065] This invention proposes a general deepfake detection method based on spatial-frequency collaborative learning. By simultaneously extracting features from the spatial and frequency domains and leveraging a hierarchical cross-modal fusion mechanism to achieve multi-stage feature interaction, it enables efficient and universal deepfake detection. The proposed space-frequency collaborative learning framework can accurately capture frequency domain artifacts, analyze and preserve local spectral features, and avoid the loss of details in global transformation; it effectively preserves scale-invariant spectral features and reduces high-frequency degradation caused by resampling. The hierarchical cross-modal fusion mechanism fully leverages the complementarity of spatial and frequency domain features. The combination of shallow attention enhancement and deep dynamic modulation effectively models the spatial-frequency interaction relationship, resolving the issue of feature heterogeneity between modalities. This invention significantly outperforms existing technologies in detection accuracy and generalization ability, adapting to different types and qualities of deepfake content. This invention is a highly efficient and accurate general-purpose deepfake detection method based on spatial-frequency collaborative learning.
[0066] Figure 3 This is a schematic diagram of the structure of a general-purpose deepfake detection device provided in an embodiment of the present invention, as shown below. Figure 3 As shown, a general-purpose deepfake detection device may include the above-mentioned Figure 3 The illustrated general-purpose deepfake detection device is based on space-frequency collaborative learning. Optionally, the general-purpose deepfake detection device 310 may include a first processor 2001.
[0067] Optionally, the general-purpose deepfake detection device 310 may also include a memory 2002 and a transceiver 2003.
[0068] The first processor 2001, memory 2002, and transceiver 2003 can be connected via a communication bus.
[0069] The following is combined Figure 3A detailed introduction to each component of the general-purpose deepfake detection device 310: The first processor 2001 is the control center of the general-purpose deepfake detection device 310. It can be a single processor or a collective term for multiple processing elements. For example, the first processor 2001 can be one or more central processing units (CPUs), application-specific integrated circuits (ASICs), or one or more integrated circuits configured to implement embodiments of the present invention, such as one or more digital signal processors (DSPs), or one or more field-programmable gate arrays (FPGAs).
[0070] Optionally, the first processor 2001 can perform various functions of the general-purpose deepfake detection device 310 by running or executing software programs stored in the memory 2002 and calling data stored in the memory 2002.
[0071] In a specific implementation, as one example, the first processor 2001 may include one or more CPUs, for example... Figure 3 CPU0 and CPU1 are shown in the diagram.
[0072] In a specific implementation, as one example, the general-purpose deepfake detection device 310 may also include multiple processors, for example... Figure 3 The first processor 2001 and the second processor 2004 are shown in the diagram. Each of these processors can be a single-core processor or a multi-core processor. Here, a processor can refer to one or more devices, circuits, and / or processing cores used to process data (such as computer program instructions).
[0073] The memory 2002 is used to store the software program that executes the present invention, and is controlled by the first processor 2001 to execute it. The specific implementation method can be referred to the above method embodiment, and will not be repeated here.
[0074] Optionally, the memory 2002 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto. The memory 2002 may be integrated with the first processor 2001 or may exist independently, and may be connected via the interface circuit of the general-purpose deepfake detection device 310. Figure 3 (Not shown in the image) is coupled to the first processor 2001, and this embodiment of the invention does not specifically limit this.
[0075] The transceiver 2003 is used to communicate with network devices or with terminal devices.
[0076] Alternatively, transceiver 2003 may include a receiver and a transmitter. Figure 3 (Not shown separately). The receiver is used to implement the receiving function, and the transmitter is used to implement the transmitting function.
[0077] Optionally, the transceiver 2003 can be integrated with the first processor 2001, or it can exist independently and be connected to the interface circuit of the general-purpose deepfake detection device 310. Figure 3 (Not shown in the image) is coupled to the first processor 2001, and this embodiment of the invention does not specifically limit this.
[0078] It should be noted that, Figure 3 The structure of the general-purpose deepfake detection device 310 shown does not constitute a limitation on the router. Actual general-purpose deepfake detection devices may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0079] Furthermore, the technical effect of the general-purpose deepfake detection device 310 can be referred to the technical effect of the general-purpose deepfake detection method based on space-frequency collaborative learning described in the above method embodiments, and will not be repeated here.
[0080] It should be understood that the first processor 2001 in the embodiments of the present invention may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or it may be any conventional processor, etc.
[0081] It should also be understood that the memory in the embodiments of the present invention can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous linked DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0082] The above embodiments can be implemented, in whole or in part, by software, hardware (such as circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.
[0083] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.
[0084] In this invention, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of a single item or a plurality of items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be a single item or multiple items.
[0085] It should be understood that, in various embodiments of the present invention, the order of the above-mentioned process numbers does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0086] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0087] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices, apparatuses, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0088] In the several embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0089] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0090] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0091] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0092] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A general deepfake detection method based on spatio-frequency collaborative learning, characterized in that, The method includes: Acquire the image to be detected; perform image recognition using a preset target detection method based on the image to be detected to obtain the first target image information; perform scale unification processing on the first target image information to obtain the second target image information; Based on the first target image information and the second target image information, a joint feature extraction is performed using a space-frequency collaborative learning framework to obtain space-frequency collaborative features. Based on the spatial-frequency co-operational features, a classifier is used to perform deep forgery detection and obtain the image forgery judgment result.
2. The general deepfake detection method based on spatio-frequency collaborative learning according to claim 1, characterized in that, The space-frequency collaborative learning framework includes a space pipeline branch, a frequency pipeline local branch, and a frequency pipeline global branch. The spatial pipeline branch includes an efficient convolutional backbone network, a frequency-aware attention enhancement module, and a hybrid cross-modal attention fusion module; The local branches of the frequency pipeline include a discrete cosine transform module, a spectral band convolution module, and a frequency backbone network; the frequency backbone network is built based on the improved Xception framework. The global branch of the frequency pipeline includes a derivative feature extraction module, a feature statistics module, and a feature fusion module.
3. The general deepfake detection method based on spatio-frequency collaborative learning according to claim 2, characterized in that, The step of extracting joint features using a space-frequency collaborative learning framework based on the first target image information and the second target image information to obtain space-frequency collaborative features includes: Based on the first target image information, global frequency statistical coding is performed through the global branch of the frequency pipeline to obtain global frequency features; Based on the second target image information, cross-band feature refinement is performed through local branches of the frequency pipeline to obtain local frequency features; Local frequency features are injected into the spatial branch; based on the hierarchical cross-modal fusion mechanism, the spatial features are refined by the spatial branch according to the information of the second target image and the local frequency features to obtain the spatial features. Based on multi-head attention mechanism and dynamic gating mechanism, collaborative modulation fusion is performed according to spatial features, global frequency features and local frequency features to obtain spatial-frequency collaborative features.
4. The general deepfake detection method based on spatio-frequency collaborative learning according to claim 3, characterized in that, The step of obtaining global frequency features by performing global frequency statistical coding through a global branch of a frequency pipeline based on the first target image information includes: Based on the inter-block row derivative, inter-block column derivative, and intra-block derivative, scale-invariant derivative features are calculated using the information of the first target image. Based on the mean, standard deviation, skewness, and kurtosis, global statistical characteristics are obtained by performing global statistical calculations according to derivative characteristics. Flattened fusion coding is performed based on global statistical features to obtain global frequency features.
5. The general deepfake detection method based on spatio-frequency collaborative learning according to claim 3, characterized in that, The step of refining cross-band features through local branches of the frequency pipeline based on the second target image information to obtain local frequency features includes: The second target image information is transformed into the spatial-frequency domain to obtain the DCT coefficient matrix; Cross-spectral band correlation is captured based on the DCT coefficient matrix to obtain spectral-semantic features; High-dimensional frequency features are refined based on spectral-semantic features to obtain local frequency features.
6. The general deepfake detection method based on spatio-frequency collaborative learning according to claim 3, characterized in that, The hierarchical cross-modal fusion mechanism, based on the second target image information and local frequency features, refines frequency-aware spatial features through spatial branching to obtain spatial features, including: Multi-scale feature extraction is performed based on the information from the second target image to obtain shallow texture features and deep semantic features; Based on the attention mechanism, frequency-aware attention enhancement processing is performed according to shallow texture features and local frequency features to obtain frequency-spatial joint enhancement features. Spatial features are obtained by cross-layer feature fusion based on frequency-spatial joint enhancement features and deep semantic features.
7. The general deepfake detection method based on spatio-frequency collaborative learning according to claim 3, characterized in that, The method, based on multi-head attention and dynamic gating mechanisms, performs collaborative modulation fusion according to spatial features, global frequency features, and local frequency features to obtain spatial-frequency collaborative features, including: Embedded spatial mapping is performed based on spatial features and local frequency features to obtain mapped spatial features and mapped global frequency features of a unified dimension. Based on the multi-head attention mechanism, residual connections are introduced, and weighted fusion is performed according to the mapped spatial features and the mapped frequency global features to obtain cross-modal fusion features. Based on a dynamic gating mechanism, cross-modal fusion features are adaptively modulated according to global frequency characteristics to obtain space-frequency cooperative features.
8. A general-purpose deepfake detection device based on spatio-frequency collaborative learning, wherein the general-purpose deepfake detection device based on spatio-frequency collaborative learning is used to implement the general-purpose deepfake detection method based on spatio-frequency collaborative learning as described in any one of claims 1-7, characterized in that, The device includes: The image information acquisition module is used to acquire the image to be detected; based on the image to be detected, it performs image recognition using a preset target detection method to obtain first target image information; and performs scale unification processing on the first target image information to obtain second target image information. The space-frequency collaborative learning module is used to perform joint feature extraction based on the first target image information and the second target image information using the space-frequency collaborative learning framework to obtain space-frequency collaborative features. The image forgery identification module is used to perform deep forgery identification using a classifier based on spatial-frequency co-location features, and obtain the image forgery judgment result.
9. A universal deepfake detection device, characterized in that, The general-purpose deepfake detection device includes: processor; A memory storing computer-readable instructions that, when executed by the processor, implement the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium contains program code that can be invoked by a processor to execute the method as described in any one of claims 1 to 7.