A method and system for identifying deepfake videos

CN121640168BActive Publication Date: 2026-09-01HARBIN INST OF TECH AT WEIHAI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511838617.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-08
Publication Date
2026-09-01
Estimated Expiration
2045-12-08

AI Technical Summary

Technical Problem

[0003]当前主流的深度伪造检测方法普遍面临两大核心挑战:首先是泛化能力不足的问题,伪造技术本身在快速迭代演进,不断涌现的新型伪造手段使得检测模型在面对“未知”或“未见过的”伪造类型时,性能显著下降,难以保持稳健的鉴别能力;其次是现有方法大多依赖于简单的二元(真/伪)分类学习范式,这种范式存在固有缺陷:它将本质上连续、渐变的伪造强度(如融合程度、编辑痕迹的明显程度)强行离散化为非黑即白的标签,导致了“信息坍缩”,具体而言,模型在学习过程中丢失了关于伪造程度细粒度差异的信息,这不仅限制了其判别精度,更容易导致模型在训练数据上过拟合,从而进一步损害其在真实复杂场景下的泛化性能

Benefits of technology

1、通过创新的实例重分级策略,将伪造程度量化为连续的监督信号,使模型能够学习伪造强度的细粒度差异,而非进行简单的真伪二分。这有效避免了因信息坍缩导致的特征表征能力退化与过拟合问题,使模型在面对未知或新型伪造技术时,仍能保持稳定且准确的检测性能。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121640168B_ABST
    Figure CN121640168B_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for identifying deepfake videos, primarily relating to the field of deepfake video detection technology. The method includes: acquiring a real image and applying spatial perturbation to it to generate a corresponding fake image; generating synthetic images with different degrees of forgery based on the real and fake images using an instance reclassification strategy; performing low-frequency enhancement processing on the synthetic images to obtain frequency-enhanced images; extracting spatial and frequency domain features from the frequency-enhanced images; fusing the extracted spatial and frequency domain features to obtain fused discriminative features; and classifying the input video frames as genuine or fake and determining the forgery intensity level based on the discriminative features. The beneficial effects of this invention are: it effectively alleviates the overfitting problem caused by the binary classification paradigm and, based on this, constructs a spatial-frequency domain feature fusion method, thereby improving the model's adaptability to diverse forgery types and its robustness in detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of deepfake video detection technology, specifically a method and system for identifying deepfake videos based on forged synthetic additional feature maps. Background Technology

[0002] With the rapid development of generative artificial intelligence technologies such as generative adversarial networks and diffusion models, deepfake technology can synthesize images and videos with extremely realistic visual effects and high deception capabilities. Although this technology has potential application value, its misuse has seriously threatened public trust, personal privacy, and social security, making robust deepfake detection technology an urgent need.

[0003] Current mainstream deepfake detection methods generally face two major challenges: First, insufficient generalization ability. Forgery technology itself is rapidly iterating and evolving, and the emergence of new forgery methods causes the detection model to significantly degrade in performance when faced with "unknown" or "unseen" forgery types, making it difficult to maintain robust discrimination ability. Second, most existing methods rely on a simple binary (true / false) classification learning paradigm, which has inherent defects: it forcibly discretizes the essentially continuous and gradual forgery intensity (such as the degree of fusion and the obviousness of editing traces) into black-and-white labels, resulting in "information collapse." Specifically, the model loses information about fine-grained differences in the degree of forgery during the learning process, which not only limits its discrimination accuracy but also makes it more likely to overfit the model on the training data, thereby further impairing its generalization performance in real complex scenarios.

[0004] To improve generalization, recent studies have proposed a data synthesis-based detection approach, constructing training data by actively simulating edge artifacts and texture inconsistencies generated during forgery. However, these methods still have significant limitations: first, their core paradigm remains within the framework of binary classification, failing to effectively utilize the continuous properties of forgery strength; second, their feature extraction is overly focused on the spatial domain (such as pixel-level texture and color in images), generally neglecting the analysis of frequency domain features. Research shows that forgery operations leave more consistent "fingerprints" or traces across methods in the frequency domain (especially in low-frequency bands). The lack of modeling for frequency domain features makes it difficult for models to capture the common patterns behind different forgery methods, thus limiting further improvements in generalization ability.

[0005] Therefore, there is an urgent need for a method and system for identifying deepfake videos based on forged synthetic additional feature maps to solve the above problems. Summary of the Invention

[0006] The purpose of this invention is to provide a method and system for deep forgery video discrimination based on forged synthetic additional feature maps. It can effectively alleviate the overfitting problem caused by the binary classification paradigm, and on this basis, a spatial-frequency domain feature fusion method is constructed to improve the model's adaptability to diverse forgery types and its robustness in detection.

[0007] To achieve the above objectives, the present invention employs the following technical solution: On the one hand, the present invention provides a method for identifying deepfake videos, comprising the following steps: Step S1: Obtain the real image and apply spatial perturbation to the real image to generate the corresponding fake image; Step S2: Based on real and fake images, a composite image with different degrees of forgery is generated using an instance reclassification strategy, wherein the instance reclassification strategy is implemented through a controllable mixing ratio. Quantify the degree of forgery into continuous monitoring signals; Step S3: Perform low-frequency enhancement processing on the synthesized image to enhance the low-frequency structural components in the image, thereby obtaining a frequency-enhanced image; Step S4: Extract the spatial domain features and frequency domain features of the frequency-enhanced image, respectively; Step S5: Fuse the extracted spatial domain features with the frequency domain features to obtain the fused discriminative features; Step S6: Based on the discriminative features, classify the input video frames as genuine or fake and determine the level of forgery.

[0008] Preferably, in step S2, generating synthetic images with different degrees of forgery using an instance reclassification strategy specifically involves: Real images Compared with forged images generated through spatial perturbation According to the preset mixing ratio Blend to generate a composite image , is represented as: ; in, Indicates pixel-by-pixel multiplication. A binary mask defining the blending transition region; blending ratio The larger the value, the more likely it is to be a composite image. The higher the strength of the forged signal; Constructing a joint loss function The model is supervised, and the joint loss function includes a binary cross-entropy loss used to distinguish between real and fake samples. and classification cross-entropy loss used to learn fine-grained differences between different levels of forgery. Its definition is: ; in, and For loss weights, and satisfying .

[0009] Preferably, in step S3, the low-frequency enhancement processing of the synthesized image specifically includes: For the synthesized image Wavelet decomposition is performed to decouple the components containing low frequencies. Horizontal high-frequency components Vertical high frequency components and diagonal high frequency components Frequency components; The wavelet decomposition is achieved through four-directional convolutional filters. Implementation, each filter It is used to capture frequency features in a specific direction, and the decomposition process is expressed as: ; in, This represents the convolution operation; Utilizing the low-frequency component Through transpose convolution operation Restore its spatial resolution and compare it with the original synthetic image. Perform pixel-level averaging to generate a frequency-enhanced image. , is represented as: ; in, This indicates pixel-level averaging.

[0010] Preferably, in step S4, extracting the spatial domain features of the frequency-enhanced image specifically involves: The frequency-enhanced image is processed using a pre-trained convolutional neural network as the spatial backbone network. Encode and extract high-level spatial feature representations, i.e., spatial domain features. The process is represented as follows: ; in, This refers to the encoder of the spatial backbone network.

[0011] Preferably, in step S4, extracting the frequency domain features of the frequency-enhanced image specifically involves: Spatial domain features Two-stage wavelet transform decomposition is performed to obtain the set of frequency domain components for direction and scale; For spatial domain features Each channel The set of bands after its two-level wavelet decomposition Defined as: ; spatial domain features The band sets of all channels are stitched together to obtain the initial frequency domain features. , is represented as: ; in, For the spatial domain features The total number of channels; The initial frequency domain features Input a frequency domain encoder consisting of multiple convolutional layers to obtain discriminative frequency domain features. The frequency domain encoder includes at least three layers, each of which performs convolution, batch normalization, and activation function processing in sequence, and includes a squeeze-excitation module to model the dependencies between channels.

[0012] Preferably, the processing procedure of the frequency domain encoder is as follows: For input The discriminative frequency domain features are obtained by sequentially going through the first encoding stage, the second encoding stage, and the third encoding stage. The process is represented as follows: ; ; ; in, , This indicates the squeeze-excitation module. This represents the GELU activation function. , , This indicates a batch normalization operation. Indicates the kernel size as Convolution operation, Indicates the kernel size as And a convolution operation with a stride of 2.

[0013] Preferably, in step S5, the extracted spatial domain features and frequency domain features are fused, specifically as follows: spatial domain features and frequency domain features Each is mapped to a unified embedding dimension through a projection layer; The mapped features are concatenated into a token sequence. ; The token sequence Input is based on a Transformer-based multi-head self-attention module for cross-domain information interaction; Enhancement tokens corresponding to spatial feature parts are extracted from the sequence output by the multi-head self-attention module and then passed through the output projection layer with adjustable weights. Features of the original spatial domain Perform residual connections to obtain the fused discriminative features. The process is represented as follows: ; in, The discriminative features after fusion To output the projection layer, This represents the first enhanced token extracted from the token sequence.

[0014] Preferably, in step S6, the model output includes a binary true / false classification result and a quantized value representing the forgery strength level, the quantized value being a blending ratio used in the instance reclassification strategy. Related.

[0015] On the other hand, the present invention provides a system for identifying deepfake videos, used to implement the deepfake video identification method described above, comprising: The image synthesis module is used to acquire real images, apply spatial perturbations to generate fake images, and generate synthetic images with different degrees of forgery based on an instance reclassification strategy; The low-frequency enhancement module is used to perform wavelet decomposition on the synthesized image, extract and enhance its low-frequency structural components, and generate a frequency-enhanced image. The feature extraction module is used to extract the spatial domain features and frequency domain features of the frequency-enhanced image, respectively. The feature fusion module is used to fuse the spatial domain features and frequency domain features to obtain discriminative features; The classification and determination module is used to output the authenticity classification result and the forgery intensity level of the input video frame based on the discrimination features.

[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. Through an innovative instance reclassification strategy, the degree of forgery is quantified as a continuous supervisory signal, enabling the model to learn fine-grained differences in forgery strength, rather than simply performing a true / false binary classification. This effectively avoids the degradation of feature representation capabilities and overfitting problems caused by information collapse, allowing the model to maintain stable and accurate detection performance when facing unknown or novel forgery techniques.

[0017] 2. By introducing a low-frequency enhancement module, low-frequency components representing global structural information in the image are actively extracted and enhanced, effectively compensating for the shortcomings of existing methods that over-rely on spatial domain texture details. Combined with the subsequently designed spatial-frequency domain feature fusion network, it can simultaneously utilize the local artifact sensitivity of the spatial domain and the cross-method stability of the frequency domain to achieve a more comprehensive and robust representation of forgery traces.

[0018] 3. Unlike traditional methods that only output binary results indicating whether something is true or false, the method of this invention can further output a quantitative level or risk assessment result reflecting the strength of forgery. This not only provides richer decision-making information, which is helpful for content hierarchical management, but also makes the model's decision-making process more interpretable.

[0019] 4. The method proposed in this invention, through efficient network architecture design (such as using a lightweight backbone network and a carefully designed fusion module), achieves high-precision detection and strong generalization capabilities while maintaining low computational complexity and model parameter quantity, thus enabling real-time or efficient batch processing detection of video streams and possessing good prospects for practical deployment. Attached Figure Description

[0020] Figure 1 This is a flowchart of a method according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the system structure according to an embodiment of the present invention; Figure 3 This is a schematic diagram illustrating the authenticity identification results and forgery detection results of an online video platform according to an embodiment of the present invention. Detailed Implementation

[0021] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. Furthermore, it should be understood that after reading the teachings of this invention, those skilled in the art can make various alterations or modifications to the invention, and these equivalent forms also fall within the scope defined in this application.

[0022] In this invention, terms such as "upper," "lower," "left," "right," "front," "back," "vertical," "horizontal," "side," and "bottom" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. These terms are only used to facilitate the description of the structural relationships of the various components or elements of this invention and do not specifically refer to any component or element in this invention. They should not be construed as limiting the invention.

[0023] Example: like Figure 1 As shown, this embodiment provides a method for identifying deepfake videos, including the following steps: Step S1: Obtain the real image and apply spatial perturbation to the real image to generate the corresponding fake image; Step S2: Based on real and fake images, a composite image with different degrees of forgery is generated using an instance reclassification strategy, wherein the instance reclassification strategy is implemented through a controllable mixing ratio. Quantify the degree of forgery into continuous monitoring signals; Step S3: Perform low-frequency enhancement processing on the synthesized image to enhance the low-frequency structural components in the image, thereby obtaining a frequency-enhanced image; Step S4: Extract the spatial domain features and frequency domain features of the frequency-enhanced image, respectively; Step S5: Fuse the extracted spatial domain features with the frequency domain features to obtain the fused discriminative features; Step S6: Based on the discriminative features, classify the input video frames as genuine or fake and determine the level of forgery.

[0024] Specifically: This embodiment proposes an instance reclassification method for forged samples to capture the diversity of forgery levels. This classification strategy explicitly incorporates the forgery level as a supervisory signal into the model, which helps mitigate information loss and promotes the learning of refined discriminative representations. Before the reclassification and mixing operation, this embodiment first applies a series of spatial perturbations to the original real image to enhance the model's sensitivity to local structural and texture changes. For the input real sample... The corresponding forged image obtained through spatial domain operations It can be represented as: ; in Indicates the first Subspace transformation.

[0025] IRGS (Instance Reclassification Strategy) aims to construct a reclassification framework for quantifying forgery severity, facilitating the model's learning of multi-level forgery features, and thus improving its generalization ability to different forgery methods. Its core innovation lies in transforming forgery severity into a computable quantitative indicator, namely, the mixture ratio. Original images can be used in the quantization synthesis process. With re-grading blended images The contribution weight, its mathematical expression is: ; in This represents pixel-by-pixel element-wise multiplication. For definition and Binary mask for transition boundaries. As can be seen from the above formula, the mixing ratio... The larger the value, the more likely it is to appear in the output image. The higher the proportion, the more significant the forged signal strength. This mechanism enables... It becomes a controllable and quantifiable index of forgery strength, which can be used to evaluate the detection behavior of the model under conditions of progressively increasing forgery strength.

[0026] IRGS is essentially an extension of the binary classification learning paradigm, aiming to introduce more refined modeling of forgery levels to improve discrimination capabilities. To this end, this embodiment designs a joint optimization mechanism that incorporates the basic binary cross-entropy loss used to distinguish between real and fake samples. Cross-entropy loss with classification In combination, the latter guides the model to learn fine-grained hierarchical differences between fake samples, as defined below: ; in For real labels, These are predicted values. Indicates the first Each sample in category The real labels in Indicates the first Each sample belongs to category The predicted probability. Therefore, the total loss is defined as: ; in and They are and The loss weights are calculated using IRGS. IRGS transforms the forgery intensity into a supervisory signal, guiding the model to learn multi-level forgery features, thereby improving its ability to detect forgeries.

[0027] Existing forgery detection methods typically focus on spatial domain operations, resulting in limited ability to describe the structural features of the generated images. To address this, this embodiment proposes a low-frequency enhancement module to enhance the low-frequency structural components in forged images generated by a re-hierarchical mixing operation. Specifically, the LFEM (Low-Frequency Enhancement Module) first applies a four-directional convolutional filter to perform wavelet decomposition, the expression of which is as follows: ; in Represents a filter set. Each Capture frequency characteristics in a specific direction. (Symbol) This represents the convolution operation. Among the components obtained from decoupling, the low-frequency components... Encoding the structural information of an image is the primary domain for characterizing forgery traces. Therefore, LFEM utilizes... Generate frequency-enhanced images The calculation method is as follows: ; in This represents the transpose convolution operation, which restores the spatial resolution to match the input. This indicates pixel-level averaging.

[0028] This embodiment proposes a spatial-frequency feature fusion network that integrates spatial and frequency domain information to enhance the model's sensitivity to forgeries.

[0029] To extract rich spatial information, this embodiment uses EfficientNet-B5 as the spatial backbone network. Specifically, for the input image... The spatial feature extractor encodes it into a high-level representation: ; in Indicates encoder, This represents the extracted spatial features.

[0030] To supplement the frequency domain representation of spatial features, SFFF (Spatial-Frequency Feature Fusion Network) designs a dedicated frequency domain stream. Input features First, a two-stage wavelet transform decomposition is used to generate multi-directional frequency domain components. For each channel... Band set It can be defined as: ; Then output features It is the splicing of all channels, which can be defined as: ; To be Encoding is a discriminative frequency representation. This embodiment designs a three-stage convolutional encoder. Each stage includes a convolutional layer, batch normalization, and a GELU activation function, and is enhanced by a squeeze-and-excitation (SE) module to model inter-channel dependencies. Specifically, for the input... The encoding process is defined as follows: ; in This represents the GELU activation function. This represents a convolution operation with a kernel size of . The step size is 2.

[0031] To simultaneously acquire complementary information in both the spatiotemporal domains, this embodiment designs a Transformer-based fusion module. Given spatial features... With frequency characteristics First, it is projected onto a unified embedding space and concatenated into a token sequence: ; After multi-head self-attention encoding, an enhanced spatial representation is obtained from the first token and fused with the original features in the following manner: ; in The fusion weights are adjustable hyperparameters. This fusion mechanism, through cross-domain attention modeling, enables spatial representations to be adaptively enriched by frequency-aware cues.

[0032] like Figure 2 As shown, this embodiment also provides a system for identifying deepfake videos, including: The image synthesis module is used to acquire real images, apply spatial perturbations to generate fake images, and generate synthetic images with different degrees of forgery based on an instance reclassification strategy; The low-frequency enhancement module is used to perform wavelet decomposition on the synthesized image, extract and enhance its low-frequency structural components, and generate a frequency-enhanced image. The feature extraction module is used to extract the spatial domain features and frequency domain features of the frequency-enhanced image, respectively. The feature fusion module is used to fuse the spatial domain features and frequency domain features to obtain discriminative features; The classification and determination module is used to output a authenticity classification result and a forgery strength level for the input video frame based on the discriminant features. This embodiment is deployed in the video review backend of an online content platform for deepfake detection and risk assessment of user-uploaded real-time video streams. The FAIR (Frequency Domain Enhanced Instance Reclassification) method proposed in this invention is integrated into the real-time review process. The system's operation flow is as follows: 1. Instance reclassification generation and learning strategy: Collect real videos from online video systems as real samples, and apply spatial perturbations to them. Generating fake images By mixing ratio Control the original image With fake images The weights are used to transform the degree of forgery into a calculable quantitative indicator. Then, a multi-level supervised model based on an online system is designed to guide the platform's model to learn the subtle differences between forged samples, thereby improving the model's generalization ability.

[0033] 2. Low-frequency enhancement and feature supplementation: Existing detection methods primarily focus on the spatial domain, easily overlooking forgery traces in the frequency domain. This system first detects forged images... Wavelet decomposition is performed to decouple low-frequency components to encode the structural information of the image. Then, transposed convolution is used to restore its spatial resolution and generate a frequency-enhanced image, enabling the recognition model of the online platform to capture the overall forgery traces in the video.

[0034] 3. Spatial-frequency domain feature fusion network: A spatial-frequency domain feature fusion network is deployed on an online platform, using EfficientNet-B5 as the spatial backbone network to extract spatial features. Frequency domain features are extracted through multi-level wavelet transform and convolutional encoder. .based on and The system is designed with a Transformer-based fusion module, which improves the detection accuracy and stability of the system under various forgery methods.

[0035] 4. System test results: In this embodiment, the video review system based on the FAIR method can not only achieve accurate binary classification of true and false content in real-time detection on the content platform, but also... Figure 3 As shown in Table 1, the system can also output fine-grained forgery strength levels, enabling multi-level risk assessment of video content. Under four commonly used forgery detection test sets, the authenticity detection performance reaches 96.2%, 99.3%, 84.8%, and 92.5%, respectively, representing an average improvement of over 4% compared to other detection methods. Furthermore, this system overcomes the limitations of the binary classification paradigm through IRGS, achieving a processing speed of 4065 frames per second with only 14.4MB of model parameters, processing 127 videos per second. This achieves a high level of generalization and high accuracy in real-time deep forgery detection of video.

[0036] Table 1 Performance Comparison of FAIR with Other Methods

[0037] The above is a detailed description of the preferred embodiments of the present invention, but the present invention is not limited to the embodiments described. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the present invention, and these equivalent modifications or substitutions are all included within the scope defined by the claims of this application.

Claims

1. A method for identifying deepfake videos, characterized in that, Includes the following steps: Step S1: Obtain the real image and apply spatial perturbation to the real image to generate the corresponding fake image; Step S2: Based on real and fake images, a composite image with different degrees of forgery is generated using an instance reclassification strategy, wherein the instance reclassification strategy is implemented through a controllable mixing ratio. Quantify the degree of forgery into continuous monitoring signals; Step S3: Perform low-frequency enhancement processing on the synthesized image to enhance the low-frequency structural components in the image, thereby obtaining a frequency-enhanced image; Step S4: Extract the spatial domain features and frequency domain features of the frequency-enhanced image, respectively; Step S5: Fuse the extracted spatial domain features with the frequency domain features to obtain the fused discriminative features; Step S6: Based on the discriminative features, classify the input video frames as genuine or fake and determine the level of forgery. In step S2, the specific steps for generating synthetic images with different degrees of forgery using an instance reclassification strategy are as follows: Real images Compared with forged images generated through spatial perturbation According to the preset mixing ratio Blend to generate a composite image , is represented as: ; in, Indicates pixel-by-pixel multiplication. A binary mask defining the blending transition region; blending ratio The larger the value, the more likely it is to be a composite image. The higher the strength of the forged signal; Constructing a joint loss function The model is supervised, and the joint loss function includes a binary cross-entropy loss used to distinguish between real and fake samples. And classification cross-entropy loss used to learn fine-grained differences between different levels of forgery. Its definition is: ; in, and For loss weights, and satisfying .

2. The method for identifying deepfake videos according to claim 1, characterized in that, In step S3, the low-frequency enhancement processing of the synthesized image specifically includes: For the synthesized image Wavelet decomposition is performed to decouple the components containing low frequencies. Horizontal high-frequency components Vertical high frequency components and diagonal high frequency components Frequency components; The wavelet decomposition is achieved through four-directional convolutional filters. Implementation, each filter It is used to capture frequency features in a specific direction, and the decomposition process is expressed as: ; in, This represents the convolution operation; Utilizing the low-frequency component Through transpose convolution operation Restore its spatial resolution and compare it with the original synthetic image. Perform pixel-level averaging to generate a frequency-enhanced image. , is represented as: ; in, This indicates pixel-level averaging.

3. The method for identifying deepfake videos according to claim 1, characterized in that, In step S4, the extraction of spatial domain features from the frequency-enhanced image specifically involves: The frequency-enhanced image is processed using a pre-trained convolutional neural network as the spatial backbone network. Encode and extract high-level spatial feature representations, i.e., spatial domain features. The process is represented as follows: ; in, This refers to the encoder of the spatial backbone network.

4. The method for identifying deepfake videos according to claim 3, characterized in that, In step S4, the extraction of frequency domain features from the frequency-enhanced image specifically involves: Spatial domain features Two-stage wavelet transform decomposition is performed to obtain the set of frequency domain components for direction and scale; For spatial domain features Each channel The set of bands after its two-level wavelet decomposition Defined as: ; spatial domain features The band sets of all channels are stitched together to obtain the initial frequency domain features. , is represented as: ; in, For the spatial domain features The total number of channels; The initial frequency domain features Input a frequency domain encoder consisting of multiple convolutional layers to obtain discriminative frequency domain features. The frequency domain encoder includes at least three layers, each of which performs convolution, batch normalization, and activation function processing in sequence, and includes a squeeze-excitation module to model the dependencies between channels.

5. The method for identifying deepfake videos according to claim 4, characterized in that, The specific processing procedure of the frequency domain encoder is as follows: For input The discriminative frequency domain features are obtained by sequentially going through the first encoding stage, the second encoding stage, and the third encoding stage. The process is represented as follows: ; ; ; in, , This indicates the squeeze-excitation module. This represents the GELU activation function. , , This indicates a batch normalization operation. Indicates the kernel size as Convolution operation, Indicates the kernel size as And a convolution operation with a stride of 2.

6. The method for identifying deepfake videos according to claim 1, characterized in that, In step S5, the extracted spatial domain features and frequency domain features are fused, specifically as follows: spatial domain features and frequency domain features Each is mapped to a unified embedding dimension through a projection layer; The mapped features are concatenated into a token sequence. ; The token sequence Input is based on a Transformer-based multi-head self-attention module for cross-domain information interaction; Enhancement tokens corresponding to spatial feature parts are extracted from the sequence output by the multi-head self-attention module and then passed through the output projection layer with adjustable weights. Features of the original spatial domain Perform residual connections to obtain the fused discriminative features. The process is represented as follows: ; in, The discriminative features after fusion To output the projection layer, This represents the first enhanced token extracted from the token sequence.

7. The method for identifying deepfake videos according to claim 1, characterized in that, In step S6, the model output includes a binary true / false classification result and a quantized value representing the level of forgery strength, the quantized value being mixed with the mixing ratio used in the instance reclassification strategy. Related.

8. A system for identifying deepfake videos, used to implement the method for identifying deepfake videos as described in any one of claims 1-7, characterized in that, include: The image synthesis module is used to acquire real images, apply spatial perturbations to generate fake images, and generate synthetic images with different degrees of forgery based on an instance reclassification strategy; The low-frequency enhancement module is used to perform wavelet decomposition on the synthesized image, extract and enhance its low-frequency structural components, and generate a frequency-enhanced image. The feature extraction module is used to extract the spatial domain features and frequency domain features of the frequency-enhanced image, respectively. The feature fusion module is used to fuse the spatial domain features and the frequency domain features to obtain discriminative features; The classification and determination module is used to output the authenticity classification result and the forgery intensity level of the input video frame based on the discrimination features.

Citation Information

Patent Citations

  • Deep forgery detection method based on self-mixing

    CN119672816A

  • Machine-learning-based detection of fake videos

    US20250166358A1