A deep fake face image detection method based on reconstruction difference detection

CN122821607APending Publication Date: 2026-09-25TIANJIN UNIVERSITY OF TECHNOLOGY +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611214849.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-11
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

尽管有这些不同的探索,但现有的方法有一个基本的局限性:它们仅将重建用作辅助工具或解纠缠约束,最终目标仍然是提取更好的静态表示或明确的视觉缺陷

Benefits of technology

本发明公开了一种基于重构差异探测的深度伪造人脸图像检测方法,本发明通过受控重构过程对输入图像进行主动探测,并结合重构差异定性分析、区域统计量定量分析以及多步迭代重构残差衰退分析,表征真实图像与伪造图像在相同重构条件下产生的非对称重构响应;进一步地,本发明构建重构关系差异特征提取机制,对原始图像与重构图像在特征关系空间中的差异变化进行建模,并将原始语义特征与重构关系差异特征融合用于真假判别;由此,本发明能够有效捕捉真实样本与伪造样本之间更具判别性和迁移性的重构差异线索,提高模型对细微伪造痕迹的检测能力,并增强其在未知伪造方法、跨数据集及复杂篡改场景下的泛化性能;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821607A_ABST
    Figure CN122821607A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of computer vision and artificial intelligence security, and particularly relates to a deep fake face image detection method based on reconstruction difference detection, which specifically comprises the following steps: firstly, controlled reconstruction is performed on an input image, a difference map, a difference heat map and regional statistical indicators between the original image and the reconstructed image are constructed, and the asymmetric reconstruction response difference between the real image and the fake image under the same reconstruction condition is analyzed; then, the original image and the reconstructed image are input into a shared feature extractor fine-tuned by LoRA to obtain original features and reconstructed features respectively, and reconstruction difference features are constructed according to the original features and the reconstructed features; finally, the original features and the reconstruction difference features are fused, a deep fake probability is output, and true or false discrimination is completed. Through quantitative analysis and qualitative analysis, the asymmetric response of true or false pictures generated by the reconstruction module can be analyzed, and through the deep fake detection method based on the captured difference, the detection performance of the model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision and artificial intelligence security technology, and in particular to a method for detecting deepfake face images based on reconstruction difference detection. Background Technology

[0002] The rapid development of generative models, particularly generative adversarial networks (GANs) and diffusion models, has significantly improved the realism and diversity of synthetic images. While these techniques enable powerful image generation capabilities, they also facilitate the creation of highly realistic deepfakes that are difficult to distinguish from real content. Such manipulated media poses a serious risk to information security, social trust, and personal reputation. As generative technologies continue to evolve, forged images are becoming increasingly diverse and complex, making reliable deepfake detection a critical task. In particular, developing detectors that can generalize to unknown manipulation methods remains a fundamental challenge for the computer vision community.

[0003] To mitigate these challenges, existing deepfake detection methods have explored several directions, including hybrid artifacts, frequency domain analysis, data augmentation strategies, and feature deentanglement. Recently, some methods have leveraged large-scale pre-trained models to learn more general and semantically rich representations. These methods typically follow two directions: introducing adapters to fine-tune CLIPs, or optimizing cues to better utilize their representational capabilities. While these methods improve generalization to some extent, they remain fundamentally passive observation paradigms, relying on static representations of single images for forgery detection. This strategy is prone to overfitting to generator-specific artifacts, thus limiting further improvements in generalization.

[0004] To mitigate overfitting caused by static feature extraction, several approaches have explored reconstruction learning for deep forgery detection. Early work focused on exposing manipulation traces through input reconstruction. For example, Cao et al. amplified local artifacts by reconstructing a classification framework, while Wang et al. used reconstruction differences to decouple identity from forgery cues. Recent methods aim to learn generalized priors or unentangled representations. RAM enhances generalization by restoring distorted faces to their original appearance, while forgetting-based keyframes employ multi-scale reconstruction to filter out dataset-specific artifacts. Furthermore, DiffusionFake utilizes diffusion priors to guide reconstruction and extract unentangled forgery features. Despite these diverse explorations, existing methods share a fundamental limitation: they merely use reconstruction as an auxiliary tool or unentanglement constraint, with the ultimate goal remaining the extraction of better static representations or explicit visual defects. Consequently, their detection process remains trapped in a passive observation paradigm.

[0005] To address the aforementioned problems, this invention proposes a fundamental paradigm shift from passive observation to active exploration, and presents a deepfake face image detection method based on reconstruction difference detection to solve these issues. This invention discovers an asymmetric response phenomenon between real and fake images during controlled reconstruction, and based on this, proposes a modeling of the representational changes caused by the controlled reconstruction process, providing a new perspective for learning more general discriminative cues. Summary of the Invention

[0006] To address the shortcomings of existing technologies, this invention proposes a deepfake face image detection method based on reconstruction difference detection. Through quantitative and qualitative analysis, the asymmetric responses generated by the reconstruction module of real and fake images can be analyzed. By using a deepfake detection method based on capturing differences, the detection performance of the model can be improved.

[0007] The technical solution of this invention to solve the technical problem is a method for detecting deepfake face images based on reconstruction difference detection, comprising the following steps: S1. Obtain the image to be detected, and perform detection, alignment, cropping, and size normalization on the face regions in the image to obtain the input image. ; S2. The controlled reconstruction module inputs the input image to the freeze parameters and generates a reconstructed image that is semantically consistent with the input image. ; S3, Input Image and reconstructed image Pixel-level residual statistics and multi-step iterative reconstruction dynamic analysis were performed, and statistical significance tests were conducted to determine whether the difference in response between real and fake images during the reconstruction process objectively exists. S4. Input image With reconstructed image Each input is a feature extractor with shared parameters, and the corresponding output is the original feature vector. and reconstructing feature vectors The feature extractor uses the CLIP ViT-L / 14 visual encoder as the backbone network. LoRA low-rank adaptation modules are injected into the self-attention layers of the backbone network, and the original parameters of the backbone network are kept frozen. S5. Convert the original feature vector and reconstructing feature vectors Input the difference construction module, output the reconstructed difference feature vector. ; S6. Convert the original feature vector and reconstructing the difference feature vector The features are concatenated and fused to obtain a fused feature vector. ; fuse feature vectors Input a classification header and output the predicted probability that the image to be detected is a fake image; S7. During the model training phase, the binary cross-entropy loss is calculated based on the predicted probability and the corresponding real label. The LoRA parameter in the feature extractor and all parameters of the classification head are updated through backpropagation of the loss. During the model inference phase, the deep forgery detection result is output based on the predicted probability.

[0008] S1 is as follows: If the input is video data, then the video is first extracted frame by frame to obtain the image to be detected; The face detection algorithm is used to locate the face region in the image to be detected. Geometric alignment is performed based on facial key points. The aligned face region is then cropped and scaled to a uniform size to obtain the input image. .

[0009] The controlled reconstruction module in S2 is a parameter-frozen variational autoencoder; The process of generating a reconstructed image is specifically as follows: taking the input image... The encoder maps the image to the latent space, outputting the encoded mean and encoded variance. Only the encoded mean is used as a deterministic latent variable input to the decoder to generate the reconstructed image. .

[0010] S3 is as follows: (1) Input image and reconstructed image Perform pixel-level subtraction on corresponding channels, and take the absolute value of the difference to obtain the absolute difference image. By analyzing the input image and reconstructed image The residuals of each color channel are averaged to calculate the difference heatmap. , in, Represents the RGB color channels. , Table of RGB color channel indexes, Indicates the first Input image with 1 channel, Indicates the first Reconstructed image of each channel, Indicates taking the absolute value; Preset the region of interest for the face, based on the input image Reconstructing images and difference heatmap Calculate the mean absolute error and high response residual density within the region of interest (ROI) of the face; define a spatial binary mask for extracting the ROI. , Indicates altitude, Represents width, spatial binary mask Limited to the input image The core facial area; Mean Absolute Error and high response residual density The calculation formula is as follows: , , in, Represents pixels, This represents the total number of valid pixels within the face mask. Indicates the total number of channels. Indicates the channel index. Indicates the input image The Middle Each channel pixel The value at that location, Represents the reconstructed image The Middle Each channel pixel The value at that location, For indicator functions, Heatmap of differences medium pixel The value at that location, High response threshold; (2) Input image Perform multi-step iterative reconstruction, letting the initial input... , No. The next iteration projection image is defined as , ,in Indicates the controlled reconfiguration module; calculates the first... Mask during the next iteration Iterative residual energy within : ; The residual energies from multiple consecutive iterations are concatenated in the order of iteration to form a dynamic residual feature vector; (3) Calculate the mean absolute error based on the real image set and the fake image set respectively. and high response residual density The sample mean, sample variance, and sample size were analyzed for significant differences using the Welch test. Let the sample mean of the real image set be... l. The sample variance is Sample size is The sample mean of the forged image set is The sample variance is Sample size is Calculate the test statistic With Satterthwaite's Degree of Freedom : , ; Let the null hypothesis be... The real image and the fake image have the same expected reconstruction residual; Under the original hypothesis Below, based on the degrees of freedom students Distribution calculation two-sided value: , in, express value; Indicates the null hypothesis Under the condition that it holds true, it follows the order of degrees of freedom. students The test statistic of the distribution of a random variable; Represents a probability function; This represents a given condition in conditional probability; If the calculation yields If the value is less than the preset significance level, it is determined that there is a significant difference between the real image and the fake image in the reconstruction residual, that is, the objective existence of the reconstruction difference is verified. Finally, the mean absolute error High response residual density The dynamic residual eigenvectors and statistical test results are used together as the output of this step to support the determination of the existence of differences.

[0011] S4 is as follows: LoRA low-rank adaptation module is injected into the query projection layer and value projection layer of the Transformer self-attention layer in the backbone network; Let the original projection matrix in the self-attention layer be... The projection matrix after LoRA injection satisfy: , , in, Represents a dimension reduction matrix. ; Represents an upgraded matrix. ; Denotes the rank of LoRA This represents the input and output feature dimensions of the projection layer, and .

[0012] Step S5, the difference construction module, includes two processing schemes, as follows: Option 1: Convert the original feature vector With reconstructing feature vectors By performing element-wise subtraction, we obtain the reconstructed difference feature vector. : , in, , Represents the feature dimension, the original feature vector. and reconstructing feature vectors The same dimensions ; Option 2: Use a shared dimensionality reduction layer to perform dimensionality reduction mapping on the original feature vector and the reconstructed feature vector respectively, to obtain the corresponding low-dimensional original features. With low-dimensional reconstruction features The shared dimensionality reduction layer consists of a linear mapping layer and a batch normalization layer connected in series, and the original feature vector and the reconstructed feature vector share the same set of dimensionality reduction layer parameters. Calculate the self-similarity matrix of the low-dimensional original features respectively. Self-similarity matrix of low-dimensional reconstructed features The calculation formula is as follows: , , in, Indicates the feature dimension after dimensionality reduction: The difference between two self-similar matrices is the relation difference matrix. , ; The relational difference matrix is ​​flattened into a one-dimensional vector, normalized, and then input into a multilayer perceptron for mapping, outputting a reconstructed difference feature vector. The calculation formula is as follows: , in, Indicates the flattening operation. Representation layer normalization, This represents a multilayer perceptron mapping.

[0013] S6 is detailed below: The classification head consists of a first linear layer, a batch normalization layer, a ReLU activation layer, a Dropout layer, a second linear layer, and a Softmax layer connected in series. fuse feature vectors The input to the classification header yields the predicted probability that the image to be detected is a fake image. The calculation formula is as follows: , in , representing a linear mapping layer; Indicates the batch normalization layer; Represents a non-linear activation function; Indicates a randomly deactivated layer; This represents a normalized classification function; This represents the predicted probability corresponding to the counterfeit category.

[0014] S7 binary cross-entropy loss Specifically as follows: , in, Indicates the true label, , This indicates a forged image. Represents a real image; This represents the natural logarithm function. The effects provided in the invention description are merely those of the embodiments, and not all the effects of the invention. The above technical solution has the following advantages or beneficial effects: This invention discloses a deepfake face image detection method based on reconstruction difference detection. The invention actively detects the input image through a controlled reconstruction process and combines qualitative analysis of reconstruction differences, quantitative analysis of region statistics, and multi-step iterative reconstruction residual decay analysis to characterize the asymmetric reconstruction responses of real and fake images under the same reconstruction conditions. Furthermore, the invention constructs a reconstruction relationship difference feature extraction mechanism to model the difference changes between the original and reconstructed images in the feature relationship space, and fuses the original semantic features with the reconstruction relationship difference features for real / fake discrimination. Therefore, this invention can effectively capture more discriminative and transferable reconstruction difference clues between real and fake samples, improve the model's ability to detect subtle forgery traces, and enhance its generalization performance in unknown forgery methods, cross-dataset scenarios, and complex tampering scenarios. Unlike existing methods that rely primarily on passive observation of static forgery traces in a single image for judgment, this invention introduces a controlled reconstruction process to actively detect the input face image, thereby transforming deepfake detection from "passively observing existing artifacts in the image" to a new detection paradigm of "actively stimulating and modeling differences in reconstruction response". By analyzing the difference maps, difference heatmaps, region of interest statistical indicators, and residual energy of multi-step iterative reconstruction between the original and reconstructed images, this invention discovers that real and forged images exhibit asymmetric reconstruction responses under the same controlled reconstruction conditions: real images are more prone to significant representational drift and residual changes during reconstruction, while forged images, due to their proximity to the generative model distribution, show relatively weaker and faster-stabilizing difference responses before and after reconstruction. Based on these findings, this invention further constructs a reconstruction relationship difference modeling framework, inputting the original and reconstructed images into a shared feature extractor, and modeling the difference changes between the two in the feature relationship space through a relationship difference feature extraction module. Then, the original semantic features and the reconstructed relationship difference features are fused for real-fake discrimination.

[0015] Through the above technical solution, the present invention can utilize stable difference cues induced by controlled reconstruction to reduce the model's dependence on artifacts of specific forgery methods or datasets, thereby improving the accuracy, robustness and cross-scene generalization ability of deepfake face image detection. Attached Figure Description

[0016] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof.

[0017] Figure 1 This is a schematic diagram of the method flow of the present invention.

[0018] Figure 2 A visual diagram illustrating the qualitative differences in reconstruction.

[0019] Figure 3 To reconstruct a quantitative visualization of the differences, Figure (a) shows the mean absolute error distribution and high response residual density distribution based on the CDFv2 dataset, and Figure (b) shows the mean absolute error distribution and high response residual density distribution based on the SD2.1 dataset.

[0020] Figure 4 The decay process of VAE iterative reconstruction images is shown in Figure (a), which is the residual energy decay curve of VAE multi-step iterative reconstruction based on dataset CDFv2, and Figure (b) is the residual energy decay curve of VAE multi-step iterative reconstruction based on dataset SD2.1. Detailed Implementation

[0021] To clearly illustrate the technical features of this solution, the invention will be described in detail below through specific implementation methods and in conjunction with the accompanying drawings.

[0022] Example 1 like Figure 1As shown, a method for detecting deepfake face images based on reconstruction difference detection includes the following steps: S1. Obtain the image to be detected, and perform detection, alignment, cropping, and size normalization on the face regions in the image to obtain the input image. ; S2. The controlled reconstruction module inputs the input image to the freeze parameters and generates a reconstructed image that is semantically consistent with the input image. ; S3, Input Image and reconstructed image Pixel-level residual statistics and multi-step iterative reconstruction dynamic analysis were performed, and statistical significance tests were conducted to determine whether the difference in response between real and fake images during the reconstruction process objectively exists. S4. Input image With reconstructed image Each input is a feature extractor with shared parameters, and the corresponding output is the original feature vector. and reconstructing feature vectors The feature extractor uses the CLIP ViT-L / 14 visual encoder as the backbone network. LoRA low-rank adaptation modules are injected into the self-attention layers of the backbone network, and the original parameters of the backbone network are kept frozen. S5. Convert the original feature vector and reconstructing feature vectors Input the difference construction module, output the reconstructed difference feature vector. ; S6. Convert the original feature vector and reconstructing the difference feature vector The features are concatenated and fused to obtain a fused feature vector. ; fuse feature vectors Input a classification header and output the predicted probability that the image to be detected is a fake image; S7. During the model training phase, the binary cross-entropy loss is calculated based on the predicted probability and the corresponding real label. The LoRA parameter in the feature extractor and all parameters of the classification head are updated through backpropagation of the loss. During the model inference phase, the deep forgery detection result is output based on the predicted probability.

[0023] In a specific implementation, S1 is as follows: To obtain the image to be detected, if the input is video data, the video is first uniformly frame-by-frame extracted, and the key frames are used as the image to be detected. For each frame of the image to be detected, face detection algorithms such as MTCNN (Multi-Task Cascaded Convolutional Neural Networks) or RetinaFace (Retina Face Detection Network) are used to locate the face region. Based on the detected facial key points (such as eyes, nose, and corners of the mouth), affine transformations are performed to achieve geometric alignment. Then, the aligned face region is cropped and scaled to a uniform size to obtain the input image. .

[0024] In a specific implementation, S2 is as follows: The controlled reconstruction module is a parameter-frozen variational autoencoder; it converts the input image... The input parameters are frozen and the variational autoencoder (VAE) is reconstructed. The process of generating the reconstructed image is as follows: The VAE encoder takes the input image... The mapping is represented by the encoded mean and encoded variance in the latent space. To eliminate the uncertainty introduced by random sampling, only the encoded mean is used as a deterministic latent variable input to the decoder to generate the reconstructed image. During this process, all parameters of the VAE remain frozen and do not participate in any training updates, ensuring the stability and controllability of the reconstruction process.

[0025] In a specific implementation, S3 is as follows: (1) Input image and reconstructed image Perform pixel-level subtraction on corresponding channels, and take the absolute value of the difference to obtain the absolute difference image. By analyzing the input image and reconstructed image The residuals of each color channel are averaged to calculate the difference heatmap. , in, Represents the RGB color channels. , Table of RGB color channel indexes, Indicates the first Input image with 1 channel, Indicates the first Reconstructed image of each channel, Indicates taking the absolute value; Preset the region of interest for the face, based on the input image Reconstructing images and difference heatmap Calculate the mean absolute error (MAE) and high response residual density within the region of interest (ROI) of the face; define a spatial binary mask for extracting the ROI. (Covering the main facial features such as eyes, nose, and mouth) Indicates altitude, Represents width, spatial binary mask Limited to the input image The core facial area; Mean Absolute Error and high response residual density The calculation formula is as follows: , , in, Represents pixels, This represents the total number of valid pixels within the face mask. Indicates the total number of channels. Indicates the channel index. Indicates the input image The Middle Each channel pixel The value at that location, Represents the reconstructed image The Middle Each channel pixel The value at that location, For indicator functions, Heatmap of differences medium pixel The value at that location, For high response threshold (threshold) Set to 1.5 times the global mean of the difference heatmap). like Figure 2 As shown, by comparing the original image, reconstructed image, absolute difference image, and difference heatmap of the real image and the forged image, it can be clearly observed that the residual response (highlighted areas in the difference image and heatmap) generated by the real image after reconstruction is significantly stronger than that of the forged image, while the reconstruction residual distribution of the forged image is more uniform and the intensity is lower. This visualization result intuitively proves that there is an asymmetric reconstruction response between the real image and the forged image under the same controlled reconstruction conditions.

[0026] (2) Input image Perform multi-step iterative reconstruction, letting the initial input... ,right The result of the first reconstruction And then Reconstruction yields This process is repeated K times for reconstruction. Let K=5, and the residual energy of each iteration is calculated. The dynamic residual eigenvector is obtained. This is used to analyze the difference in residual fading between real and fake images during the iterative reconstruction process; The next iteration projection image is defined as , ,in Indicates the controlled reconfiguration module; calculates the first... Mask during the next iteration Iterative residual energy within : ; The residual energies from multiple consecutive iterations are concatenated in the order of iteration to form a dynamic residual feature vector; like Figure 4 As shown, based on two representative datasets used to demonstrate the objective existence of differences—CDFv2 (Celebrity Deepfake Dataset Version 2) and SD2.1 (Stable Diffusion Model Version 2.1)—iterative residual energy decay curves for real and fake images are plotted. Significantly different decay patterns can be observed: the residual energy of real images decreases slowly with increasing iterations, maintaining a high residual level in each reconstruction round; while the residual energy of fake images rapidly decreases to a minimum after the first reconstruction round, with subsequent iterations producing almost no significant residuals. This dynamic difference indicates that fake images, due to their proximity to the manifold distribution of the generative model, can quickly converge to a stable state during controlled reconstruction, while real images, deviating from this manifold, continuously exhibit significant reconstruction drift.

[0027] (3) Statistical significance test: Collect the mean absolute error of real images and fake images on the training set respectively. and high response residual density The statistics were calculated, including the sample mean, variance, and sample size. The Welch test was used to test the significance of the two-sample differences. Let the sample mean of the real image set be... l. The sample variance is Sample size is The sample mean of the forged image set is The sample variance is Sample size is Calculate the test statistic With Satterthwaite's Degree of Freedom : , ; Let the null hypothesis be... The real image and the fake image have the same expected reconstruction residual; Under the original hypothesis Below, based on the degrees of freedom students Distribution calculation two-sided value: , in, express value; Indicates the null hypothesis Under the condition that it holds true, it follows the order of degrees of freedom. students The test statistic of the distribution of a random variable; Represents a probability function; This represents a given condition in conditional probability; The preset significance level is 0.05. If the calculated significance level is... If the value is less than 0.05, the null hypothesis is rejected, and it is determined that there is a significant difference between the reconstruction residuals of the real image and the forged image; like Figure 3 As shown, statistical significance tests were conducted using two representative datasets, CDFv2 (Celebrity Deepfake Dataset Version 2) and SD2.1 (Stable Diffusion Model Version 2.1), to demonstrate the objective existence of the differences. Figure 3 The middle box plot clearly displays the true image. and τ The numerical distribution is significantly higher than that of the fake image, and there is almost no overlap between the two sets of data. The Welch test gives... The value is much less than 0.05 (on multiple datasets). Values ​​can be as low as 10 -9 ~10 -69 The magnitude of the difference (in terms of magnitude) has been statistically rigorously verified that this difference is not due to random fluctuations, but rather to the essential difference in the reconstruction dynamics between the two types of images.

[0028] Finally, the mean absolute error High response residual density The dynamic residual feature vector and statistical test results are combined as outputs of this step to support the determination of the existence of differences. The determination result of this step serves as the internal mechanism supporting the effectiveness of the method and does not participate in subsequent feature fusion and classification processes.

[0029] pass Figure 2 Qualitative visualization intuitively demonstrates the asymmetric response of real and fake images to the reconstruction residuals; Figure 3 Quantitative statistics and Welch's test rigorously confirmed the significance of this difference from a statistical perspective. Figure 4 The iterative reconstruction residual decay curve further reveals the essential difference between the two in reconstruction dynamics, namely, the real image continues to drift while the fake image converges quickly; these phenomena together constitute the core theoretical support for the method of this invention.

[0030] In a specific implementation, S4 is as follows: Input image and reconstructed image Input the CLIP ViT-L / 14 (Contrastive Language-Image Pre-Training VisionTransformer Large with patch size 14) visual encoder with shared parameters. This encoder injects LoRA (Low-Rank Adaptation) modules into the query projection layer and value projection layer of the Transformer self-attention layer. All parameters of the original encoder are frozen. Let the original projection matrix in the self-attention layer be... The projection matrix after LoRA injection satisfy: , , in, Represents a dimension reduction matrix. ; Represents an upgraded matrix. ; Denotes the rank of LoRA This represents the input and output feature dimensions of the projection layer, and ,set up .

[0031] In a specific implementation, the difference construction module in step S5 includes two processing schemes, as follows: Option 1: Convert the original feature vector With reconstructing feature vectors By performing element-wise subtraction, we obtain the reconstructed difference feature vector. : , in, , Represents the feature dimension, the original feature vector. and reconstructing feature vectors The same dimensions ; Obtain the original feature vector output in step S4. and reconstructing feature vectors The original feature vector and the reconstructed feature vector have the same feature dimension, and take the following form: , in, This represents the feature dimension of the feature extractor's output. Since the input image and the reconstructed image are encoded using the same shared feature extractor, their feature vectors reside in the same feature representation space. Subtracting the original feature vector from the reconstructed feature vector element by element yields the reconstructed difference feature vector: , in, The reconstructed differential feature vector is used to describe the representational changes in the feature space produced by the controlled reconstruction of the input image; For real images, the changes to natural textures, imaging noise, and local details during the reconstruction process will generate corresponding responses in the reconstruction difference features; for forged images, the feature response patterns before and after reconstruction differ from those of real images; therefore, the reconstruction difference features can provide feature basis for distinguishing between real and forged images. Option 2: Use a shared dimensionality reduction layer to perform dimensionality reduction mapping on the original feature vector and the reconstructed feature vector respectively, to obtain the corresponding low-dimensional original features. With low-dimensional reconstruction features The shared dimensionality reduction layer consists of a linear mapping layer and a batch normalization layer connected in series, and the original feature vector and the reconstructed feature vector share the same set of dimensionality reduction layer parameters. Calculate the self-similarity matrix of the low-dimensional original features respectively. Self-similar matrix of low-dimensional reconstructed features The calculation formula is as follows: , , in, Indicates the feature dimension after dimensionality reduction: The difference between two self-similar matrices is the relation difference matrix. , ; The relational difference matrix is ​​flattened into a one-dimensional vector, normalized, and then input into a multilayer perceptron for mapping, outputting a reconstructed difference feature vector. The calculation formula is as follows: , in, Indicates the flattening operation. Representation layer normalization, This represents a multilayer perceptron mapping.

[0032] In a specific implementation, S6 is as follows: The original feature vector Reconstructing the difference feature vector By concatenating along the channel dimension, a fused feature vector is obtained. ; The classification heads are sequentially composed of the first linear layer (which will...) It consists of a 512-dimensional mapping layer, a batch normalization layer, a ReLU activation layer, a Dropout layer (with a dropout rate of 0.3), a second linear layer (mapping 512 to 2 dimensions), and a Softmax layer connected in series. fuse feature vectors The input to the classification header yields the predicted probability that the image to be detected is a fake image. The calculation formula is as follows: , in , representing a linear mapping layer; Indicates the batch normalization layer; Represents a non-linear activation function; Indicates a randomly deactivated layer; This represents a normalized classification function; This represents the predicted probability corresponding to the counterfeit category.

[0033] In a specific implementation, the binary cross-entropy loss in S7 Specifically as follows: , in, Indicates the true label, , This indicates a forged image. Represents a real image; This represents the natural logarithm function.

[0034] Example 2 To comprehensively evaluate the generalization ability of this method, comparative experiments were conducted on several commonly used deepfake datasets. The experimental setup is as follows: the model was trained on the FaceForensics++ dataset (FF++ represents a widely used benchmark dataset for deepfake face detection), and then evaluated directly on cross-dataset test sets such as CDFv1Celeb-DF version 1 (the first version of the celebrity deepfake dataset), CDFv2 (Celeb-DF version 2, the second version of the celebrity deepfake dataset), DFDDeepfake Detection (Google deepfake detection dataset), DFDCDeepfake Detection Challenge (deepfake detection challenge dataset), and DFDCP (Deepfake Detection Challenge Preview dataset), without any domain adaptation or fine-tuning. The frame-level AUC (Area Under the ROC Curve, a metric for evaluating the performance of binary classification models, with values ​​closer to 1 indicating better performance) evaluation results are shown in Table 1, the video-level AUC evaluation results are shown in Table 2, and the ablation experiments are shown in Table 3.

[0035] Table 1 Frame-level AUC Evaluation Results Compared to current state-of-the-art generalization-oriented methods (Xception, EfficientB4, F3Net, Face X-ray, FFD, SPSL, SRM, Recce, SBI, UCF, LSDA, ACMF, ADA-FInfer, ForAda, FreqDebias, FIA-USA), the method of this invention achieves the best average performance (average AUC of 90.48%) across all benchmark tests, consistently outperforming the compared methods. This demonstrates stronger robustness to unknown forgery distributions. Existing methods, even with techniques such as frequency domain debiasing, artifact synthesis, or domain adaptation, still rely on static feature representations; while the method of this invention models the differences between the original image and its reconstructed image, enabling the capture of more fundamental and transferable forgery clues, thus achieving superior generalization ability. This result also confirms... Figures 2 to 4 The revealed asymmetric reconstruction difference phenomenon is due to the fact that real images and fake images have essentially separable dynamic features in their reconstruction responses, which enables the detection model based on reconstruction differences to obtain stronger cross-domain generalization performance.

[0036] Table 2 Video-level AUC Evaluation Results Table 2 presents the video-level AUC results across datasets. Compared to recent top-tier methods such as seeABLE, TALL++, LAA-NET, LSDA, CFM, ForAda, FIA-USA, and Effort, the method of this invention achieves state-of-the-art performance on all evaluation datasets, with an average AUC of 93.79%. ForAda, as the strongest baseline, has an average AUC of 91.93%, and the method of this invention improves upon it by +1.86% on average. Furthermore, it achieves stable performance improvements on the CDFv2, DFD, and DFDC datasets. These results further validate the excellent cross-dataset generalization ability of the method of this invention and also corroborate the discriminative effectiveness of the reconstructed differential features from the perspective of downstream task performance.

[0037] Table 3 Ablation Experiment In the ablation experiments, LoRA-CLIP indicates whether the LoRA adaptation module is injected into the feature extractor, and PRD indicates whether the reconstructed difference feature proposed in this invention (i.e., the relational difference feature in step S5) is introduced. Experimental results show that using LoRA-CLIP alone (without introducing PRD) significantly improves upon the baseline model (without LoRA or PRD), with the average AUC increasing from 82.29% to 87.85%, demonstrating that the CLIP pre-trained model, after fine-tuning with LoRA, can extract more effective forgery detection features. Further introducing PRD to reconstruct the difference feature further increases the average AUC from 87.85% to 91.07%, achieving consistent performance improvements across four cross-dataset test sets. This verifies that the reconstructed difference detection mechanism proposed in this invention can effectively capture discriminative clues related to forgery, effectively complementing existing static feature extraction methods. These ablation results are consistent with... Figure 3 Consistent with quantitative analysis, the reconstructed differential features themselves have a strong ability to distinguish between classes, and fusing them with the original semantic features can significantly improve detection performance.

[0038] In this embodiment, cross-dataset experiments and ablation experiments verify from an engineering perspective that the detection method designed based on the above findings can indeed achieve superior generalization performance and stable improvement, thus completing a complete closed loop from phenomenon discovery to mechanism verification, and then to method design and performance verification.

[0039] Although the specific embodiments of the invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the invention. Based on the technical solutions of the invention, various modifications or variations that can be made by those skilled in the art without creative effort are still within the scope of protection of the invention.

Claims

1. A method for detecting deepfake face images based on reconstruction difference detection, characterized in that, Includes the following steps: S1. Obtain the image to be detected, and perform detection, alignment, cropping, and size normalization on the face regions in the image to obtain the input image. ; S2. The controlled reconstruction module inputs the input image to the freeze parameters and generates a reconstructed image that is semantically consistent with the input image. ; S3, Input Image and reconstructed image Pixel-level residual statistics and multi-step iterative reconstruction dynamic analysis were performed, and statistical significance tests were conducted to determine whether the difference in response between real and fake images during the reconstruction process objectively exists. S4. Input image With reconstructed image Each input is a feature extractor with shared parameters, and the corresponding output is the original feature vector. and reconstructing feature vectors The feature extractor uses the CLIP ViT-L / 14 visual encoder as the backbone network. LoRA low-rank adaptation modules are injected into the self-attention layers of the backbone network, and the original parameters of the backbone network are kept frozen. S5. Convert the original feature vector and reconstructing feature vectors Input the difference construction module, output the reconstructed difference feature vector. ; S6. Convert the original feature vector and reconstructing the difference feature vector The features are concatenated and fused to obtain a fused feature vector. ; fuse feature vectors Input a classification header and output the predicted probability that the image to be detected is a fake image; S7. During the model training phase, the binary cross-entropy loss is calculated based on the predicted probability and the corresponding real label. The LoRA parameter in the feature extractor and all parameters of the classification head are updated through backpropagation of the loss. During the model inference phase, the deep forgery detection result is output based on the predicted probability.

2. The method for detecting deepfake face images based on reconstruction difference detection according to claim 1, characterized in that, Step S1 is as follows: If the input is video data, then the video is first extracted frame by frame to obtain the image to be detected; The face detection algorithm is used to locate the face region in the image to be detected. Geometric alignment is performed based on facial key points. The aligned face region is then cropped and scaled to a uniform size to obtain the input image. .

3. The method for detecting deepfake face images based on reconstruction difference detection according to claim 1, characterized in that, The controlled reconstruction module in step S2 is a parameter-frozen variational autoencoder; The process of generating a reconstructed image is specifically as follows: taking the input image... The encoder maps the image to the latent space, outputting the encoded mean and encoded variance. Only the encoded mean is used as a deterministic latent variable input to the decoder to generate the reconstructed image. .

4. The method for detecting deepfake face images based on reconstruction difference detection according to claim 1, characterized in that, Step S3 is as follows: (1) Input image and reconstructed image Perform pixel-level subtraction on corresponding channels, and take the absolute value of the difference to obtain the absolute difference image. By analyzing the input image and reconstructed image The residuals of each color channel are averaged to calculate the difference heatmap. , in, Represents the RGB color channels. , Table of RGB color channel indexes, Indicates the first Input image with 1 channel, Indicates the first Reconstructed image of each channel, Indicates taking the absolute value; Preset the region of interest for the face, based on the input image Reconstructing images and difference heatmap Calculate the mean absolute error and high response residual density within the region of interest (ROI) of the face; define a spatial binary mask for extracting the ROI. , Indicates altitude, Represents width, spatial binary mask Limited to the input image The core facial area; Mean Absolute Error and high response residual density The calculation formula is as follows: , , in, Represents pixels, This represents the total number of valid pixels within the face mask. Indicates the total number of channels. Indicates the channel index. Indicates the input image The Middle Each channel pixel The value at that location, Represents the reconstructed image The Middle Each channel pixel The value at that location, For indicator functions, Heatmap of differences medium pixel The value at that location, High response threshold; (2) Input image Perform multi-step iterative reconstruction, letting the initial input... , No. The next iteration projection image is defined as , ,in Indicates the controlled reconfiguration module; calculates the first... Mask during the next iteration Iterative residual energy within : ; The residual energies from multiple consecutive iterations are concatenated in the order of iteration to form a dynamic residual feature vector; (3) Calculate the mean absolute error based on the real image set and the fake image set respectively. and high response residual density The sample mean, sample variance, and sample size were analyzed for significant differences using the Welch test. Let the sample mean of the real image set be... l. The sample variance is Sample size is The sample mean of the forged image set is The sample variance is Sample size is Calculate the test statistic With Satterthwaite's Degree of Freedom : , ; Let the null hypothesis be... The real image and the fake image have the same expected reconstruction residual; Under the original hypothesis Below, based on the degrees of freedom students Distribution calculation two-sided value: , in, express value; Indicates the null hypothesis Under the condition that it holds true, it follows the order of degrees of freedom. students The test statistic of the distribution of a random variable; Represents a probability function; This represents a given condition in conditional probability; If the calculation yields If the value is less than the preset significance level, it is determined that there is a significant difference between the real image and the fake image in the reconstruction residual, that is, the objective existence of the reconstruction difference is verified. Finally, the mean absolute error High response residual density The dynamic residual eigenvectors and statistical test results are used together as the output of this step to support the determination of the existence of differences.

5. The method for detecting deepfake face images based on reconstruction difference detection according to claim 1, characterized in that, Step S4 is as follows: LoRA low-rank adaptation module is injected into the query projection layer and value projection layer of the Transformer self-attention layer in the backbone network; Let the original projection matrix in the self-attention layer be... The projection matrix after LoRA injection satisfy: , , in, Represents a dimension reduction matrix. ; Represents an upgraded matrix. ; Denotes the rank of LoRA This represents the input and output feature dimensions of the projection layer, and .

6. The method for detecting deepfake face images based on reconstruction difference detection according to claim 1, characterized in that, Step S5, the difference construction module, includes two processing schemes, as follows: Option 1: Convert the original feature vector With reconstructing feature vectors By performing element-wise subtraction, we obtain the reconstructed difference feature vector. : , in, , Represents the feature dimension, the original feature vector. and reconstructing feature vectors The same dimensions ; Option 2: Use a shared dimensionality reduction layer to perform dimensionality reduction mapping on the original feature vector and the reconstructed feature vector respectively, to obtain the corresponding low-dimensional original features. With low-dimensional reconstruction features The shared dimensionality reduction layer consists of a linear mapping layer and a batch normalization layer connected in series, and the original feature vector and the reconstructed feature vector share the same set of dimensionality reduction layer parameters. Calculate the self-similarity matrix of the low-dimensional original features respectively. Self-similar matrix of low-dimensional reconstructed features The calculation formula is as follows: , , in, Indicates the feature dimension after dimensionality reduction: The difference between two self-similar matrices is the relation difference matrix. , ; The relational difference matrix is ​​flattened into a one-dimensional vector, normalized, and then input into a multilayer perceptron for mapping, outputting a reconstructed difference feature vector. The calculation formula is as follows: , in, Indicates the flattening operation. Representation layer normalization, This represents a multilayer perceptron mapping.

7. The method for detecting deepfake face images based on reconstruction difference detection according to claim 1, characterized in that, Step S6 is as follows: The classification head consists of a first linear layer, a batch normalization layer, a ReLU activation layer, a Dropout layer, a second linear layer, and a Softmax layer connected in series. fuse feature vectors The predicted probability of the image to be detected being a fake image is obtained from the input classification header. The calculation formula is as follows: , in , representing a linear mapping layer; Indicates the batch normalization layer; Represents a nonlinear activation function; Indicates a randomly deactivated layer; This represents the normalized classification function; This represents the predicted probability corresponding to the counterfeit category.

8. The method for detecting deepfake face images based on reconstruction difference detection according to claim 1, characterized in that, Binary cross-entropy loss in step S7 Specifically as follows: , in, Indicates the true label, , Indicates a forged image. Represents a real image; This represents the natural logarithm function.