Artifact decomposition based false image detection method

By employing a fake image detection method based on artifact decomposition, and utilizing a three-branch artifact-aware encoder and a cross-dimensional fusion module, the method solves the problem of unified processing of various fake content detection methods in existing technologies, and achieves efficient and interpretable fake image detection.

CN122289773APending Publication Date: 2026-06-26SICHUAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SICHUAN UNIV
Filing Date
2026-03-30
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

Existing fake content detection systems cannot handle multiple types of fake content within a unified framework, resulting in low detection efficiency, complex system deployment, difficulties in knowledge transfer, limited generalization capabilities, and complex maintenance and updates.

Method used

A fake image detection method based on artifact decomposition is adopted. The method uses a three-branch artifact perception encoder with scene consistency, imaging authenticity and signal naturalness, combined with a cross-dimensional gating collaborative fusion module and a classification head for end-to-end training. A unified artifact embedding representation is generated by using a multi-objective joint loss function and a weakly supervised adaptive course learning mechanism.

Benefits of technology

It enables the simultaneous handling of multiple fake content detection tasks within a single model framework, reducing deployment complexity, improving cross-domain generalization ability and robustness, and providing interpretable detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122289773A_ABST
    Figure CN122289773A_ABST
Patent Text Reader

Abstract

This invention discloses a fake image detection method based on artifact decomposition. The method includes acquiring the image to be detected and inputting it into a trained three-branch artifact-aware encoder (scene consistency branch, imaging realism branch, and signal naturalness branch) for feature extraction, resulting in three artifact reflection maps. A trained cross-dimensional gated collaborative fusion module is used to fuse the features of the three artifact reflection maps, generating a unified artifact embedding representation. This unified artifact embedding representation is then input into a trained classification head, and the label classification head of the classification head outputs the true / false prediction probability of the image to be detected. This fake image detection method solves the problems of existing technologies, such as inability to uniformly detect multi-source heterogeneous fake images, poor generalization ability, and susceptibility to overfitting.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image detection technology, and specifically to a method for detecting fake images based on artifact decomposition. Background Technology

[0002] With the rapid development of deep learning and generative artificial intelligence technologies, the generation and dissemination of fake content has become increasingly easier, posing a serious threat to social information security, judicial evidence collection, and media credibility. Currently, the field of fake content detection mainly includes four important research directions: deepfake detection, document forgery detection (DOC), AI-generated content recognition (AIGC), and image tampering detection and localization (IMDL-Benco).

[0003] In practical applications, such as in the field of judicial evidence collection, investigators may need to deal with multiple types of false evidence simultaneously: forged surveillance videos (Deepfake), altered contract documents (DOC), AI-generated fake photos (AIGC), and partially modified image evidence (IMDL). However, existing detection systems typically can only detect a single type of false content and cannot handle multiple types of false content within a unified framework, resulting in low detection efficiency and complex system deployment.

[0004] Although these four areas all fall under the category of fake content detection, they exhibit a clear "domain silo" phenomenon in practical applications. Each domain has its own independent datasets, evaluation criteria, and detection models. Researchers typically design specialized algorithms for single tasks, lacking a unified detection framework. This fragmentation leads to the following problems: 1. High model deployment cost: In practical applications, multiple independent models need to be deployed to cover different detection needs, and each model requires independent computing resources and storage space. For example, a complete forensic system may need to deploy a Deepfake detection model, a document forgery detection model, an AIGC recognition model, and an image tampering detection model at the same time, which will increase the system resource consumption many times over and make it difficult to deploy on resource-constrained mobile devices or edge computing devices; 2. Difficulty in knowledge transfer: Detection experience and feature representations across different domains cannot be effectively shared, leading to redundant research and development and wasted resources. For example, the generated trace features learned in Deepfake detection are similar to the generated model fingerprints in AIGC recognition, but due to differences in model architecture and training methods, this common knowledge cannot be transferred and reused across different tasks. 3. Limited generalization ability: Models trained for a single task show a significant drop in detection performance when faced with cross-domain or mixed-type fake content. For example, a model trained on a specific Deepfake dataset will have a significantly lower detection accuracy when faced with new generation methods or AIGC-generated face images, requiring data collection and model retraining. 4. Complex maintenance and updates: Maintaining, updating, and optimizing multiple independent models requires significant human and material resources, resulting in high system complexity. When a generation technology is updated, multiple detection models need to be updated separately, increasing the workload and technical difficulty of system maintenance.

[0005] Therefore, there is an urgent need for a unified detection method that can simultaneously handle multiple domain-specific fake content detection tasks, such as Deepfake detection, document forgery detection, AIGC recognition, and image tampering detection, within a single model framework. This would break down domain silos, enable knowledge sharing and collaborative optimization, reduce system deployment and maintenance costs, and improve cross-domain generalization capabilities. Summary of the Invention

[0006] To address the aforementioned shortcomings in existing technologies, the artifact decomposition-based fake image detection method provided by this invention solves the problems of existing technologies being unable to uniformly detect multi-source heterogeneous fake images, having poor generalization ability, and being prone to overfitting.

[0007] To achieve the above-mentioned objectives, the technical solution adopted by this invention is as follows: A method for detecting fake images based on artifact decomposition is provided, which includes the following steps: S1. Acquire the image to be detected and input it into the scene consistency branch, imaging authenticity branch and signal naturalness branch of the trained three-branch artifact perception encoder for feature extraction, and obtain three artifact response maps respectively. S2. The trained cross-dimensional gated collaborative fusion module is used to perform feature fusion on the three artifact impact maps to generate a unified artifact embedding representation. S3. The unified artifact embedding representation is input to the trained classification head, and the label classification head of the classification head is used to output the true and false prediction probabilities of the image to be detected.

[0008] Furthermore, the scene consistency branch is a Vision Transformer encoder or a deep convolutional neural network, the imaging realism branch is a convolutional neural network initialized based on an SRM filter or a convolutional neural network initialized based on a learnable high-pass filter bank, and the signal naturalness branch is a multi-scale convolutional neural network.

[0009] Furthermore, an end-to-end joint training approach is adopted to train the three modules: the three-branch artifact perception encoder, the cross-dimensional gating collaborative fusion module, and the classification head. The training is then optimized based on a weakly supervised adaptive course learning mechanism. The training method includes: T1. Obtain the training set. Each training sample in the training set includes an image, image-level true and false labels, and weakly supervised localization pseudo labels. T2. Input the training samples into the three-branch artifact perception encoder to obtain three artifact images, and then input them into the cross-dimensional gated collaborative fusion module to obtain fused features. T3. Input the fused features into the classification head, and use the label classification head and heat map classification head of the classification head to obtain the image-level prediction probability and the forgery confidence heat map of the training samples, respectively. T4. Calculate the spatial entropy of the fake confidence heatmap as the sample difficulty, and use an exponential decay sampling strategy to dynamically schedule the courses based on the sample difficulty to obtain the sampling probability of each training step. T5. The multi-objective joint loss function is calculated using image-level prediction probability and fake confidence heatmap, and the loss function is weighted by sample difficulty weight. Then, the weighted loss function is used for backpropagation to update all learnable parameters of the three modules. T6. Repeat steps T2 to T5 until the model consisting of the three modules converges.

[0010] Furthermore, the expression for calculating the sample difficulty is: in, for The sample difficulty of the corresponding sample; To fabricate confidence heatmaps; for The xth spatial position; The expression for calculating the sampling probability is: in, For the first The sampling probability of each training step; It is an exponential function; For the first Temperature parameters for each training step; The initial temperature; The attenuation coefficient is... <1.

[0011] Furthermore, the expression for the multi-objective joint loss function is: , in, For multi-objective joint loss function; For classification loss; To locate the loss; To incorporate sparse regularization terms; and y is the loss weight hyperparameter; y is the ground truth label; p is the image-level prediction probability. The pseudo-labels are generated under weak supervision, M is the fake confidence heatmap; H and W are the height and width of the fake confidence heatmap, respectively; For pixels Regularization terms encourage gating weights; It is an extremely small positive constant to prevent for This leads to floating-point precision issues.

[0012] Furthermore, the expression for weighting the loss function using sample difficulty weights is as follows: , in, The weighted loss function; Batch size; , and These are the classification loss, localization loss, and fusion sparsity regularization term for the nth training sample, respectively. The sample difficulty weights for the nth training sample; A heatmap of the fake confidence scores for the nth training sample; for The sample difficulty of the corresponding sample; For the first Temperature parameters for each training step; It is an exponential function.

[0013] The beneficial effects of the above technical solution are as follows: This solution, through the collaborative optimization of a multi-objective joint loss function, enables the three-branch artifact-aware encoder, the cross-dimensional gating collaborative fusion module, the classification head, and the auxiliary localization head to be trained end-to-end under a unified framework, with each module having a clear function and cooperating with each other; through a differentiated loss constraint mechanism, the three-branch artifact-aware encoder focuses on extracting artifact features of different dimensions (achieving functional differentiation through different network architectures), and the cross-dimensional gating collaborative fusion module focuses on learning sparse adaptive fusion weights (through... (Regularization constraints), the classification head focuses on true / false discrimination (through...) (Supervision), the auxiliary positioning head focuses on generating fake confidence heatmaps (through...) (Supervision) Each module is collaboratively optimized in end-to-end training; the heatmap entropy-driven adaptive course learning mechanism realizes a progressive training strategy of "from easy to difficult". In the early stage of training, easy samples are processed first, and high-entropy difficult samples are gradually introduced as the model converges, which effectively alleviates the interference of weak supervision noise labels and improves generalization ability and robustness.

[0014] Furthermore, when the scene consistency branch is a Vision Transformer encoder, its method for extracting pseudo-effect maps includes: The semantic illusion intensity M1 and local geometric consistency of the image to be detected are extracted using a feature extractor. The shadow angle deviation M2 and structural compliance score M3 are calculated, and CAM backpropagation, Sigmoid normalization, and Gaussian kernel diffusion are applied to M1, M2, and M3 respectively to generate feature maps. ; For feature maps Concatenate the components and use them as input tensors. Input the Transform-vit module to input tensors Divided into The patch is generated, and the dimensions are mapped from 1024 to 192 through linear projection to obtain the token embedding, and position encoding is added. The feature input with added position encoding is sequentially connected to 12 Transformer Blocks to obtain the Token embedding, which is then processed by the Attention Rollout mechanism to output a high-dimensional pseudo-effect map with scene consistency. Each Transformer Block comprises, in sequence, a layer normalization layer, a multi-head self-attention mechanism module, another layer normalization layer, and a feedforward neural network. The feedforward neural network includes two linear transformation layers and a GELU activation function. The dimensionality transformation process is as follows: Furthermore, residual connections are introduced after the feedforward neural network.

[0015] The beneficial effects of the above technical solution are as follows: The scene consistency branch mainly focuses on capturing the multimodal contradictions of images at the physical and logical levels. By integrating semantic, depth geometry, lighting and shadow and text layout information, it can effectively identify unavoidable physical law violations (such as lighting and shadow logic errors), geometric perspective distortion and unnatural artifacts in structural layout in the generated image.

[0016] Furthermore, when the imaging realism branch is a convolutional neural network initialized with an SRM filter, its method for extracting the pseudo-image distortion map includes: A feature extractor is used to extract the PRNU noise fingerprint, local linear prediction residual, and CFA interpolation trace from the image to be detected. These three feature maps are then normalized, subjected to autocorrelation analysis, and processed by channel response, respectively, and finally concatenated to form the input tensor. ; Input tensor The input convolutional neural network is initialized using an SRM filter, followed by four sequentially connected convolutional processing blocks to obtain a pseudo-image of the imaging authenticity fused with multi-source evidence features. Each convolutional processing block includes sequentially connected... The convolutional layers, batch normalization layers, and ReLU activation function.

[0017] The beneficial effects of the above technical solution are as follows: the imaging realism branch mainly focuses on capturing the underlying physical traces left over from the image generation or tampering process, including the destruction of the sensor noise mode (PRNU), CFA interpolation anomalies caused by resampling, and changes in the statistical properties of local textures. These underlying clues are crucial for detecting subtle pixel-level artifacts, unnatural high-frequency noise residues, and the heterogeneity of image sources in deepfakes.

[0018] Furthermore, the method for extracting the pseudo-effect map from the natural branch of the signal includes: Local spectral anomalies, fourth-order Laplacian difference features, and neighborhood quantization step size consistency are extracted from the image to be detected. These features are then processed by bilinear interpolation, ACF spatial function processing, and mapping interpolation to generate feature maps. And concatenate them into the input tensor ; feature tensor Input a three-way parallel convolutional module, and use respectively... , and The convolutional kernels are used for multi-receptive field feature extraction, and each branch includes batch normalization and ReLU activation; The features extracted by the three convolutional modules are concatenated and then input into a single... The convolutional layer performs channel fusion and dimensionality reduction to obtain a pseudo-effect map of the signal's naturalness.

[0019] The beneficial effects of the above technical solution are as follows: the signal naturalness branch mainly focuses on unnatural traces at the image signal level, such as frequency domain artifacts (e.g., checkerboard effect) generated by generative models (GANs / Diffusion), resampling traces caused by nonlinear operations, and double JPEG quantization anomalies caused by multiple compressions. These signal-level clues are often difficult for the human eye to detect, but they provide strong statistical evidence in distinguishing naturally captured images from images synthesized or tampered with by algorithms.

[0020] Furthermore, step S2 further includes: S21. The artifacts extracted from the scene consistency branch, imaging realism branch, and signal naturalness branch should be visualized. , and Perform separately Convolutional dimensionality reduction yields the pseudo-impact graph after dimensionality reduction. , and ; S22. The pseudo-influence diagram after dimensionality reduction should be used. , and By concatenating along the channel dimension, the concatenated features are obtained. ; splicing features Input a multi-head self-attention mechanism and calculate the credibility weights in three dimensions: in, , and Positions The weights of scene consistency, image realism, and signal naturalness are considered. MHSA is a multi-head self-attention mechanism; Softmax is a normalization function. S23, According to position Weight at point , and , regarding position The pseudo-effect should be shown in the figure. , and Weighted fusion is performed to generate a unified artifact embedding representation. : in, This means taking all elements along the channel dimension.

[0021] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) This solution is the first to construct a unified detection framework covering four major forgery types: AIGC, Deepfake, image tampering, and document forgery. It achieves "universal detection and isomorphism" of multi-source heterogeneous fake images within a single architecture. Instead of modeling each type of forgery separately, this framework enables the model to learn the common forgery rules across categories during the training phase by sharing the underlying representation and high-level discrimination mechanism. This significantly reduces deployment complexity and provides core algorithmic support for building the next generation of digital content security infrastructure.

[0022] (2) The artifact decoupling mechanism based on the real physical imaging chain proposed in this scheme fundamentally breaks through the path dependence of traditional fake image detection methods on data-driven surface statistical features. By decoupling the complex forgery artifacts into three orthogonal and interpretable essential dimensions—scene consistency, imaging authenticity, and signal naturalness—and through feature orthogonalization design under physical mechanism constraints, it actively suppresses interference signals introduced by non-forgery factors such as content semantics, illumination changes, and compression distortion, thereby accurately focusing on the real imaging physical residuals that cannot be perfectly reproduced by the forgery process. This leap from "phenomenon fitting" to "mechanism modeling" not only significantly improves the model's generalization ability in unknown forgery techniques and open scenarios, but also endows the detection system with interpretability and forward-looking defense potential.

[0023] (3) The "heatmap entropy-driven adaptive course learning mechanism" of this scheme abandons the traditional static difficulty division and innovatively introduces the spatial entropy based on the model's own prediction of the forgery confidence heatmap as a dynamic difficulty measurement standard. It perceives the sample complexity in real time and evolves in tandem with the model's current cognitive state, guiding it to learn the essential characteristics of unified artifacts from easy to difficult and from surface to core, rather than simply memorizing the surface statistical patterns of specific data distributions or forgery types. This strategy systematically alleviates the model overfitting problem and significantly improves its robustness and generalization ability in unknown forgery types and cross-domain scenarios.

[0024] (4) This scheme achieves pixel-level adaptive weighted fusion of three orthogonal artifact dimensions through a cross-dimensional gating collaborative fusion mechanism. This mechanism can dynamically adjust the contribution weight of each dimension according to the forgery characteristics of different spatial locations, enhance scene consistency features in semantically contradictory regions, dominate imaging authenticity features in sensor noise abnormal regions, and highlight signal naturalness features near splicing boundaries, thereby achieving content-aware dimension selection and significantly improving the accuracy and robustness of detection.

[0025] (5) This scheme adopts a weakly supervised learning paradigm, which only requires image-level labels and coarse annotations of fake regions to train the model, greatly reducing the cost of data annotation. Through multi-objective joint optimization, the model can not only accurately distinguish between genuine and fake images, but also generate a heatmap of fake confidence, providing users with interpretable detection results and enhancing the credibility and practicality of the system. Attached Figure Description

[0026] Figure 1 This is a flowchart of a fake image detection method based on artifact decomposition.

[0027] Figure 2 This is a network architecture diagram of the model corresponding to the fake image detection method based on artifact decomposition.

[0028] Figure 3Network architecture diagram for the scenario consistency branch.

[0029] Figure 4 Network architecture diagram for the imaging realism branch.

[0030] Figure 5 This is a network architecture diagram representing the natural branch of the signal.

[0031] Figure 6 This is a network architecture diagram for a cross-dimensional gating collaborative fusion module. Detailed Implementation

[0032] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.

[0033] refer to Figure 1 , Figure 1 A flowchart of a fake image detection method based on artifact decomposition is shown; as follows: Figure 1 As shown, the method S includes steps S1 to S3.

[0034] In step S1, the image to be detected is acquired. The image to be detected can be any type of image, such as an AIGC-generated image, a Deepfake face image, a tampered image, or a document image. The image to be detected is input into the scene consistency branch, imaging realism branch, and signal naturalness branch of the trained three-branch artifact-aware encoder for feature extraction, resulting in three artifact response maps.

[0035] In this scheme, the scene consistency branch is a Vision Transformer encoder or a deep convolutional neural network, the imaging realism branch is a convolutional neural network initialized based on SRM filter or a convolutional neural network initialized based on a learnable high-pass filter bank, and the signal naturalness branch is a multi-scale convolutional neural network.

[0036] like Figure 3 As shown, the scene consistency branch employs a combination of multi-source physical feature fusion and the Transform-vit architecture to extract and integrate multi-dimensional consistency features of the image. The methods used by the scene consistency branch to extract pseudo-effect maps include: The semantic illusion intensity M1 and local geometric consistency of the image to be detected are extracted using a feature extractor. The shadow angle deviation M2 and structural compliance score M3 are calculated, and CAM backpropagation, Sigmoid normalization, and Gaussian kernel diffusion are applied to M1, M2, and M3 respectively to generate feature maps. ; For feature maps Concatenate the components and use them as input tensors. Input the Transform-vit module to input tensors Divided into The patch is generated, and the dimensions are mapped from 1024 to 192 through linear projection to obtain the token embedding, and position encoding is added. The feature input with added position encoding is sequentially connected to 12 Transformer Blocks to obtain the Token embedding, which is then processed by the Attention Rollout mechanism to output a high-dimensional pseudo-effect map with scene consistency. Each Transformer Block comprises, in sequence, a layer normalization layer, a multi-head self-attention mechanism module, another layer normalization layer, and a feedforward neural network. The feedforward neural network includes two linear transformation layers and a GELU activation function. The dimensionality transformation process is as follows: Furthermore, residual connections are introduced after the feedforward neural network.

[0037] When the scene consistency branch is a deep convolutional neural network, it still extracts four sets of feature maps—semantic illusion intensity, local geometric consistency, shadow angle deviation, and structural compliance score—and concatenates them into an input tensor when extracting the pseudo-illusion response map. However, the subsequent Transformer Block is replaced with a deep residual convolution module. To compensate for the limitations of convolutional operations in the global receptive field, dilated convolutions or non-local attention modules can be introduced into the network to capture long-range spatial dependencies. This alternative has advantages in computational efficiency and is suitable for deployment scenarios with high inference speed requirements.

[0038] like Figure 4 As shown, the imaging realism branch employs a multi-stream convolutional neural network architecture initialized with an SRM (Steganalysis Rich Model) filter, aiming to mine the invisible low-level imaging fingerprints in images; the method for extracting pseudo-images in the imaging realism branch includes: A feature extractor is used to extract the PRNU noise fingerprint, local linear prediction residual, and CFA interpolation trace from the image to be detected. These three feature maps are then normalized, subjected to autocorrelation analysis, and processed by channel response, respectively, and finally concatenated to form the input tensor. ; Input tensor The input convolutional neural network is initialized using an SRM filter, followed by four sequentially connected convolutional processing blocks to obtain a pseudo-image of the imaging authenticity fused with multi-source evidence features. Each convolutional processing block includes sequentially connected... The convolutional layers, batch normalization layers, and ReLU activation function.

[0039] When the image realism branch uses a convolutional neural network initialized with a learnable high-pass filter bank, a set of learnable convolutional kernels (such as 32 5×5 convolutional kernels) is initialized. During training, the ability to extract noise residuals is guaranteed by constraining the sum of the convolutional kernel weights to zero (high-pass constraint), while allowing the network to adaptively adjust the shape and response characteristics of the filter according to task requirements. Alternatively, other filtering schemes with noise residual extraction capabilities can be used to replace the SRM filter, such as Bayesian constrained convolution (BayarConv), wavelet transform high-frequency subband decomposition, etc. As long as they can effectively suppress image content information and preserve the underlying noise residual features, they can be used as the front-end feature extraction module of the image realism branch.

[0040] like Figure 5 As shown, the signal naturalness branch employs a multi-scale parallel convolutional neural network architecture, aiming to capture subtle anomalies in the frequency domain and statistical signal level of images; the method for extracting pseudo-effect maps by the signal naturalness branch includes: Local spectral anomalies, fourth-order Laplacian difference features, and neighborhood quantization step size consistency are extracted from the image to be detected. These features are then processed by bilinear interpolation, ACF spatial function processing, and mapping interpolation to generate feature maps. And concatenate them into the input tensor ; feature tensor Input a three-way parallel convolutional module, and use respectively... , and The convolutional kernels are used for multi-receptive field feature extraction, and each branch includes batch normalization and ReLU activation; The features extracted by the three convolutional modules are concatenated and then input into a single... The convolutional layer performs channel fusion and dimensionality reduction to obtain a pseudo-effect map of the signal's naturalness.

[0041] Before extracting local spectral anomalies, fourth-order Laplacian difference features, and neighborhood quantization step size consistency, this scheme preferably uses Discrete Cosine Transform (DCT) or Fast Fourier Transform (FFT) to transform the image to be detected to the frequency domain. FFT can provide frequency domain information in terms of both amplitude and phase, and has a richer frequency domain representation capability than DCT. Subsequently, multi-scale convolutional neural networks can be used to further extract and encode the frequency domain features.

[0042] In step S2, the trained cross-dimensional gated collaborative fusion module is used to perform feature fusion on the three artifact embedding images to generate a unified artifact embedding representation; the network architecture of the cross-dimensional gated collaborative fusion module can be referenced. Figure 6 .

[0043] In implementation, the preferred step S2 of this scheme further includes: S21. The artifacts extracted from the scene consistency branch, imaging realism branch, and signal naturalness branch should be visualized. , and Perform separately Convolutional dimensionality reduction uniformly reduces the number of channels to 256, yielding the dimensionality-reduced pseudo-impact graph. , and ; S22. The pseudo-influence diagram after dimensionality reduction should be used. , and By concatenating along the channel dimension, the concatenated features are obtained. Its dimensions are .

[0044] splicing features Input a multi-head self-attention mechanism and calculate the credibility weights in three dimensions: in, , and Positions The weights of scene consistency, image realism, and signal naturalness are considered. MHSA is a multi-head self-attention mechanism; Softmax is a normalization function.

[0045] The Multi-Head Self-Attention (MHSA) mechanism contains eight attention heads, each with a dimension of 96. Through self-attention computation, the model is able to capture long-range dependencies between different spatial locations and learn the relative importance of the three artifact dimensions at each location.

[0046] In the cross-dimensional gated collaborative fusion module, a cross-attention fusion mechanism or a channel attention fusion mechanism can be used to replace the multi-head self-attention mechanism (MHSA). When the cross-dimensional gated collaborative fusion module adopts the cross-attention mechanism, the three dimensionality-reduced artifact response maps are used as different combinations of Query, Key, and Value, respectively. Through multiple rounds of cross-attention calculation, each artifact dimension can explicitly perceive the information of the other two dimensions and adjust its own feature representation and fusion weight accordingly. Compared with the self-attention mechanism, this scheme can more directly model the interaction relationship between artifact dimensions.

[0047] When the cross-dimensional gated collaborative fusion module employs a channel attention mechanism, it concatenates the three dimensionality-reduced pseudo-effect maps along the channel dimension, compresses the spatial dimension through global average pooling, and then generates a channel-level weight vector by passing it through two fully connected layers and Sigmoid activation. This vector is then used for channel-level weighted fusion of the concatenated features. This approach has lower computational overhead and a simpler structure, making it suitable for scenarios with high computational efficiency requirements.

[0048] S23, According to position Weight at point , and , regarding position The pseudo-effect should be shown in the figure. , and Weighted fusion is performed to generate a unified artifact embedding representation. : in, This means taking all elements along the channel dimension.

[0049] The weighted fusion operation is equivalent to channel-dimensional weighted fusion. The result not only preserves the spatial resolution of the original features but also achieves content-aware dimension selection: in semantically contradictory regions, Automatic enhancement; in areas with abnormal sensor noise, Dominant integration; while near the splicing boundary, Significant increase.

[0050] In this scheme, weighted fusion can also adopt a structure similar to Feature Pyramid Network (FPN), performing three-dimensional fusion at multiple resolution scales, and then integrating the fusion results from different scales through top-down paths and lateral connections. This scheme can simultaneously capture artifact feature interactions at different spatial scales and is suitable for processing tampered regions of different scales.

[0051] This solution achieves adaptive weighted fusion of three orthogonal artifact dimensions through a cross-dimensional gating collaborative fusion mechanism. In semantically contradictory regions, scene consistency weights are automatically enhanced; in sensor noise anomalous regions, image realism weights are enhanced. Dominant fusion; while near the splicing boundary, the signal's naturalness weight. Significantly improved. This content-aware dimensional selection mechanism allows the model to dynamically adjust the contribution of each dimension according to the characteristics of different forgery types, thereby improving the accuracy and robustness of detection.

[0052] In step S3, the unified artifact embedding representation is used as the input to the trained classification head, and the label classification head of the classification head is used to output the true or false prediction probability of the image to be detected.

[0053] In one embodiment of the present invention, the classification head includes a label classification head and a thermal classification head, and the label classification head includes a global average pooling layer, a fully connected layer and an output layer; The global average pooling layer compresses the spatial dimension of the unified artifact embedding representation, resulting in a global feature vector: in, for A global feature vector of dimension ; H and W are the height and width of the feature map, respectively; express In spatial location Extract the feature vectors of all channels; Fully connected layers perform non-linear transformations on global feature vectors: in, for Hidden feature vectors of dimension ; This is the weight matrix; For bias terms; For batch normalization; For activation functions; Output layer generates true / false prediction probabilities: in, This represents the probability that the image to be detected is a forgery. For the Sigmoid function; This is the output layer weight vector; This is the output layer bias.

[0054] like Figure 2As shown, an end-to-end joint training method is used to train three modules: the three-branch artifact perception encoder, the cross-dimensional gating collaborative fusion module, and the classification head. The training is then optimized based on a weakly supervised adaptive course learning mechanism. The training method includes: T1. Obtain the training set. Each training sample in the training set includes an image, image-level true and false labels, and weakly supervised localization pseudo labels. T2. Input the training samples into the three-branch artifact perception encoder to obtain three artifact images, and then input them into the cross-dimensional gated collaborative fusion module to obtain fused features. T3. Input the fused features into the classification head, and use the label classification head and heat map classification head of the classification head to obtain the image-level prediction probability and the forgery confidence heat map of the training samples, respectively. T4. Calculate the spatial entropy of the fake confidence heatmap as the sample difficulty: in, for The sample difficulty of the corresponding sample; To fabricate confidence heatmaps; for The xth spatial location.

[0055] Based on the sample difficulty, an exponential decay sampling strategy is used for dynamic course scheduling to obtain the sampling probability for each training step: in, For the first The sampling probability of each training step; It is an exponential function; For the first Temperature parameters for each training step; The initial temperature; The attenuation coefficient is... <1.

[0056] T5. The multi-objective joint loss function is calculated using image-level prediction probability and fake confidence heatmap, and the loss function is weighted by sample difficulty weight. Then, the weighted loss function is used for backpropagation to update all learnable parameters of the three modules. T6. Repeat steps T2 to T5 until the model consisting of the three modules converges.

[0057] In one embodiment of the present invention, the expression for the multi-objective joint loss function is: , in, For multi-objective joint loss function; For classification loss; To locate the loss; To incorporate sparse regularization terms; and The loss weight hyperparameters control the optimization intensity of localization capability and fusion sparsity, respectively; y is the ground truth label; p is the image-level prediction probability. The pseudo-labels are generated under weak supervision, M is the fake confidence heatmap; H and W are the height and width of the fake confidence heatmap, respectively; For pixels Regularization terms encourage gating weights; It is a very small positive constant, preventing for This leads to floating-point precision issues.

[0058] The expression for weighting the loss function using sample difficulty weights is as follows: , in, The weighted loss function; Batch size; , and These are the classification loss, localization loss, and fusion sparsity regularization term for the nth training sample, respectively. The sample difficulty weights for the nth training sample; A heatmap of the fake confidence scores for the nth training sample; for The sample difficulty of the corresponding sample; For the first Temperature parameters for each training step; It is an exponential function.

[0059] In the multi-objective joint loss function, the classification loss is used to supervise the real and fake discrimination ability of the classification head. This loss is propagated backward through the gradient of the classification head to the cross-dimensional gated collaborative fusion module and the three-branch artifact perception encoder, driving the entire network to learn artifact representations that are conducive to real and fake discrimination. The localization loss is propagated forward through the gradient of the auxiliary localization head and back to the cross-dimensional gated collaborative fusion module and the three-branch artifact perception encoder, which prompts the network to learn artifact features with spatial discrimination ability. At the same time, the generated heat map M is used to calculate the sample difficulty metric H(M). Incorporating sparse regularization terms encourages gate weights , and By activating only a few dimensions at each pixel location, the interpretability and stability of the model are improved. This loss only applies to the parameters of the multi-head self-attention mechanism in the cross-dimensional gated collaborative fusion module and does not affect the parameter updates of the three-branch artifact perception encoder and the classification head.

[0060] In this scheme, the scene consistency branch, image realism branch, and signal naturalness branch share the classification loss during training. and positioning loss The supervisory signal is used to extract artifact features from different representation spaces of the input image (spatial semantics, noise residuals, frequency statistics) by adopting different network architectures (Vision Transformer, CNN initialized with SRM filter, multi-scale CNN), naturally forming functional differentiation; the parameters of the three branches are updated independently, without sharing weights, and each learns its own artifact pattern.

[0061] To verify the effectiveness of the fake image detection method based on artifact decomposition provided in this embodiment, experiments were conducted on multiple public datasets.

[0062] The training dataset uses the OpenMMSec open multimodal AI security benchmark dataset, containing 100,000 samples covering four major forgery types. The test datasets include: AIGC detection: DiffusionForensics, Chameleon; Deepfake detection: FaceForensics++, DF40; Image tampering detection: IMD2020, Autosplice; Document forgery detection: RealTextManipulation, T-SROIE.

[0063] The experiment adopts a three-level progressive evaluation protocol: (1) Intra-domain testing: Test on the OpenMMSec validation set to verify the basic fitting ability of the model. The validation results can be found in Table 1. (2) Cross-domain testing: After training on OpenMMSec, no fine-tuning was performed, and the results were evaluated directly on 8 public test sets. The test results can be found in Table 2. (3) Leave-one-out test: Each time, a certain type of fake data is excluded for training, and the model's ability to identify zero samples of completely unseen fake types is tested. The test results can be found in Table 3.

[0064] For ease of description, the three-branch artifact decomposition architecture of this scheme is referred to as ADSNet (Artifact Decomposition Structure Network) in Tables 1-3.

[0065] For in-domain testing, Table 1 lists various existing models and ADSNet provided in this application in the first column. The various methods are compared on five metrics: Precision, Recall, Accuracy, F1-score, and AUC.

[0066] Table 1. In-domain test results (OpenMMSec validation set test results) As shown in Table 1, under the domain setting, the proposed ADSNet outperforms existing methods in all five metrics: Precision, Recall, Accuracy, F1-score, and AUC. Its core advantage lies in enhancing sensitivity to subtle forgery traces by decomposing artifacts. Especially in high-fidelity generated images, Recall is significantly improved, indicating a lower false negative rate; AUC reaches 94.5%, an improvement of +8.1% over the current best baseline, Effort (86.4%); F1-score is 92.6%, an improvement of +7.4% over the best, UnivFD (85.2%); and all metrics (Precision / Recall / Accuracy) exceed 90%, demonstrating that the proposed ADSNet possesses extremely strong fitting ability and discrimination accuracy in the same distribution scenario.

[0067] The comparison in Table 1 shows that the three-branch artifact decomposition architecture can effectively extract forgery clues in three orthogonal dimensions: scene consistency, imaging authenticity, and signal naturalness. It can also generate a highly discriminative unified representation through gating fusion, which is significantly better than single-branch or implicit feature learning methods.

[0068] For cross-domain testing, after training on OpenMMSec, no fine-tuning is performed, and the tests are directly evaluated on the eight public test sets shown in the first row of Table 2; the first column of Table 2 lists various existing models and the ADSNet provided in this application.

[0069] Table 2 Cross-domain test results As shown in Table 2, ADSNet in this scheme leads in all four domains, with significant improvements, especially in Doc and IMDL. All methods show performance degradation in cross-domain scenarios, but ADSNet consistently maintains its lead, achieving an average AUC of 87.1%, a 6.0 percentage point improvement over the second-best method, Effort (81.1%). In the most challenging IMDL domain (Autosplice: 84.4%, IMD2020: 85.1%), ADSNet outperforms Effort by +7.1~8.0%, demonstrating its stronger sensitivity to local editing-type forgeries. Compared to TruFor (dependent on noise maps) and UnivFD (general feature extraction), AD-UniFD explicitly models multi-scale artifact residuals (including JPEG block artifacts, interpolation artifacts, and generator fingerprints), enabling it to capture common forgery signals even in unknown domains.

[0070] The data shown in Table 2 demonstrates that the ADSNet proposed in this scheme exhibits excellent generalization stability: the artifact decoupling mechanism enables the model to no longer rely on specific forged priors to a certain extent, but instead learns the common physical imaging residual laws across domains, fundamentally solving the "domain silo" problem.

[0071] For the leave-one-out test, each time a certain type of fake data is removed from OpenMMSec for training, the model is trained with the remaining three types plus real images, and then evaluated on the test subset of the removed domain to test the model's ability to identify zero samples of completely unseen fake types. The test results are shown in Table 3.

[0072] Table 3. Results of the Leave-One-out Method Test As shown in Table 3, in the leave-one-out test (completely excluding a certain type of forgery during training), ADSNet still maintains excellent detection performance on completely unseen forgery categories. ADSNet significantly outperforms existing methods in all cases. Even if a certain type of forgery mechanism has never been seen before (e.g., AIGC images have never been encountered during training), ADSNet can still achieve an AUC of 81.5% in this unknown domain, which is significantly better than all existing baseline methods.

[0073] The AIGC domain achieves the highest zero-shot performance (87.5% AUC) due to the strong cross-domain generality of its typical artifacts, which can be effectively captured by the model's frequency domain decomposition module. Doc is the most challenging, but ADSNet still surpasses the current strongest baseline Effort (78.1%) with an AUC of 80.9%, leading by +2.8%, demonstrating its high sensitivity to local anomalous signals. The average AUC reaches 83.4%, an improvement of +5.6% over the current strongest baseline Effort (77.8%).

[0074] Even after training away from Deepfake types, the ADSNet model still achieves an AUC of 83.8% on Deepfake test sets such as FF++ / DFDC, indicating that its scene consistency modeling mechanism can be effectively transferred to new face generation paradigms.

[0075] The comparison shows that the proposed solution uses an adaptive learning mechanism to guide the ADSNet model to focus on essential artifacts rather than superficial statistical patterns, enabling it to make reliable judgments based on physical consistency criteria when faced with unknown forgery techniques, thus possessing a forward-looking defense capability.

Claims

1. A method for detecting fake images based on artifact decomposition, characterized in that, Including the following steps: S1. Acquire the image to be detected and input it into the scene consistency branch, imaging authenticity branch and signal naturalness branch of the trained three-branch artifact perception encoder for feature extraction, and obtain three artifact response maps respectively. S2. The trained cross-dimensional gated collaborative fusion module is used to perform feature fusion on the three artifact impact maps to generate a unified artifact embedding representation. S3. The unified artifact embedding representation is input to the trained classification head, and the label classification head of the classification head is used to output the true and false prediction probabilities of the image to be detected.

2. The fake image detection method according to claim 1, characterized in that, The scene consistency branch is a Vision Transformer encoder or a deep convolutional neural network, the imaging realism branch is a convolutional neural network initialized based on an SRM filter or a convolutional neural network initialized based on a learnable high-pass filter bank, and the signal naturalness branch is a multi-scale convolutional neural network.

3. The fake image detection method according to claim 1, characterized in that, An end-to-end joint training approach is used to train three modules: a three-branch artifact-aware encoder, a cross-dimensional gating collaborative fusion module, and a classification head. The training is then optimized based on a weakly supervised adaptive learning mechanism. The training method includes: T1. Obtain the training set. Each training sample in the training set includes an image, image-level true and false labels, and weakly supervised localization pseudo labels. T2. Input the training samples into the three-branch artifact perception encoder to obtain three artifact images, and then input them into the cross-dimensional gated collaborative fusion module to obtain fused features. T3. Input the fused features into the classification head, and use the label classification head and heat map classification head of the classification head to obtain the image-level prediction probability and the forgery confidence heat map of the training samples, respectively. T4. Calculate the spatial entropy of the fake confidence heatmap as the sample difficulty, and use an exponential decay sampling strategy to dynamically schedule the courses based on the sample difficulty to obtain the sampling probability of each training step. T5. The multi-objective joint loss function is calculated using image-level prediction probability and fake confidence heatmap, and the loss function is weighted by sample difficulty weight. Then, the weighted loss function is used for backpropagation to update all learnable parameters of the three modules. T6. Repeat steps T2 to T5 until the model consisting of the three modules converges.

4. The fake image detection method according to claim 3, characterized in that, The expression for calculating sample difficulty is: in, for The sample difficulty of the corresponding sample; To fabricate confidence heatmaps; for The xth spatial position; The expression for calculating the sampling probability is: in, For the first The sampling probability of each training step; It is an exponential function; For the first Temperature parameters for each training step; The initial temperature; The attenuation coefficient is... <1.

5. The fake image detection method according to claim 3, characterized in that, The expression for the multi-objective joint loss function is: , in, For multi-objective joint loss function; For classification loss; To locate the loss; To incorporate sparse regularization terms; and y is the loss weight hyperparameter; y is the ground truth label; p is the image-level prediction probability. The pseudo-labels are generated under weak supervision, M is the fake confidence heatmap; H and W are the height and width of the fake confidence heatmap, respectively; For pixels Regularization terms encourage gating weights; It is an extremely small positive constant to prevent for This leads to floating-point precision issues.

6. The fake image detection method according to claim 5, characterized in that, The expression for weighting the loss function using sample difficulty weights is as follows: , in, The weighted loss function; Batch size; , and These are the classification loss, localization loss, and fusion sparsity regularization term for the nth training sample, respectively. The sample difficulty weights for the nth training sample; A heatmap of the fake confidence scores for the nth training sample; for The sample difficulty of the corresponding sample; For the first Temperature parameters for each training step; It is an exponential function.

7. The fake image detection method according to claim 2, characterized in that, When the scene consistency branch is a Vision Transformer encoder, its method for extracting pseudo-effect maps includes: The semantic illusion intensity M1 and local geometric consistency of the image to be detected are extracted using a feature extractor. The shadow angle deviation M2 and structural compliance score M3 are calculated, and CAM backpropagation, Sigmoid normalization, and Gaussian kernel diffusion are applied to M1, M2, and M3 respectively to generate feature maps. ; For feature maps Concatenate the components and use them as input tensors. Input the Transform-vit module to input tensors Divided into The patch is generated, and the dimensions are mapped from 1024 to 192 through linear projection to obtain the token embedding, and position encoding is added. The feature input with added position encoding is sequentially connected to 12 Transformer Blocks to obtain the Token embedding, which is then processed by the Attention Rollout mechanism to output a high-dimensional pseudo-effect map with scene consistency. Each Transformer Block comprises, in sequence, a layer normalization layer, a multi-head self-attention mechanism module, another layer normalization layer, and a feedforward neural network. The feedforward neural network includes two linear transformation layers and a GELU activation function. The dimensionality transformation process is as follows: Furthermore, residual connections are introduced after the feedforward neural network.

8. The fake image detection method according to claim 2, characterized in that, When the imaging realism branch is a convolutional neural network initialized with an SRM filter, the method for extracting the pseudo-image rendering map includes: A feature extractor is used to extract the PRNU noise fingerprint, local linear prediction residual, and CFA interpolation trace from the image to be detected. These three feature maps are then normalized, subjected to autocorrelation analysis, and processed by channel response, respectively, and finally concatenated to form the input tensor. ; Input tensor The input convolutional neural network is initialized using an SRM filter, followed by four sequentially connected convolutional processing blocks to obtain a pseudo-image of the imaging authenticity fused with multi-source evidence features. Each convolutional processing block includes sequentially connected... The convolutional layers, batch normalization layers, and ReLU activation function.

9. The fake image detection method according to claim 2, characterized in that, The method for extracting pseudo-effect maps from the natural branches of the signal includes: Local spectral anomalies, fourth-order Laplacian difference features, and neighborhood quantization step size consistency are extracted from the image to be detected. These features are then processed by bilinear interpolation, ACF spatial function processing, and mapping interpolation to generate feature maps. And concatenate them into the input tensor ; feature tensor Input a three-way parallel convolutional module, and use respectively... , and The convolutional kernels are used for multi-receptive field feature extraction, and each branch includes batch normalization and ReLU activation; The features extracted by the three convolutional modules are concatenated and then input into a single... The convolutional layer performs channel fusion and dimensionality reduction to obtain a pseudo-effect map of the signal's naturalness.

10. The fake image detection method according to claim 1, characterized in that, Step S2 further includes: S21. The artifacts extracted from the scene consistency branch, imaging realism branch, and signal naturalness branch should be visualized. , and Perform separately Convolutional dimensionality reduction yields the pseudo-impact graph after dimensionality reduction. , and ; S22. The pseudo-influence diagram after dimensionality reduction should be used. , and By concatenating along the channel dimension, the concatenated features are obtained. ; splicing features Input a multi-head self-attention mechanism and calculate the credibility weights in three dimensions: in, , and Positions The weights of scene consistency, image realism, and signal naturalness are considered. MHSA is a multi-head self-attention mechanism; Softmax is a normalization function. S23, According to position Weight at point , and , regarding position The pseudo-effect should be shown in the figure. , and Weighted fusion is performed to generate a unified artifact embedding representation. : in, This means taking all elements along the channel dimension.