Dual-attention-based fake image detection method and device, and electronic device

CN118570484BActive Publication Date: 2026-09-25NAT UNIV OF DEFENSE TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410710716.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-04
Publication Date
2026-09-25
Estimated Expiration
2044-06-04

AI Technical Summary

Technical Problem

[0005]本发明要解决的技术问题是现有技术由于在多特征融合上处理方式过于单一且仅能得到单一固定分辨率的特征图,对伪造图像的检测和定位效果较差,为了解决上述问题,本发明提供一种基于双注意力的伪造图像检测方法、装置及电子设备

Benefits of technology

[0045]在本发明实施例中,为了克服伪造区域不一致带来的不良影响,采用了一个分层渐进网络捕捉图像不同尺度上的伪造伪影而实现检测和定位。本申请实施例依靠双注意力机制将多模态图像特征进行自适应深度融合,而后由多分支交互网络将不同尺度的图像特征进行充分交互,并依靠层次之间的依赖关系改善检测器的性能。此外,还通过提取更敏感的噪声指纹以获得更加显著的伪造区域的伪影特征。通过上述方法,提高了对伪造图像的检测和定位的效果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118570484B_ABST
    Figure CN118570484B_ABST
Patent Text Reader

Abstract

The application belongs to a forged image detection method, and particularly relates to a forged image detection method and device based on double attention and electronic equipment, which comprises the following steps: inputting a to-be-detected image into a forged feature enhancement module to extract image noise features, RGB features and image frequency features; using a double attention fusion module to fuse the image noise features and spatial domain RGB features, and then performing feature concatenation with the image frequency features to obtain fusion features; inputting the fusion features into a multi-scale feature interaction module to perform information interaction between different resolution features to obtain M feature maps with different resolutions; inputting the M feature maps with different resolutions into corresponding positioning detection branches respectively; using the output result of an r-1 hierarchical positioning detection branch as a prior to constrain the result of an r hierarchical positioning detection branch, and outputting a prediction mask and a prediction label obtained by each hierarchical positioning detection branch. The method has good forged image detection and positioning effects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method for detecting forged images, specifically a method, apparatus, and electronic device for detecting forged images based on dual attention. Background Technology

[0002] With technological advancements, generative models can generate highly realistic images, making it impossible for the naked eye to distinguish between real and fake images. Traditional detection models also struggle to make accurate judgments. This lack of identification methods means that the false information they carry can cause incalculable losses to individuals and society. In fact, image forgery detection and localization has long been a key research area in artificial intelligence security. For a forged image, researchers not only want to detect its forgery but also to pinpoint which areas within the image are fake. Unlike traditional image manipulation methods such as copying, splicing, and moving, the application of large-scale image generation models in image editing in recent years has resulted in increasingly higher image quality and blurred boundaries of tampered areas, significantly increasing the difficulty of forgery detection and localization.

[0003] Artifacts in generated images are mainly concentrated in the frequency and spatial domains. Spatial domain artifacts manifest in two ways: firstly, abnormal fluctuations in the color space of the forged image; and secondly, discrepancies in noise levels between the forged and real regions. However, current techniques for multi-feature fusion are too simplistic. Since forged regions in an image often vary significantly in location and scale, a fixed-resolution feature map may overlook forged features in some areas. Variations in the scale of forged regions negatively impact detector performance, and current techniques typically only yield a single, fixed-resolution feature map.

[0004] Therefore, it can be seen that the existing technology has poor detection and localization effects on forged images because its multi-feature fusion processing method is too simple and can only obtain a single fixed resolution feature map. Summary of the Invention

[0005] The technical problem to be solved by the present invention is that the existing technology has a poor effect on the detection and localization of forged images because the processing method of multi-feature fusion is too simple and can only obtain a single fixed resolution feature map. In order to solve the above problems, the present invention provides a forged image detection method, device and electronic device based on dual attention.

[0006] The content of this invention includes:

[0007] In a first aspect, embodiments of the present invention provide a method for detecting forged images based on dual attention, comprising:

[0008] The image to be detected is input into the forgery feature enhancement module for feature extraction, resulting in image noise features, spatial domain RGB features, and image frequency features.

[0009] The image noise features and the spatial domain RGB features are fused using a dual attention fusion module, and then concatenated with the image frequency features to obtain fused features. The dual attention fusion module includes threshold-limited channel attention and position attention.

[0010] The fused features are input into the multi-scale feature interaction module to exchange information between features of different resolutions, resulting in M ​​feature maps of different resolutions, where M is a positive integer.

[0011] The M feature maps with different resolutions are respectively input into their corresponding localization detection branches, wherein the M localization detection branches are divided into M levels from low to high according to the resolution of their corresponding feature maps;

[0012] The output of the (r-1)th level localization detection branch is used as a priori to constrain the result of the rth level. The predicted mask and predicted label obtained by the localization detection branch at each level are output, where r is an integer greater than 1 and less than or equal to M.

[0013] Optionally, the forgery feature enhancement module includes a noise module, an RGB module, and a frequency module. The step of inputting the image to be detected into the forgery feature enhancement module for feature extraction to obtain image noise features, spatial domain RGB features, and image frequency features includes:

[0014] The image to be detected is input into the noise module for feature extraction to obtain the image noise features; the image to be detected is input into the RGB module for feature extraction to obtain the spatial domain RGB features; and the image to be detected is input into the frequency module for feature extraction to obtain the image frequency domain features.

[0015] Optionally, the step of fusing the image noise features and the spatial domain RGB features using a dual attention fusion module, and then concatenating them with the image frequency features to obtain fused features, includes:

[0016] The image noise features and RGB features are input together into the dual attention fusion module to obtain spatial domain fusion features;

[0017] The spatial domain fusion features are concatenated with the image frequency features to obtain the fusion features.

[0018] Optionally, the multi-scale feature interaction module includes four extraction branches. The step of inputting the fused features into the multi-scale feature interaction module to obtain M feature maps of different resolutions includes:

[0019] The fused features are input into the four extraction branches respectively, and the outputs of the different feature branches are interacted through a fully connected layer to obtain four feature maps with different resolutions.

[0020] Optionally, the step of using the output of the (r-1)th level localization and detection branch as a priori to constrain the result of the rth level, and outputting the prediction mask and prediction label obtained by the localization and detection branch at each level, includes:

[0021] The output of the positioning detection branch at the first level from bottom to top is used as the prior input to the positioning detection branch at the second level to obtain the output result of the second level;

[0022] The output of the positioning detection branch at level 2 is used as the prior input to the positioning detection branch at level 3 to obtain the output result at level 3.

[0023] The output of the localization detection branch at level 3 is used as the prior input to the localization detection branch at level 4 to obtain the output result at level 4, which includes the prediction mask and the prediction label.

[0024] Optionally, the method further includes:

[0025] The detection model is iteratively trained using a training dataset with a loss function as the optimization objective to obtain the detection model. The detection model includes the forgery feature enhancement module, the multi-scale feature interaction module, and the progressive detection and localization module. The progressive detection and localization module includes M localization detection branches.

[0026] The loss function includes classification loss, localization loss, and edge loss. The classification loss is the sum of the classification losses on the M localization detection branches, and the classification loss on the b-th localization detection branch satisfies:

[0027]

[0028] Where N represents the number of samples, y i Represents the true category label, p(y) b |X) represents the category prediction probability on the localization detection branch b;

[0029] The positioning loss is the sum of the positioning losses on the M positioning detection branches, and the positioning loss on the b-th positioning detection branch satisfies:

[0030]

[0031] Among them, H b and W bLet represent the length and width of the image on the localization detection branch b, respectively. The label represents the actual mask position (i,j) on the localization detection branch b. This represents the predicted probability of the input image at position (i,j) on the localization detection branch b;

[0032] The edge loss satisfies:

[0033]

[0034] Where w is the hyperparameter to be learned to control the edge loss, and S x and used for S y Let M represent the Sobel convolution kernels in the x and y directions, M represent the true mask image, and m represent the predicted mask image.

[0035] Optionally, before iteratively training the detection model to be trained using the training dataset with the loss function as the optimization objective to obtain the detection model, the method further includes:

[0036] A first forged image is generated based on an input image and / or input text, and a second forged image is obtained by partially modifying a target image based on guiding conditions. The training dataset includes the first forged image and the second forged image.

[0037] Secondly, embodiments of the present invention also provide a forged image detection device based on dual attention, comprising:

[0038] The feature extraction module is used to input the image to be detected into the forgery feature enhancement module for feature extraction, and obtain image noise features, spatial domain RGB features and image frequency features;

[0039] The feature fusion module is used to fuse the image noise features and the spatial domain RGB features using the dual attention fusion module, and then concatenate them with the image frequency features to obtain fused features. The dual attention fusion module includes threshold-limited channel attention and position attention.

[0040] The feature interaction module is used to input the fused features into the multi-scale feature interaction module to exchange information between features of different resolutions, so as to obtain M feature maps of different resolutions, where M is a positive integer;

[0041] The localization detection module is used to input the M feature maps of different resolutions into their corresponding localization detection branches, wherein the M localization detection branches are divided into M levels from low to high according to the resolution of their corresponding feature maps;

[0042] The output module is used to use the output of the (r-1)th level localization detection branch as a priori to constrain the result of the rth level, and outputs the prediction mask and prediction label obtained by the localization detection branch at each level, where r is an integer greater than 1 and less than or equal to M.

[0043] Thirdly, embodiments of the present invention provide an electronic device, including: a memory, a processor, and a program stored in the memory and executable on the processor; the processor is configured to read the program in the memory to implement the steps in the dual-attention-based forged image detection method as described in the first aspect.

[0044] Fourthly, embodiments of the present invention provide a readable storage medium for storing a program, which, when executed by a processor, implements the steps in the dual-attention-based forged image detection method described in the first aspect.

[0045] In this embodiment of the invention, to overcome the adverse effects of inconsistent forgery regions, a hierarchical progressive network is employed to capture forgery artifacts at different scales of the image for detection and localization. This embodiment relies on a dual-attention mechanism to adaptively and deeply fuse multimodal image features, and then a multi-branch interactive network fully interacts the image features at different scales, improving detector performance through the dependencies between layers. Furthermore, more sensitive noise fingerprints are extracted to obtain more significant artifact features of forgery regions. Through these methods, the detection and localization of forged images are improved. Attached Figure Description

[0046] Appendix Figure 1 A flowchart illustrating a dual-attention-based forged image detection method provided in an embodiment of the present invention;

[0047] Appendix Figure 2a This is a view of the DA-HFNet network structure provided in an embodiment of the present invention;

[0048] Appendix Figure 2b For the appendix Figure 2a A schematic diagram of the structure of the forgery feature enhancement module in the middle;

[0049] Appendix Figure 2c For the appendix Figure 2a A schematic diagram of the structure of the multi-scale feature interaction module;

[0050] Appendix Figure 2d For the appendix Figure 2a A schematic diagram of the progressive detection and positioning module;

[0051] Appendix Figure 3 This is a schematic diagram of the dual attention feature fusion module structure provided in an embodiment of the present invention;

[0052] Appendix Figure 4 This is a partial fake region location result provided in an embodiment of the present invention;

[0053] Appendix Figure 5 A schematic diagram of a dual-attention-based forged image detection device provided in an embodiment of the present invention;

[0054] Appendix Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0055] In this application's embodiments, the term "and / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship. In this application's embodiments, the term "multiple" refers to two or more, and other quantifiers are similar.

[0056] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0057] Please see Figure 1 , Figure 1 This is a flowchart illustrating the dual-attention-based forged image detection method provided in an embodiment of the present invention. The method specifically includes the following steps:

[0058] Step 101: Input the image to be detected into the forgery feature enhancement module for feature extraction to obtain image noise features, spatial domain RGB features and image frequency domain features.

[0059] Step 102 is used to perform feature fusion on the image noise features and the spatial domain RGB features using a dual attention fusion module, and then perform feature concatenation with the image frequency features to obtain fused features. The dual attention fusion module includes threshold-limited channel attention and position attention.

[0060] Step 103 is used to input the fused features into the multi-scale feature interaction module to exchange information between features of different resolutions, so as to obtain M feature maps of different resolutions, where M is a positive integer.

[0061] Step 104: Input the M feature maps of different resolutions into their corresponding localization detection branches respectively. The M localization detection branches are divided into M levels from low to high according to the resolution of their corresponding feature maps.

[0062] Step 105: Use the output of the (r-1)th level localization detection branch as a priori to constrain the result of the rth level, and output the prediction mask and prediction label obtained by the localization detection branch at each level, where r is an integer greater than 1 and less than or equal to M.

[0063] In this embodiment, to overcome the adverse effects of inconsistent forgery regions, a hierarchical progressive network is employed to capture forgery artifacts at different scales of the image for detection and localization. This embodiment relies on a dual-attention mechanism to adaptively and deeply fuse multimodal image features, and then a multi-branch interactive network fully interacts the image features at different scales, improving detector performance through the dependencies between layers. Furthermore, more sensitive noise fingerprints are extracted to obtain more significant artifact features of forgery regions. Through these methods, the detection and localization of forged images are improved.

[0064] The detection model provided in this embodiment of the invention can also be called the DA-HFNet model. Please refer to [link / reference]. Figures 2a-3 The DA-HFNet model includes a fake feature enhancement module, a multi-scale feature interaction module, and a progressive detection and localization module. The progressive detection and localization module includes M localization and detection branches.

[0065] Unlike existing methods that use Steganalysis Rich Model (SRM) to obtain image noise residuals, this application uses a fake feature enhancement module to extract features, thereby obtaining image noise features and spatial domain RGB features.

[0066] Optionally, the forgery feature enhancement module includes a noise module, an RGB module, and a frequency module, and step 101 includes:

[0067] The image to be detected is input into the noise module for feature extraction to obtain the image noise features; the image to be detected is input into the RGB module for feature extraction to obtain the spatial domain RGB features; and the image to be detected is input into the frequency module for feature extraction to obtain the image frequency domain features.

[0068] Step 102 includes:

[0069] The image noise features and RGB features are input together into the dual attention fusion module to obtain spatial domain fusion features;

[0070] The spatial domain fusion features are concatenated with the image frequency features to obtain the fusion features.

[0071] In this embodiment, more significant artifact features of the forged region can be obtained by extracting more sensitive noisy fingerprints.

[0072] Because there are obvious frequency inconsistencies in images generated by different methods, this application adds an image frequency domain feature extraction branch to obtain image frequency domain features, and uses a combination of multiple features to obtain a stronger feature representation of the image forgery region.

[0073] Optionally, in some embodiments, step 103 includes:

[0074] The fused features are concatenated with the image frequency domain features and then input into the multi-scale feature interaction module to obtain M feature maps with different resolutions.

[0075] In the feature fusion stage, please refer to Figures 2a-3 This application utilizes a dual attention fusion (DAM) module to deeply fuse RGB features and noise features in the spatial domain. For example... Figure 3 As shown, the DAM module consists of two parts: threshold-limited channel attention (CA) and position attention (PA). CA focuses on the relationships between all channels and sets a threshold-limited mechanism to adjust channel weights, which helps to achieve more accurate segmentation results. PA utilizes the correlation between two features to enhance the expressive power of features by building rich contextual relationships on local features.

[0076] This application proposes a learnable weighted fusion method to fuse the output features of CA and PA, establishing a connection between the image context and features at different locations. For example... Figure 2a and Figure 2b As shown, the fused features obtained by the DAM module are further fused with the image frequency domain features and then input into the multi-scale feature interaction module.

[0077] Because the tampered regions in an image vary in size according to the generative model, and using a feature map with a fixed resolution may result in missing information, this application utilizes feature maps of different resolutions to capture richer local and global features. While conventional network structures concatenate different resolutions, this application provides a parallel feature interaction method.

[0078] It should be understood that the number of extraction branches (i.e., the value of M) included in the multi-scale feature interaction module is not limited here. Each extraction branch can extract a feature map with a fixed resolution. The number of extraction branches can be set to extract a feature map with a different resolution.

[0079] In this embodiment, a dual-attention mechanism is used to adaptively and deeply fuse multimodal image features. Then, a multi-branch interactive network fully interacts with image features at different scales, and the performance of the detector is improved by relying on the dependencies between layers. Specifically, this embodiment also utilizes multiple image features and enhances feature fusion. To learn richer feature representations, this embodiment performs dual-attention adaptive fusion processing on features of different categories. Specifically, this embodiment uses RGB features and noise features in the image spatial domain for fusion, while using frequency domain features as a supervision branch, and weighting the spatial and frequency domain features of the image to enhance the expressive power of the features.

[0080] Alternatively, as a specific embodiment, such as Figure 2c As shown, the multi-scale feature interaction module includes four extraction branches. The fused features are input into the multi-scale feature interaction module to obtain M feature maps of different resolutions, including:

[0081] The fused features are input into the four extraction branches respectively, and the outputs of the different feature branches are interacted through a fully connected layer to obtain four feature maps with different resolutions.

[0082] In this embodiment, the extraction branch is represented as θ b b∈(1...4). Each extraction branch can extract feature maps at a specific resolution, and full feature interaction is achieved between different extraction branches through fully connected connections. The output of each extraction branch is obtained by fusing the outputs of all extraction branches. For example, extraction branch θ3, which is downsampled by 4 times, is obtained by adding the outputs of extraction branch θ1 (downsampled by 4 times), extraction branch θ2 (downsampled by 2 times), extraction branch θ4 (upsampled by 2 times), and θ3 itself.

[0083] In this embodiment, a hierarchical network is used to achieve coarse-to-fine image-level forgery detection and pixel-level forgery localization. Specifically, M feature maps of different resolutions are obtained in the aforementioned steps. Then, starting from the lowest resolution feature map, category prediction and region prediction for that level are performed in a detection module and a localization module, respectively. The detection module uses a simple linear classifier, and the localization module employs a self-attention mechanism to map the feature map of that level to a mask. b .

[0084] After obtaining the result at the current level, the classification and localization prediction results from the previous level are introduced as priors to constrain the results at the next level. This embodiment employs a progressive network to improve our detection and localization results layer by layer. For detection tasks, considering that lower-resolution features may lack information and fail to achieve fine-grained classification, a coarser classification category is set in this embodiment. Similarly, for localization tasks, this embodiment uses the coarsely estimated fake regions obtained from the low-resolution branch to guide the next layer in improving the localization effect.

[0085] Since forged regions in an image often vary significantly in location and scale, a fixed-resolution feature map may overlook forged features in some regions. Variations in the scale of forged regions negatively impact detector performance, a point largely ignored in existing work. Therefore, this embodiment utilizes a hierarchical network to fully exchange information between features of different resolutions, and then performs detection and localization at different scales from image features of varying resolutions. Furthermore, this embodiment employs an edge loss mechanism to correct the localization of forged regions, further improving model performance.

[0086] Optionally, in some embodiments, M is set to 4, and the output of the (r-1)th level localization detection branch is used as a priori to constrain the result of the rth level, outputting the prediction mask and prediction label obtained by the localization detection branch at each level, including:

[0087] The output of the positioning detection branch at the first level from bottom to top is used as the prior input to the positioning detection branch at the second level to obtain the output result of the second level;

[0088] The output of the positioning detection branch at level 2 is used as the prior input to the positioning detection branch at level 3 to obtain the output result at level 3.

[0089] The output of the localization detection branch at level 3 is used as the prior input to the localization detection branch at level 4 to obtain the output result at level 4, which includes the prediction mask and the prediction label.

[0090] In practice, the output of the highest-level localization and detection branch is the final predicted mask and predicted label. Specifically, the above process can be described as follows:

[0091] Given image The features extracted by the above four extraction branches are respectively represented as follows: Then we have the location spoofing area as:

[0092]

[0093] mask b+1 =softmax(F b+1 +Γ(mask b ·F b+1 / 2)).

[0094] in, Let Γ(·) represent the localization head of extraction branch b, and let Γ(·) represent a linear interpolator that expresses the class prediction results and prediction probabilities of different extraction branches as θ. b and p(y b |X), the probability prediction satisfies:

[0095]

[0096] p(y b+1 |X)=softmax(θ b+1 (X)⊙(1+p(y b |X))).

[0097] in, This indicates that the category detection head on branch b is extracted.

[0098] To train the DA-HFNet model, this embodiment sets different loss functions as optimization objectives for the detection and localization tasks in each localization and detection branch. Optionally, the method further includes:

[0099] The detection model to be trained is iteratively trained using a training dataset with a loss function as the optimization objective to obtain the detection model. The detection model includes the forgery feature enhancement module, the multi-scale feature interaction module, and the progressive detection and localization module.

[0100] The loss function includes classification loss, localization loss, and edge loss.

[0101] Specifically, for the task of detecting forgery categories, the optimization objective of each branch is expressed as follows: the classification loss on the b-th localization detection branch satisfies:

[0102]

[0103] Where N represents the number of samples, y i Represents the true category label, p(y) b |X) represents the predicted category probability on the localization detection branch b; the classification loss is the sum of the classification losses on the M localization detection branches, in such a case... Figure 2a and Figure 2d In the example shown, the classification loss for the four branches is expressed as:

[0104]

[0105] For the task of locating fake regions, the optimization objective of each branch is expressed using the binary cross-entropy loss function as follows: the localization loss on the b-th localization detection branch satisfies:

[0106]

[0107] Among them, H b and W b These represent the length and width of the image on branch b, respectively. The label representing the actual mask position (i,j) on branch b. This represents the predicted probability of the input image at position (i,j) on branch b; the localization loss is the sum of the localization losses on the M localization detection branches, such as... Figure 2a and Figure 2d In the embodiment shown, the localization loss of the four branches is expressed as:

[0108]

[0109] The localization of forged regions falls under the category of image segmentation. Considering that segmentation tasks often suffer from misclassification of forged pixels and real pixels at boundary positions, and that AI-edited images do not show obvious boundary marks in the edited region, this embodiment sets up a learnable edge loss. Since this embodiment focuses more on the final result of forged region localization, the edge loss of all branches is not calculated, but only the edge loss of the highest resolution branch is calculated.

[0110] The edge loss satisfies:

[0111]

[0112] Where w is the hyperparameter to be learned to control the edge loss, and S x and used for S y Let M represent the Sobel convolution kernels in the x and y directions, M represent the true mask image, and m represent the predicted mask image.

[0113] In this embodiment, the sum of classification loss, localization loss, and edge loss is used as the final loss function. Therefore, the loss function Loss during iterative training satisfies:

[0114]

[0115] Optionally, before iteratively training the detection model to be trained using the training dataset with the loss function as the optimization objective to obtain the detection model, the method further includes:

[0116] A first forged image is generated based on an input image and / or input text, and a second forged image is obtained by partially modifying a target image based on guiding conditions. The training dataset includes the first forged image and the second forged image.

[0117] The first forged image is generated using the full image generation mode. As an optional implementation, models such as CycleGAN, LinkGAN, and DDPM are used to generate a complete forged image from an image as input. As another optional implementation, models such as StyleGAN, GigaGAN, GLIDE, and StableDiffusion are used to generate a complete forged image from text as input.

[0118] The second type of forged image is a forged image generated in a partial editing mode. The partial editing mode only modifies a certain part of the image. It can use the image as a guiding condition to guide the model to partially modify the target image, such as EditGAN, DragGAN, HD-painter, Paint-byexample, etc., as well as use text as a guide to modify a part of the input image, such as StyleCLIP, Imagic, Ranni, etc.

[0119] These generative models produce highly realistic forged images, making them difficult for detection models to accurately identify. In particular, image manipulation based on generative models differs from traditional stitching methods because real and forged pixels coexist within the same image, and the forged boundaries are more concealed. Therefore, such images are more difficult for models to successfully detect and locate. The method described above yields a first and a second forged image. Using these images to train the detection model can improve training effectiveness, resulting in a better-performing model.

[0120] In this embodiment, to address the diversity of generative models and generative modes, a DA-HFNet fake image training dataset covering the most mainstream generative models was constructed. This dataset includes text-guided image editing and image-guided image editing methods, which solves the problem of the lack of image editing methods in current AI-generated image datasets. This improves the completeness of the training dataset and thus enhances the model training effect.

[0121] To verify the effectiveness of this method in locating and detecting forged images, experimental tests were conducted in this embodiment, and the specific instructions are as follows.

[0122] First, to address the lack of corresponding datasets for some advanced image editing methods, this embodiment constructs an image generation and editing dataset based on GANs and DMs, which includes both image and text modes in terms of guidance. The DA-HFNet dataset is constructed in this embodiment primarily for two reasons: 1) Forged images based on generative models are not only becoming increasingly realistic in terms of generation quality, but the generation methods are also becoming more diverse. In particular, some image editing methods no longer have clear editing boundaries, but existing datasets do not yet cover this area. 2) The detection and localization of forged images requires research on various types of forgery methods, and datasets containing various methods are essential for conducting such research.

[0123] The relevant composition structure of the dataset in this embodiment is shown in Table 1. The image-guided generated images are from BigGAN, DDPM, and PaintbyExample. BigGAN is a representative GAN-based whole-image generation method; DDPM is a representative diffusion-based whole-image generation method; PaintbyExample can use a diffusion model to draw a specified region of a given image onto another image. The text-guided generated images are from FuseDream, GLIDE, InpaintAnything, and StyleCLIP. FuseDream can generate an image that matches a given text using GAN; GLIDE can generate an image that matches a given text using a diffusion model; InpaintAnything uses a diffusion model to change a specified region in a real image into an image that matches the text description; StyleCLIP uses GAN to edit an image into an image that matches the text description. For real images, 10k and 10k images were randomly selected from COCO2017 and ImageNet, respectively.

[0124] In this embodiment, accuracy (ACC) and F1 score (F1) are selected as the evaluation metrics for our experiment. The basic experimental setup is as follows: DA-HFNet is implemented in PyTorch and trained on two NVIDIA 3090 GPUs. The image input size is set to 256×256, the initial learning rate is set to 0.0002 and periodically decays to 1e-08, and the number of training epochs is set to 50. We apply common data augmentation methods to the training process, including rotation and inversion.

[0125] Table 1. Composition structure of the DA-HFNet dataset

[0126]

[0127] This embodiment tests the model on the DA-HFNet dataset constructed in this paper. First, the method in this embodiment is compared with the baseline method in image-level classification and pixel-level localization. Table 2 reports the results of different methods in image-level attribute category classification. Specifically, Table 2 shows that the pre-trained detectors generally perform poorly on our dataset, which is directly related to the feature representation ability of the parameters learned by the detectors. Trufor uses features including noise features and RGB features; PSCC-Net uses the RGB features of the image and employs a hierarchical structure. HIFI-IFDL uses RGB features and frequency domain features, also employing a hierarchical structure. PSCC-Net and HIFI-IFDL are representative methods in hierarchical networks. On our dataset, our method, trained with the detectors trained on it, shows the best ACC score on both GAN-based and Diffusionmodel-based forged images, and its F1 score is also highly competitive. Table 3 then reports the performance of different detectors in pixel-level forged region localization. Our baseline model uses the same detector used in image-level forged category classification. Table 3 shows that the pre-trained detectors do not perform well on our dataset. This is partly because the segmented objects performed by the pre-trained PSCC-Net and HIFI-IFDL are mostly from manually edited forged images, which have significant fingerprint differences compared to AIGC-edited forged images. We then trained the PSCC and IFDL detectors using our DA-HFNet dataset, which showed significant improvements in both ACC and F1 scores. Furthermore, we observed that for forged region localization, our method achieved the best results in both ACC and F1 scores, with an average ACC score improvement of 3.93% and an average F1 score improvement of 2.72% compared to the second-best method. Figure 4 The image shows a comparison of our fake region location effect and comparison methods.

[0128] Table 2. Experimental Results of Image-Level Forgery Detection

[0129]

[0130] Table 3. Experimental Results of Pixel-Level Spoofed Region Location

[0131]

[0132] This embodiment further conducted ablation experiments, specifically comparing the ablation of the multi-feature dual-attention interaction module in the method. Table 4 reports the results of the ablation experiments. First, this embodiment removed the dual-attention feature fusion module from the method, fusing features only using feature cascading. The method's ACC and F1 scores decreased by 2.84% and 6.2% respectively in forgery category detection, and by 3.08% in forgery region localization, with a slight decrease in F1 score. This embodiment attributes this to insufficient feature fusion after removing the dual-attention fusion module. Subsequently, this embodiment ablated each feature used in the method. Rows 2-4 of Table 4 show that removing image features from any branch causes varying degrees of performance degradation. The most severe degradation occurs when the frequency branch is removed, resulting in a 7.68% decrease in ACC for forgery category classification and a 5.75% decrease in ACC for forgery region localization. Therefore, this embodiment concludes that the features of each branch positively impact the method's performance. Finally, this embodiment also replaces the noise features of this embodiment with the method of obtaining noise residuals by SRM filtering. The results are shown in the fifth row of Table 4. Compared with the method of this embodiment, the detector trained by SRM filtering has decreased in both ACC and F1 scores.

[0133] Table 4 Ablation Experiment Results

[0134]

[0135] In subsequent experiments, this embodiment investigated the impact of edge loss on the experiments, and the corresponding experimental results are reported in Table 5. As can be seen from Table 5, when we removed the edge loss, the model's ACC score and F1 score decreased by 6.47% and 4.88% respectively in the fake category classification task, and by 4.35% and 5.79% respectively in the fake region localization task. This demonstrates the positive effect of the edge loss we set in the experiments.

[0136] Table 5. Impact of Edge Loss on Model Detection Performance: Ablation Experiment Results

[0137]

[0138] This embodiment proposes a dual-attention-based method for detecting forged images. This method employs a progressive network with multi-feature fusion for detecting and locating high-quality AIGC forged images. It utilizes noise features, which are more sensitive to noise, and combines them with the image's RGB and frequency features to obtain richer forgery traces. A dual-attention fusion mechanism adaptively fuses multimodal image features, and a multi-branch feature interaction network enables information exchange between features of different resolutions. Furthermore, this embodiment constructs a high-quality forged image dataset containing GANs and a Diffusion model, generated using text-guided or image-guided methods. Extensive experiments on this dataset demonstrate the effectiveness of the method provided in detecting and locating high-quality forged images. Using ACC and F1 scores as evaluation metrics, compared to existing forged image detection and localization methods, the method presented in this embodiment achieves a 3.93% improvement in ACC score and a 98.37% F1 score on the forged attribute recognition task, demonstrating strong competitiveness. On the forged region localization task, the method's ACC and F1 scores are improved by 2.49% and 2.72%, respectively, compared to state-of-the-art methods.

[0139] Please see Figure 5 This invention also provides a dual-attention-based forged image detection device 500, comprising:

[0140] The feature extraction module 501 is used to input the image to be detected into the forgery feature enhancement module for feature extraction, and obtain image noise features, spatial domain RGB features and image frequency features;

[0141] The feature fusion module 502 is used to fuse the image noise features and the spatial domain RGB features using a dual attention fusion module, and then concatenate them with the image frequency features to obtain fused features. The dual attention fusion module includes threshold-limited channel attention and position attention.

[0142] The feature interaction module 503 is used to input the fused features into the multi-scale feature interaction module to exchange information between features of different resolutions, so as to obtain M feature maps of different resolutions, where M is a positive integer;

[0143] The localization detection module 504 is used to input the M feature maps of different resolutions into their corresponding localization detection branches, wherein the M localization detection branches are divided into M levels from low to high according to the resolution of their corresponding feature maps;

[0144] The output module 505 is used to use the output result of the localization detection branch at level r-1 as a priori to constrain the result at level r, and output the prediction mask and prediction label obtained by the localization detection branch at each level, where r is an integer greater than 1 and less than or equal to M.

[0145] Optionally, the forgery feature enhancement module includes a noise module, an RGB module, and a frequency module, and the feature extraction module 501 is specifically used for:

[0146] The image to be detected is input into the noise module for feature extraction to obtain the image noise features; the image to be detected is input into the RGB module for feature extraction to obtain the spatial domain RGB features; and the image to be detected is input into the frequency module for feature extraction to obtain the image frequency domain features.

[0147] Optionally, the feature fusion module 502 is specifically used for:

[0148] The image noise features and RGB features are input together into the dual attention fusion module to obtain spatial domain fusion features;

[0149] The spatial domain fusion features are concatenated with the image frequency features to obtain the fusion features.

[0150] Optionally, the multi-scale feature interaction module includes four extraction branches, and the feature interaction module 503 is specifically used for:

[0151] The fused features are input into the four extraction branches respectively, and the outputs of the different feature branches are interacted through a fully connected layer to obtain four feature maps with different resolutions.

[0152] Optionally, the output module 505 is specifically used for:

[0153] The output of the positioning detection branch at the first level from bottom to top is used as the prior input to the positioning detection branch at the second level to obtain the output result of the second level;

[0154] The output of the positioning detection branch at level 2 is used as the prior input to the positioning detection branch at level 3 to obtain the output result at level 3.

[0155] The output of the localization detection branch at level 3 is used as the prior input to the localization detection branch at level 4 to obtain the output result at level 4, which includes the prediction mask and the prediction label.

[0156] Optionally, the dual-attention-based forged image detection device 500 further includes:

[0157] The training module is used to iteratively train the detection model to be trained using the training dataset with the loss function as the optimization objective, so as to obtain the detection model. The detection model includes the forgery feature enhancement module, the multi-scale feature interaction module, and the progressive detection and localization module. The progressive detection and localization module includes M localization detection branches.

[0158] The loss function includes classification loss, localization loss, and edge loss. The classification loss is the sum of the classification losses on the M localization detection branches, and the classification loss on the b-th localization detection branch satisfies:

[0159]

[0160] Where N represents the number of samples, y i Represents the true category label, p(y) b |X) represents the category prediction probability on the localization detection branch b;

[0161] The positioning loss is the sum of the positioning losses on the M positioning detection branches, and the positioning loss on the b-th positioning detection branch satisfies:

[0162]

[0163] Among them, H b and W b Let represent the length and width of the image on the localization detection branch b, respectively. The label represents the actual mask position (i,j) on the localization detection branch b. This represents the predicted probability of the input image at position (i,j) on the localization detection branch b;

[0164] The edge loss satisfies:

[0165]

[0166] Where w is the hyperparameter to be learned to control the edge loss, and S x and used for S y Let M represent the Sobel convolution kernels in the x and y directions, M represent the true mask image, and m represent the predicted mask image.

[0167] Optionally, the dual-attention-based forged image detection device 500 further includes:

[0168] The generation module is used to generate a first forged image based on an input image and / or input text, and to guide the model to partially tamper with the target image based on guiding conditions to obtain a second forged image. The training dataset includes the first forged image and the second forged image.

[0169] The dual-attention-based forged image detection device 500 provided in this application embodiment can execute the above method embodiment, and its implementation principle and technical effect are similar, so it will not be described again here.

[0170] It should be noted that the division of units in the embodiments of this application is illustrative and only represents one logical functional division. In actual implementation, other division methods may be used. Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated units described above can be implemented in hardware or as software functional units.

[0171] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a processor-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0172] like Figure 6 As shown, this application provides an electronic device 600, including: a memory 602, a processor 601, and a program stored in the memory 602 and executable on the processor 601; the processor 601 is used to read the program in the memory 602 to implement the steps in the dual attention-based fake image detection method as described above.

[0173] This application also provides a readable storage medium storing a program that, when executed by a processor, implements the various processes of the above-described dual-attention-based forged image detection method embodiments and achieves the same technical effect. To avoid repetition, it will not be described again here. The readable storage medium can be any available medium or data storage device that the processor can access, including but not limited to magnetic storage (e.g., floppy disks, hard disks, magnetic tapes, magneto-optical disks (MO), etc.), optical storage (e.g., compact disks (CDs), digital video discs (DVDs), Blu-ray discs (BD), high-definition universal discs (HVD), etc.), and semiconductor storage (e.g., read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), non-volatile memory (NAND FLASH), solid-state disks (SSDs), etc.).

[0174] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0175] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0176] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A method for detecting forged images based on dual attention, characterized in that, include: The image to be detected is input into the forgery feature enhancement module for feature extraction, resulting in image noise features, spatial domain RGB features, and image frequency features. The image noise features and the spatial domain RGB features are fused using a dual attention fusion module, and then concatenated with the image frequency features to obtain fused features. The dual attention fusion module includes threshold-limited channel attention and position attention. The fused features are input into the multi-scale feature interaction module to exchange information between features of different resolutions, resulting in M ​​feature maps of different resolutions, where M is a positive integer. The M feature maps with different resolutions are respectively input into their corresponding localization detection branches, wherein the M localization detection branches are divided into M levels from low to high according to the resolution of their corresponding feature maps; The output of the (r-1)th level localization detection branch is used as a priori to constrain the result of the rth level. The predicted mask and predicted label obtained by the localization detection branch at each level are output, where r is an integer greater than 1 and less than or equal to M. The step of fusing the image noise features and the spatial domain RGB features using a dual attention fusion module, and then concatenating them with the image frequency features to obtain fused features, includes: The image noise features and spatial domain RGB features are input together into the dual attention fusion module to obtain spatial domain fusion features; The spatial domain fusion features are concatenated with the image frequency features to obtain the fusion features.

2. The method as described in claim 1, characterized in that, The forgery feature enhancement module includes a noise module, an RGB module, and a frequency module. The step of inputting the image to be detected into the forgery feature enhancement module for feature extraction, obtaining image noise features, spatial domain RGB features, and image frequency features, includes: The image to be detected is input into the noise module for feature extraction to obtain the image noise features; the image to be detected is input into the RGB module for feature extraction to obtain the spatial domain RGB features; and the image to be detected is input into the frequency module for feature extraction to obtain the image frequency domain features.

3. The method as described in claim 1 or 2, characterized in that, The multi-scale feature interaction module includes four extraction branches. The fused features are input into the multi-scale feature interaction module to obtain M feature maps of different resolutions, including: The fused features are input into the four extraction branches respectively, and the outputs of the different extraction branches are interacted through a fully connected layer to obtain four feature maps with different resolutions.

4. The method as described in claim 1, characterized in that, The step of using the output of the (r-1)th level localization and detection branch as a priori to constrain the result of the rth level, and outputting the prediction mask and prediction label obtained by the localization and detection branch at each level, includes: The output of the positioning detection branch at the first level from bottom to top is used as the prior input to the positioning detection branch at the second level to obtain the output result of the second level; The output of the positioning detection branch at level 2 is used as the prior input to the positioning detection branch at level 3 to obtain the output result at level 3. The output of the localization detection branch at level 3 is used as the prior input to the localization detection branch at level 4 to obtain the output result at level 4, which includes the prediction mask and the prediction label.

5. The method as described in claim 1, characterized in that, The method further includes: The detection model is iteratively trained using a training dataset with a loss function as the optimization objective to obtain the detection model. The detection model includes the forgery feature enhancement module, the multi-scale feature interaction module, and the progressive detection and localization module. The progressive detection and localization module includes M localization detection branches. The loss function includes classification loss, localization loss, and edge loss. The classification loss is the sum of the classification losses on the M localization detection branches, and the classification loss on localization detection branch b satisfies: ; Where N represents the number of samples, Indicates the true category label, This represents the predicted category probability on the localization detection branch b; The positioning loss is the sum of the positioning losses on the M positioning detection branches, and the positioning loss on positioning detection branch b satisfies: ; in, and Let represent the length and width of the image on the localization detection branch b, respectively. The label represents the actual mask position (i,j) on the localization detection branch b. This represents the predicted probability of the input image at position (i,j) on the localization detection branch b; The edge loss satisfies: ; Where w is the hyperparameter to be learned to control the edge loss. and Let M and m represent the Sobel convolution kernels in the x and y directions, respectively. M represents the true mask image, and m represents the predicted mask image.

6. The method as described in claim 5, characterized in that, Before iteratively training the detection model to be trained using the training dataset with the loss function as the optimization objective to obtain the detection model, the method further includes: A first forged image is generated based on an input image and / or input text, and a second forged image is obtained by partially modifying a target image based on guiding conditions. The training dataset includes the first forged image and the second forged image.

7. A forged image detection device based on dual attention, characterized in that, include: The feature extraction module is used to input the image to be detected into the forgery feature enhancement module for feature extraction, and obtain image noise features, spatial domain RGB features and image frequency features; The feature fusion module is used to fuse the image noise features and the spatial domain RGB features using the dual attention fusion module, and then concatenate them with the image frequency features to obtain fused features. The dual attention fusion module includes threshold-limited channel attention and position attention. The feature interaction module is used to input the fused features into the multi-scale feature interaction module to exchange information between features of different resolutions, so as to obtain M feature maps of different resolutions, where M is a positive integer; The localization detection module is used to input the M feature maps of different resolutions into their corresponding localization detection branches, wherein the M localization detection branches are divided into M levels from low to high according to the resolution of their corresponding feature maps; The output module is used to use the output result of the localization detection branch at level (r-1) as a priori to constrain the result at level r, and outputs the prediction mask and prediction label obtained by the localization detection branch at each level, where r is an integer greater than 1 and less than or equal to M. Specifically, the feature fusion module is used for: The image noise features and spatial domain RGB features are input together into the dual attention fusion module to obtain spatial domain fusion features; The spatial domain fusion features are concatenated with the image frequency features to obtain the fusion features.

8. An electronic device, comprising: A memory, a processor, and a program stored in the memory and executable on the processor; characterized in that the processor is configured to read the program from the memory to implement the steps of the dual-attention-based forgery image detection method as described in any one of claims 1 to 6.

9. A readable storage medium for storing a program, characterized in that, When the program is executed by the processor, it implements the steps in the dual-attention-based forgery image detection method as described in any one of claims 1 to 6.