A method for locating image tampering areas using a bidirectional interactive network

Through the design of a two-way interactive network model and the use of multi-scale feature information and information feedback mechanism, the problem of precise positioning of image tampering detection in existing technologies is solved, achieving higher positioning accuracy and robustness.

CN120431410BActive Publication Date: 2025-09-12NANJING UNIV OF INFORMATION SCI & TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510929447.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2025-09-12
Estimated Expiration
2045-07-07

AI Technical Summary

Technical Problem

Existing technologies have difficulty in effectively utilizing multi-scale feature information and information feedback mechanisms in image tampering detection, resulting in positioning results that are not precise or accurate enough, especially when facing complex tampering techniques and post-processing operations, and lack robustness and generalization capabilities.

Method used

A bidirectional interactive network model is adopted to perform multi-scale feature extraction through the backbone network. Combined with the downsampling feature compensation module and the overall attention enhancement module, the feature map reorganization, splicing, weighted fusion and attention guidance are realized. The focal loss function is used to optimize the network model to improve positioning accuracy and robustness.

Benefits of technology

The positioning accuracy of image tampering areas and the robustness of the model are significantly improved, and it can maintain efficient positioning effects in the face of common image interference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120431410B_ABST
    Figure CN120431410B_ABST
Patent Text Reader

Abstract

This invention belongs to image localization technology, specifically a method for locating image tampering areas using a bidirectional interactive network. The method uses a backbone network to extract multi-scale features, which are then input into a downsampling feature compensation module and an overall attention enhancement module. The downsampling feature compensation module processes the compensated feature map, which is then fused with the current feature map and the underlying prediction map to generate an attention feature map. The fused features are then decoded layer by layer using a multi-level prediction map and a hierarchical fusion module to generate the localization results for the forged areas. Through the bidirectional interactive network model design, the present invention can effectively utilize multi-scale feature information and an information feedback mechanism to accurately locate the tampered area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to information security technology, and in particular relates to a method for locating an image tampering area by utilizing a bidirectional interactive network. Background Art

[0002] With the rapid development of image editing tools and generative models, image manipulation techniques (such as splicing, object removal, and copy-and-paste) have become increasingly sophisticated, making it easy for ordinary users to forge image content. Such manipulated images can cause misleading information and security risks in fields such as news, social media, and the legal system. Therefore, accurately detecting and locating manipulated areas has become a research priority in digital image forensics.

[0003] Early image tampering detection methods mainly relied on manually designed statistical features, such as color filter array (CFA) patterns, discrete cosine transform (DCT) coefficient analysis, noise inconsistency analysis, etc. These methods performed well under controlled conditions, but had poor robustness and generalization capabilities when faced with complex tampering techniques or post-processing operations (such as JPEG compression, blurring, and resizing).

[0004] In recent years, with the development of deep learning technology, more and more research has adopted convolutional neural networks (CNNs) and Transformer architectures to automatically detect and locate image forgeries. Typical methods generally use an encoder-decoder architecture to learn the pixel-level distribution of tampered areas through end-to-end training. Kwon et al. (Kwon MJ, Nam SH, Yu IJ, et al. Learning jpeg compression artifacts for image manipulation detection and localization[J]. International Journal of Computer Vision, 2022, 130(8): 1875-1895.) and Li et al. ( Li D, Zhu J, Wang M, et al. Edge-awareregional message passing controller for image forgery localization[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and PatternRecognition. 2023: 8222-8232.) used compression artifacts and edge features as clues for locating tampered areas, respectively. This design makes it impossible to fully utilize the rich intermediate features extracted by the encoder at different stages, and lacks feedback or refinement mechanisms that can further optimize tampering localization. This defect may cause the loss of key information in the feature extraction process, ultimately affecting the accuracy and robustness of the positioning results. For example, Chinese patent application publication number CN113962941A discloses a method, device, terminal, and storage medium for locating tampered areas in tampered images. However, this solution lacks multi-scale information, making it difficult to simultaneously capture the relationship between local details and global context, resulting in limited perception of tampered areas at different scales. Chinese patent application publication number CN117853397A discloses a method and system for detecting and locating image tampering based on multi-level feature learning. However, this solution lacks an effective information feedback mechanism, making it difficult to use the current prediction results for self-correction or gradual optimization, resulting in positioning results that are not precise or accurate enough. Summary of the Invention

[0005] Purpose of the invention: The purpose of the present invention is to address the deficiencies in the prior art and provide a method for locating image tampering areas using a bidirectional interactive network. Through the design of a bidirectional interactive network model, multi-scale feature information and information feedback mechanism can be effectively utilized to accurately locate the tampered area.

[0006] Technical solution: The present invention provides a method for locating image tampering areas using a bidirectional interactive network. For an input original image, the following steps are performed:

[0007] Step 1: Use a backbone network (such as ResNet) to perform multi-scale feature extraction to obtain feature maps of different levels of semantic and detail information;

[0008] Step 2: First, the feature maps of each layer are input into the corresponding downsampling feature compensation module and overall attention enhancement module respectively;

[0009] The downsampling feature compensation module reorganizes, splices, group normalizes, and dynamically weights the feature map of the current layer, and then combines it with the feature map of the current layer to obtain a compensation feature map; the overall attention enhancement module performs attention guidance enhancement on the feature map of the current layer and the underlying prediction map, and then fuses them with the feature map of the current layer to obtain an attention feature map;

[0010] The bottom layer prediction map p2 used by the overall attention enhancement module of the first layer is obtained by inputting the compensated feature map obtained by the downsampling feature compensation module of the first layer into the second layer through training;

[0011] Then, starting from the second layer, the feature map of the current layer is spliced ​​with the compensated feature map of the previous layer to generate enhanced representation features, which are then fed into the corresponding overall attention enhancement module.

[0012] Step 3: Decode the features layer by layer through the multi-level prediction map and hierarchical fusion module to generate the positioning results of the forged area.

[0013] In order to alleviate the loss of effective information during feature downsampling, the specific implementation process of the downsampling feature compensation module is as follows:

[0014] First, the input feature map is reorganized using a phase-interval sampling strategy to generate four complementary sub-feature maps. The calculation formulas for the four sub-feature maps are as follows:

[0015] ; ;

[0016] ; ;

[0017] in, Representatives in The pixel value at position, Represents four sub-feature maps;

[0018] Next, the sub-feature map Splice by channel and sub-feature map Splicing is done in the channel direction, and the two spliced ​​feature maps are reduced in dimension through 1×1 convolution to compress redundant information and improve computational efficiency.

[0019] Subsequently, a group normalization operation is introduced and combined with a learnable convolution kernel to enhance the two feature maps after dimensionality reduction, obtaining two sets of compensatory features, thereby improving the modeling ability of local spatial structures.

[0020] Then, using the learnable dynamic factor (The dynamic factor can be adjusted automatically according to the network training) The two sets of compensation features are weightedly fused to achieve effective integration of global and local information;

[0021] Finally, the weighted fusion result is combined with the original input feature map through residual connection to generate the final output feature, thereby retaining key semantic information to the greatest extent and improving the expressiveness of the model.

[0022] Furthermore, the specific processing method of the overall attention enhancement module of each layer is as follows:

[0023] First, the underlying prediction map is Gaussian blurred and the result is normalized to generate a guiding attention weight map. Then, the obtained attention weight map is element-wise multiplied with the current layer feature map to achieve semantically guided feature enhancement and obtain enhanced features.

[0024] Then, the enhanced features are added to the original features, and the fusion results are normalized through layer normalization to obtain the final output features; the calculation formula is as follows:

[0025] ;

[0026] In the above formula, P represents the bottom layer prediction map, S represents the feature map of the current layer, LN represents layer normalization, Cg represents a convolution operation using a Gaussian kernel with a bias of zero, and f(·) is a normalization function used to map the Gaussian blurred image to the range [0, 1]. The coverage area of ​​the coarse prediction map P is expanded to reduce information loss during the filtering process.

[0027] Among them, the current layer feature map of the fourth layer is spliced ​​with the compensated feature map of the downsampling feature compensation module of the third layer, and then processed to obtain the prediction map P4, which is used as the deep prediction map as input to the overall attention enhancement module of the third layer; the features processed by the overall attention enhancement module of the first layer include the current layer feature map and the deep prediction map P2.

[0028] To reduce the mutual interference between channels, depthwise separable convolution is introduced in the hierarchical fusion module of step 3. Replacing traditional convolution with depthwise separable convolution not only effectively reduces the computational cost and parameter amount of the model, but also retains the independent structural information of each input channel, which is helpful for feature decoupling and channel importance analysis.

[0029] During the hierarchical fusion process, input features first undergo point-by-point and channel-by-channel convolutions to integrate cross-channel information and extract local features. Batch normalization is then performed to improve training stability and accelerate convergence. To enhance nonlinear expression capabilities, the GELU activation function is used. This probabilistically smooths and controls the degree of information retention, particularly by partially activating inputs close to zero, making it more advantageous in capturing fine-grained feature variations. Furthermore, to generate the final prediction, a classifier is used to perform one-dimensional convolution on the channel dimension for dimensionality reduction. Finally, a sigmoid activation function is used to output a normalized probability map, representing the spatial distribution of forged regions.

[0030] Because locating tampered image regions requires a binary classification problem, the network's final layer uses a sigmoid layer to map the network output to the range [0, 1], representing the probability of tampering. To further optimize the network model, a focal loss function is used. This function reduces the loss contribution of easily classified samples, allowing the model to focus more on difficult or minority samples.

[0031] The focal loss function is expressed as:

[0032] ;

[0033] Where P is the prediction map, G is the label, α is the category balance factor that controls the importance of positive and negative samples, and β is the focus parameter used to reduce the loss contribution of easy-to-classify samples.

[0034] Beneficial Effects: The downsampling feature compensation module of the present invention can effectively alleviate the problem of semantic information loss caused by image scaling and enhance the expressiveness of deep features. The overall attention enhancement module uses coarse-grained prediction maps as a guide to encourage shallow features to focus on the tampered area, achieving efficient fusion of semantic and detail information, thereby significantly improving positioning accuracy. Supported by the bidirectional feature interaction mechanism, the network model of the present invention exhibits stronger robustness and generalization ability when dealing with common image interference. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 Schematic diagram of the overall network model structure and positioning process in the present invention;

[0036] Figure 2 Schematic diagram of the downsampling feature compensation module in the present invention;

[0037] Figure 3 This is a comparison diagram of the image tampering area in the embodiment. DETAILED DESCRIPTION

[0038] The technical solution of the present invention is described in detail below, but the protection scope of the present invention is not limited to the embodiments.

[0039] like Figure 1 As shown, the present invention provides a method for locating image tampering areas using a bidirectional interactive network. For an input original image, the following steps are performed:

[0040] Step 1: Use the backbone network to perform multi-scale feature extraction to obtain feature maps of different levels of semantics and detail information;

[0041] Step 2: First, the feature maps of each layer are input into the corresponding downsampling feature compensation module and overall attention enhancement module respectively;

[0042] The downsampling feature compensation module reorganizes, splices, group normalizes, and dynamically weights the feature map of the current layer, and then combines it with the feature map of the current layer to obtain a compensation feature map; the overall attention enhancement module performs attention guidance enhancement on the feature map of the current layer and the underlying prediction map, and then fuses them with the feature map of the current layer to obtain an attention feature map;

[0043] The bottom layer prediction map p2 used by the overall attention enhancement module of the first layer is obtained by inputting the compensated feature map obtained by the downsampling feature compensation module of the first layer into the second layer through training;

[0044] Then, starting from the second layer, the feature map of the current layer is spliced ​​with the compensated feature map of the previous layer to generate enhanced representation features, which are used as the enhanced representation features;

[0045] Step 3: Through the multi-layer prediction map and hierarchical fusion module, the features (the features are the concatenation and fusion of the features after conv1x1 in the next layer and the features of the overall attention enhancement module in the previous layer in the channel dimension) are decoded layer by layer to generate the positioning results of the forged area.

[0046] like Figure 2 As shown, the specific execution process of the downsampling feature compensation module of this embodiment is:

[0047] First, the input feature map is reorganized using a phase-interval sampling strategy to generate four complementary sub-feature maps. The calculation formulas for the four sub-feature maps are as follows:

[0048] ; ;

[0049] ; ;

[0050] in, Representatives in The pixel value at position, Represents four sub-feature maps;

[0051] Next, the sub-feature map Splice by channel and sub-feature map Splice in the channel direction and perform dimensionality reduction on the two spliced ​​feature maps through 1×1 convolution;

[0052] Subsequently, the group normalization operation is introduced and combined with the learnable convolution kernel to enhance the two feature maps after dimensionality reduction to obtain two sets of compensation features;

[0053] Then, using the learnable dynamic factor Perform weighted fusion on the two sets of compensation features;

[0054] Finally, the weighted fusion result is combined with the original input feature map through residual connection to generate the corresponding compensated feature map.

[0055] The specific processing methods of the overall attention enhancement modules of the above layers are as follows:

[0056] First, the underlying prediction map is Gaussian blurred and the result is normalized to generate the attention weight map;

[0057] Then perform element-wise multiplication of the obtained attention weight map with the current layer feature map to obtain the enhanced features;

[0058] Then, the enhanced features are added to the original features, and the fusion results are normalized through layer normalization to obtain the final output features; the calculation formula is as follows:

[0059] ;

[0060] In the above formula, P represents the bottom layer prediction map, S represents the feature map of the current layer, LN represents layer normalization, Cg represents the convolution operation using a Gaussian kernel and setting the bias to zero, and f(·) is a normalization function used to map the Gaussian blurred image to the range [0, 1].

[0061] Among them, the current layer feature map of the fourth layer is spliced ​​with the compensated feature map of the downsampling feature compensation module of the third layer, and then processed to obtain the prediction map P4, which is used as the deep prediction map as input to the overall attention enhancement module of the third layer; the features processed by the overall attention enhancement module of the first layer include the current layer feature map and the deep prediction map P2.

[0062] In step 3, the hierarchical fusion module introduces depthwise separable convolution. During the fusion process, the input features of the hierarchical fusion module first undergo point-by-point convolution and channel-by-channel convolution, followed by batch normalization and activation using the GELU activation function. The classifier is then used to perform one-dimensional convolution on the channel dimension for dimensionality reduction and compression, and a normalized probability map is output using the Sigmoid activation function to represent the spatial distribution of the forged area.

[0063] Finally, the network model is optimized by the focal loss function, which is expressed as:

[0064] ;

[0065] Among them, P is the prediction map, G is the label, α is the category balance factor and β is the focusing parameter.

[0066] In order to further verify the technical effect of the present invention, this embodiment Figure 3 The first column of the three images at the top, middle and bottom are used to locate the tampered area. Figure 3 The second column is a result diagram of detecting the tampered area of ​​the tampered image using the present invention. Figure 3 The third column is the real tampered area map of the tampered image, through Figure 3 It can be seen from the comparison that the positioning accuracy of the present invention is higher.

Claims

1. A method for locating image tampering areas using a bidirectional interactive network, characterized in that: For the input original image, perform the following steps: Step 1: Use the backbone network to perform multi-scale feature extraction to obtain feature maps of different levels of semantics and detail information; Step 2: First, the feature maps of each layer are input into the corresponding downsampling feature compensation module and overall attention enhancement module respectively; The downsampling feature compensation module reorganizes, splices, group normalizes, and dynamically weights the feature map of the current layer, and then combines it with the feature map of the current layer to obtain a compensated feature map. The specific execution process of the downsampling feature compensation module is as follows: First, the input feature map is reorganized using a phase-interval sampling strategy to generate four complementary sub-feature maps. The calculation formulas for the four sub-feature maps are as follows: ; ; ; ; in, Representatives in The pixel value at position, Represents four sub-feature maps; Next, the sub-feature map Splice by channel and sub-feature map The two feature maps are spliced ​​in the channel direction and the dimensionality reduction of the spliced ​​two feature maps is performed by 1×1 convolution. Subsequently, the group normalization operation is introduced and the two feature maps after dimensionality reduction are enhanced in combination with the learnable convolution kernel to obtain two sets of compensation features. After that, the learnable dynamic factor is used to The two sets of compensation features are weightedly fused; finally, the weighted fusion result is combined with the original input feature map through residual connection to generate the corresponding compensation feature map; The overall attention enhancement module performs attention guidance enhancement on the current layer feature map and the bottom layer prediction map, and then fuses them with the current layer feature map to obtain the attention feature map; the prediction map used by the overall attention enhancement module of the first layer The method for obtaining is as follows: the compensated feature map obtained by the downsampling feature compensation module of the first layer is input into the second layer and obtained through training; then, starting from the second layer, the feature map of the current layer is spliced ​​with the compensated feature map of the previous layer to generate the enhanced feature; The specific processing method of the overall attention enhancement module of each layer is as follows: First, the underlying prediction map is Gaussian blurred and the result is normalized to generate an attention weight map. The obtained attention weight map is then element-wise multiplied with the current layer feature map to obtain the enhanced features. The enhanced features are then added to the original features, and the fusion results are normalized through layer normalization to obtain the final output features. The calculation formula is as follows: ;k=2,3,4; In the above formula represents the prediction map, S represents the feature map of the current layer, LN represents layer normalization, Cg represents the convolution operation using a Gaussian kernel and setting the bias to zero, and f(·) is a normalization function used to map the Gaussian blurred image to the range of [0, 1]; Among them, the current layer feature map of the fourth layer is spliced ​​with the compensation feature map of the downsampling feature compensation module of the third layer, and then processed to obtain the prediction map , the prediction graph The prediction map is used as input to the overall attention enhancement module of the third layer; the features processed by the overall attention enhancement module of the first layer include the current layer feature map and the prediction map ; Step 3: Decode the features layer by layer through the multi-level prediction map and hierarchical fusion module to generate the positioning results of the forged area.

2. The method for locating image tampering areas using a two-way interactive network according to claim 1, characterized in that: In step 3, the hierarchical fusion module introduces depthwise separable convolution. During the fusion process, the input features of the hierarchical fusion module first undergo point-by-point convolution and channel-by-channel convolution, followed by batch normalization and activation using the GELU activation function. The classifier is then used to perform one-dimensional convolution on the channel dimension for dimensionality reduction and compression, and a normalized probability map is output using the Sigmoid activation function to represent the spatial distribution of the forged area. Finally, the network model is optimized by the focal loss function, which is expressed as: ; Among them, P is the final output graph , G is the label, α is the category balance factor, and β is the focusing parameter.

Citation Information

Patent Citations

  • Method and device for positioning tampered area in tampered image, terminal and storage medium

    CN113962941A

  • Image tampering detection and positioning method and system based on multilevel feature learning

    CN117853397A

  • Image tampering positioning method fusing multi-level multi-scale and boundary information

    CN117893858A

  • Infrared weak and small target detection network design method based on feature compensation

    CN118470303A