Multi-scale Image Tampering Detection Method Based on Hybrid Attention Mechanism
The hybrid attention mechanism with multi-scale context refinement addresses the limitations of existing image forgery detection by enhancing semantic features and refining predictions, leading to improved localization of tampered regions.
Patent Information
- Application Number
- CN202210793450.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-07
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-07-07
AI Technical Summary
The prior art is difficult to effectively detect and accurately locate the tampering area under various image tampering methods. The traditional method has great limitations and is difficult to adapt to unknown tampering types.
A multi-scale image tamper detection method based on a hybrid attention mechanism is adopted. By building a hybrid attention module that integrates channels and spatial attention, semantic information is enhanced, and multi-scale features are refined. Combined with the ResNet50 basic network architecture, the image tamper detection model is trained and the tampering area mask map is output.
It improves the positioning accuracy of the tampering area, effectively eliminates false positive and false negative interference, and realizes accurate detection of various tampering methods.
Smart Images

Figure CN115578626B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of image processing and computer vision, and in particular to a multi-scale image tampering detection method based on a hybrid attention mechanism. Background Art
[0002] With the rapid development of digital technology, digital images, with their intuitive and vivid characteristics, play an important role as a carrier in the process of people obtaining and transmitting information, and have penetrated into all aspects of social life. However, with the continuous updating and development of image processing technology, a series of powerful image editing software such as Adobe Photoshop and Meitu Xiuxiu have emerged. This allows some users who have not received professional image processing technology training to easily beautify and entertain images according to their own needs, to the point where they are indistinguishable from the real thing.
[0003] Image tampering refers to the use of image editing software to process the original image in order to change the semantic information of the image. Common image tampering includes three categories: (1) copy-move tampering: refers to cropping a certain area in the original image and moving it to another area of the same image; (2) splicing tampering: copying and pasting a certain area on one image to another image; (3) removal tampering: deleting certain areas on an image and replacing them with areas that are the same as the background. In order to make the tampered area less noticeable and enhance the authenticity of the tampered image, the tampered image is often subjected to a series of post-processing operations such as rotation, compression, boundary processing, and brightness adjustment, making it difficult for the human eye to identify the tampered area. The goal of image tampering detection is to detect different types of tampering and accurately locate the tampered area in the image at the pixel level.
[0004] With the diversification of image tampering methods, it is often difficult to identify tampered images with naked eye after they have been "carefully" processed. Early researchers used traditional methods to extract features from tampered images, and achieved image tampering detection through feature comparison analysis methods, or extracted watermarks or digital signatures after embedding secret information in the original image to determine whether the image has been tampered with. However, in real life, we cannot predict the type of image tampering in advance, and these traditional methods have great limitations. They often only target a specific image attribute, and it is difficult to detect multiple or unknown tampering methods based on these features and accurately locate the tampered area. Therefore, it is of great practical significance and value to realize a universal and effective image tampering detection technology.
[0005] With the excellent achievements of deep learning in fields such as semantic segmentation and object detection, many researchers have also applied deep learning technology to image forgery detection. Image forgery detection is different from the semantic segmentation task. Image forgery detection identifies the forged regions in the image. Different forgery operations make the image features vary greatly, and the forged regions are often irregular, which may be a combination of multiple semantic-level objects or a removed region. Therefore, correctly segmenting the forged region depends more on extracting forgery features rather than semantic content. How to utilize the characteristics of the differences in color, intensity, and noise distribution between the forged region and the non-forged region to design and train a network that can accurately locate the forged region has certain challenges and research significance. Summary of the Invention
[0006] In view of this, the purpose of the present invention is to provide a multi-scale image forgery detection method based on a hybrid attention mechanism, which effectively improves the accuracy of locating the forged region.
[0007] To achieve the above purpose, the present invention adopts the following technical solutions:
[0008] A multi-scale image forgery detection method based on a hybrid attention mechanism, comprising the following steps:
[0009] Step S1: Obtain a forgery dataset and divide it into a training set and a test set, and then perform data preprocessing on the forged images and labels in the training set;
[0010] Step S2: Construct a hybrid attention module that combines channel attention and spatial attention to enhance the semantic information of the forged image and obtain an initial prediction map of the forged region;
[0011] Step S3: Construct a refinement module that integrates context information and refine the initial prediction map of the forged region using multi-scale features;
[0012] Step S4: Construct and train a multi-scale image forgery detection model based on the hybrid attention mechanism;
[0013] Step S5: Input the forged image into the trained multi-scale image forgery detection model based on the hybrid attention mechanism and output the corresponding forged region mask map.
[0014] Further, the specific steps of step S1 are as follows:
[0015] Step S11: Obtain a forgery dataset and divide it into a training set and a test set according to a preset ratio;
[0016] Step S12: Perform data augmentation on the images and labels in the training set to increase the number of samples in the dataset;
[0017] Step S13: Preprocess the data-augmented images and labels, including resizing them to a fixed size, performing normalization operations, and converting the data into a standard normal distribution.
[0018] Further, the specific steps of Step S2 are as follows:
[0019] Step S21: Take the feature map F4 with a dimension of C×H×W from the previous module as the input of the channel attention module, where C, H, and W represent the number of channels, height, and width of the feature map respectively; obtain queries Q, keys K, and values V by changing the dimension of the feature map F4, and the dimensions of Q, K, and V are all C×N, where N = H×W represents the number of pixels in the image;
[0020] Step S22: Calculate the similarity by performing matrix multiplication on the transposed matrices of Q and K to obtain the attention weights, and normalize the weights using the Softmax layer to obtain the channel attention weight matrix X. The specific expression is:
[0021]
[0022] where, represents matrix multiplication, and Softmax(·) represents the Softmax activation function; the dimension of X is C×C, where each element x ij represents the influence of the j-th channel on the i-th channel;
[0023] Then, multiply the obtained channel attention weight matrix X by the value V and adjust its dimension to C×H×W; connect the obtained result with the input feature map F4 through a residual structure to enhance the semantic features and obtain the channel attention feature map F c The specific expression is:
[0024]
[0025] where γ is a learnable scaling parameter, initialized to 1, and the dimension of F c is C×H×W;
[0026] Step S23: Take the channel attention feature map F c with a dimension of C×H×W as the input of the spatial attention module. First, apply three 1×1 convolutional layers to the input feature F c to change its number of channels and obtain three new queries Q′, keys K′, and values V′. The dimensions of Q′ and K′ are C / 8×N, and the dimension of V′ is C×N. The specific expression is:
[0027] Q′ = w1(F c) + b1
[0028] K′ = w2(F c ) + b2
[0029] V′ = w3(F c ) + b3
[0030] where w1 and b1, w2 and b2, w3 and b3 are the weights and biases of three different 1×1 convolutional layers respectively;
[0031] Step S24: Calculate the similarity by performing matrix multiplication on the transpose Q′ of Q′ T and K′ to obtain the attention weights, and normalize the weights using the Softmax layer to obtain the spatial attention weight matrix X′. The specific expression is:
[0032]
[0033] where, represents matrix multiplication, and Softmax(·) represents the Softmax activation function. The dimension of X′ is N×N, where each element x′ ij represents the influence of the j-th pixel on the i-th pixel;
[0034] Then, perform matrix multiplication on the transpose X′ of the obtained spatial attention weight matrix X′ T and the value V′, and adjust its dimension to C×H×W. Connect the obtained result and the channel attention feature map F c through a residual structure to enhance the semantic features and obtain the spatial attention feature map F p . The specific expression is:
[0035]
[0036] where γ′ is a learnable scale parameter, initialized to 1, and the dimension of F p is C×H×W;
[0037] Step S25: Concatenate the channel attention map F c obtained in Step S22 and the spatial attention map F p obtained in Step S24 with the input feature map F4. The dimension of the concatenated feature map is 3C×H×W, and then change its dimension to C×H×W through a 1×1 convolutional layer. The specific expression is:
[0038] F′ = Concat(F4, F p , F c )
[0039] F m= w3(F′) + b3
[0040] Where Concat(·) represents concatenating features in a new dimension, w3 and b3 are the weights and biases of the 1×1 convolutional layer, and F m is the feature map with enhanced semantics after passing through the hybrid attention module;
[0041] Finally, apply a convolutional layer with a kernel size of 7×7 and a padding of 3 to F m to obtain the initial prediction map of the tampered area
[0042] Furthermore, the specific steps of step S3 are as follows:
[0043] Step S31: The refinement module for fusing context information has a total of three inputs, including the feature map F h from the previous level, the feature map F l from the current level, and the prediction map from the previous level
[0044] Upsample the prediction map from the previous level so that its resolution size is the same as that of the feature map F l from the current level, and normalize it using the Sigmoid layer;
[0045] Then multiply the normalized result and its inverted result element-wise with the feature map F l from the current level to generate the foreground attention feature map F fa and the background attention feature map F ba , and the calculation formulas are as follows:
[0046]
[0047] F fa = F l × y up
[0048] F ba = F l × (1 - y up )
[0049] Where U represents bilinear upsampling, Sigmoid(·) represents the Sigmoid activation function, and × represents element-wise multiplication;
[0050] Step S32: The refinement module for fusing context information contains two context reasoning modules. Apply the foreground attention feature map F fa and the background attention feature map F baare fed into the above-mentioned context inference module in parallel to obtain false positive interference F fpd and false negative interference F fnd ;
[0051] Step S33: Adjust the dimension of the feature map F h at the previous level to be consistent with the dimension of the feature map F l at the current level. Subtract the false positive interference F fpd and the false negative interference F fnd element-wise obtained in Step S32 to eliminate false positive interference and add them element-wise to eliminate false negative interference, thereby correcting the feature map F h at the previous level to obtain a more refined feature map F r , and its calculation formula is:
[0052] F up = U(CBR(F h ))
[0053] F r = BR(F up - αF Fpd )
[0054] F r = BR(F r + βF fnd )
[0055] where α and β are learnable scale parameters, initialized to 1, CBR represents the combination of convolution, batch normalization layer and ReLU activation function, BR represents the combination of batch normalization layer and ReLU function, and U represents bilinear upsampling;
[0056] Finally, apply a convolutional layer with a convolutional kernel size of 7×7 and padding of 3 to F r to obtain a refined prediction map.
[0057] Furthermore, the context inference module consists of four context inference branches. Each branch sequentially passes the input feature map E through a 3×3 convolutional layer to reduce the number of channels to 1 / 4 of the original, a K i ×K i convolutional layer for local feature extraction, and a dilated convolutional layer with a convolutional kernel size of 3×3 and a dilation rate of r i to fuse context information. Among them, the K i and r iThey are {1, 3, 5, 7} and {1, 2, 4, 8} respectively; the output of the i-th (i = 1, 2, 3) branch will be fed into the (i + 1)-th branch, so that the feature map can be further processed under a larger receptive field; finally, the output results of the four context reasoning branches are concatenated in the channel dimension, and a 3×3 convolution is performed for feature fusion to obtain the interference map E′, and its calculation formula is:
[0058] E i_1 = w i_1 (E) + b i_1
[0059]
[0060] E i_3 = w i_3 (E i_2 ) + b i_3
[0061] E′ = w4(Concat(E 1_3 , E 2_3 , E 3_3 , E 4_3 )) + b4
[0062] Among them, E i_1 represents the feature output after passing through the 3×3 convolutional layer for channel reduction in the i-th branch, w i_1 and b i_1 correspond to its weight and bias; E i_2 represents the feature output after passing through the K i ×K i convolutional layer for local feature extraction in the i-th branch, w i_2 and b i_2 correspond to its weight and bias; E i_3 represents the feature output after passing through the dilated convolution for fusing context information in the i-th branch, w i_3 and b i_3 correspond to its weight and bias; Concat(·) represents the concatenation of features in a new dimension, w4 and b4 correspond to the weight and bias of the convolutional layer for feature fusion, and E′ represents the obtained interference map.
[0063] Furthermore, the specific steps of step S4 are as follows:
[0064] Step S41: Taking ResNet50 as the basic network architecture, using its feature extraction network to extract features from the tampered image after preprocessing in Step S1, obtaining four feature maps X1, X2, X3, and X4 with different numbers of channels; reducing the number of channels of the four-level feature maps to 1 / 4 of the original through a convolutional layer, a batch normalization layer, and a ReLU layer to obtain multi-scale feature maps F1, F2, F3, and F4; using the feature map F4 as the input of the hybrid attention module to obtain the semantically enhanced feature map F m and the initial prediction map of the tampered area Input the semantically enhanced feature map F m and the multi-scale feature maps F1, F2, and F3 into three refinement modules in a top-down manner to obtain three refined feature maps F r 1 、F r 2 、F r 3 and the tampered area prediction map Select the prediction map output by the third refinement module as the final prediction result;
[0065] Step S42: The loss function Loss of the hybrid attention module m is obtained by calculating the element-wise binary cross-entropy loss function loss between its output prediction map and the corresponding label value y bce and the element-wise intersection over union loss function loss iou ;
[0066] Step S43: The loss function Loss of the refinement module f is obtained by calculating the weighted binary cross-entropy loss function loss between its output prediction map and the corresponding label value y wbce and the weighted intersection over union loss function loss wiou ;
[0067] Step S44: The total loss function Loss of the multi-scale image tampering detection network model based on the hybrid attention mechanism all The formula is as follows:
[0068]
[0069] where Loss m represents the loss output by the hybrid attention module, represents the loss output by the i-th refinement module;
[0070] According to the total loss function of the multi-scale image tampering detection network model based on the hybrid attention mechanism, use the backpropagation method to calculate the gradients of the parameters in the image tampering detection network model;
[0071] Step S45: Repeat the above steps S41 to S44 in batches until the total loss value calculated in step S44 converges and stabilizes. Save the network parameters to complete the training process of the multi-scale image tampering detection network model based on the hybrid attention mechanism.
[0072] Furthermore, the loss function Loss of the hybrid attention module m is as follows:
[0073] Loss m = loss bce + loss iou
[0074]
[0075]
[0076] The loss function Loss of the refinement module f is as follows:
[0077] Loss f = loss wbce + loss wiou
[0078]
[0079]
[0080]
[0081] where represents the predicted map output by the k-th refinement module, and α ij ∈[0, 1] represents the weight assigned to each pixel, and A ij represents the pixels around the pixel point (i, j).
[0082] Furthermore, the specific content of step S5 is: Input the images in the test set into the trained multi-scale image tampering detection model based on the hybrid attention mechanism, and adjust their sizes to the size of the original tampered image to obtain the corresponding tampered area mask map.
[0083] The present invention has the following beneficial effects compared with the prior art:
[0084] The present invention uses a hybrid attention mechanism to enhance the semantic information of image features, locates the initial tampering area, and then refines the feature maps at different levels in a top-down manner, continuously eliminating false positive and false negative interferences, effectively improving the accuracy of tampering area location. BRIEF DESCRIPTION OF THE DRAWINGS
[0085] Figure 1 is the flowchart of the method of the present invention;
[0086] Figure 2 is the structural diagram of the network model in an embodiment of the present invention;
[0087] Figure 3 is the structural diagram of the hybrid attention module in an embodiment of the present invention;
[0088] Figure 4 is the structural diagram of the refinement module in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0089] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0090] Please refer to Figures 1-4 , the present invention provides a multi-scale image tampering detection method based on a hybrid attention mechanism, including the following steps:
[0091] Step S1: Obtain a tampering dataset and divide it into a training set and a test set, and then perform data preprocessing on the tampered images and labels in the training set;
[0092] Step S2: Construct a hybrid attention module that fuses channel attention and spatial attention to enhance the semantic information of the tampered image and obtain an initial prediction map of the tampering area;
[0093] Step S3: Construct a refinement module that fuses context information and use multi-scale features to refine the initial prediction map of the tampering area;
[0094] Step S4: Construct and train a multi-scale image tampering detection model based on the hybrid attention mechanism;
[0095] Step S5: Input the tampered image into the trained multi-scale image tampering detection model based on the hybrid attention mechanism, and output the corresponding tampering area mask map.
[0096] In this embodiment, the specific steps of step S1 include the following steps:
[0097] Step S11: Adopt the tampering dataset CASIA, use CASIAv2 as the training set, and CASIAv1 as the test set: the training set includes 5123 pictures, and the test set includes 921 pictures;
[0098] Step S12: Perform data augmentation on the images and labels in the training set to increase the number of samples in the dataset, including random horizontal flipping and color jittering;
[0099] Step S13: Preprocess the images after data augmentation in Step S12 and convert them into the input of the image forgery detection network. First, scale the size of the images to 512×512 pixels, and then perform normalization on the data to convert the data into a standard normal distribution; to ensure that the size and position of the forged area in the label correspond to the forged image, the same operations are also performed on the label while performing each step of data augmentation and image preprocessing.
[0100] In this embodiment, the specific steps of the said Step S2 include the following steps:
[0101] Step S21: Use the feature map F4 with dimensions C×H×W from the previous module as the input of the channel attention module, where C, H, and W represent the number of channels, height, and width of the feature map respectively. Obtain the query Q, key K, and value V by changing the dimensions of the feature map F4. The dimensions of Q, K, and V are all C×N, where N = H×W represents the number of pixels in the image;
[0102] Step S22: Calculate the similarity to obtain the attention weights by performing matrix multiplication on the transposed matrices of Q and K, and normalize the weights using the Softmax layer to obtain the channel attention weight matrix X. The specific expression is:
[0103]
[0104] Among them, represents matrix multiplication, and Softmax(·) represents the Softmax activation function. The dimension of X is C×C, where each element x ij represents the influence of the j-th channel on the i-th channel;
[0105] Then, multiply the obtained channel attention weight matrix X by the value V and adjust its dimension to C×H×W. Connect the obtained result and the input feature map F4 through a residual structure to enhance the semantic features and obtain the channel attention feature map F c , and the specific expression is:
[0106]
[0107] Among them, γ is a learnable scale parameter, initialized to 1, and the dimension of F c is C×H×W;
[0108] Step S23: Use the channel attention feature map F with dimensions C×H×W cAs the input of the spatial attention module, first, the input feature F c Apply three 1×1 convolutional layers respectively to change its number of channels and its dimension to obtain three new queries Q′, keys K′, and values V′, where the dimensions of Q′ and K′ are C / 8×N, and the dimension of V′ is C×N; the specific expressions are:
[0109] Q′ = w1(F c ) + b1
[0110] K′ = w2(F c ) + b2
[0111] V′ = w3(F c ) + b3
[0112] where w1 and b1, w2 and b2, w3 and b3 are the weights and biases of three different 1×1 convolutional layers respectively;
[0113] Step S24: Calculate the similarity by performing matrix multiplication on the transpose Q′ T of Q′ and K′ to obtain the attention weights, and normalize the weights using the Softmax layer to obtain the spatial attention weight matrix X′. The specific expression is:
[0114]
[0115] where, represents matrix multiplication, and Softmax(·) represents the Softmax activation function. The dimension of X′ is N×N, where each element x′ ij represents the influence of the j-th pixel on the i-th pixel;
[0116] Then, perform matrix multiplication on the transpose X′ T of the obtained spatial attention weight matrix X′ and the value V′, and adjust its dimension to C×H×W. Connect the obtained result with the channel attention feature map F c through a residual structure to enhance the semantic features and obtain the spatial attention feature map F p , and the specific expression is:
[0117]
[0118] where γ′ is a learnable scale parameter, initialized to 1, and the dimension of F p is C×H×W;
[0119] Step S25: Combine the channel attention map F c obtained in step S22 and the spatial attention map F pConcatenate with the input feature map F4. The dimension of the concatenated feature map is 3C×H×W, and then change its dimension to C×H×W through a 1×1 convolutional layer. The specific expression is:
[0120] F′ = Concat(F4, F p , F c )
[0121] F m = w3(F′) + b3
[0122] where Concat(·) represents concatenating features in a new dimension, w3 and b3 are the weights and biases of the 1×1 convolutional layer, and F m is the feature map with enhanced semantics after passing through the hybrid attention module;
[0123] Finally, apply a convolutional layer with a kernel size of 7×7 and a padding of 3 to F m to obtain the initial prediction map of the tampered area
[0124] In this embodiment, step S3 specifically includes the following steps:
[0125] Step S31: The refinement module for fusing context information has a total of three inputs, including the feature map F h from the previous level, the feature map F l at the current level, and the prediction map
[0126] Upsample the prediction map from the previous level to make its resolution size the same as that of the feature map F l at the current level, and normalize it using the Sigmoid layer;
[0127] Then multiply the normalized result and its inverted result element-wise with the feature map F l at the current level to generate the foreground attention feature map F fa and the background attention feature map F ba , and the calculation formulas are as follows:
[0128]
[0129] F fa = F l × y up
[0130] F ba = F l × (1 - y up )
[0131] Among them, U represents bilinear upsampling, Sigmoid(·) represents the Sigmoid activation function, and × represents element-wise multiplication;
[0132] Step S32: The refinement module for fusing context information contains two context reasoning modules. The foreground attention feature map F obtained in step S31 fa and the background attention feature map F ba are fed into the above-mentioned context reasoning modules in parallel to respectively obtain false positive interference F fpd and false negative interference F fnd ;
[0133] Step S33: Adjust the dimension of the feature map F at the previous level h to be consistent with the dimension of the feature map F at the current level l . Subtract false positive interference F obtained in step S32 fpd and false negative interference F fnd element-wise to eliminate false positive interference and add them element-wise to eliminate false negative interference, thereby correcting the feature map F at the previous level h to obtain a more refined feature map F r , and its calculation formula is:
[0134] F up = U(CBR(F h ))
[0135] F r = BR(F up -αF fpd )
[0136] F r = BR(F r +βF fnd )
[0137] where α and β are learnable scale parameters, initialized to 1, CBR represents the combination of convolution, batch normalization layer, and ReLU activation function, BR represents the combination of batch normalization layer and ReLU function, and U represents bilinear upsampling;
[0138] Finally, apply a convolutional layer with a kernel size of 7×7 and a padding of 3 to F r to obtain the refined prediction map.
[0139] Furthermore, the context reasoning module consists of four context reasoning branches. Each branch sequentially passes the input feature map E through a 3×3 convolutional layer to reduce the number of channels to 1 / 4 of the original, a K i ×K i convolutional layer for local feature extraction, and a convolutional layer with a kernel size of 3×3 and a dilation rate of ri The dilated convolution fuses the context information, where K of the i-th (i = 1, 2, 3, 4) branch i and r i are {1, 3, 5, 7} and {1, 2, 4, 8} respectively; the output of the i-th (i = 1, 2, 3) branch will be sent to the (i + 1)-th branch, so that the feature map is further processed under a larger receptive field; finally, the output results of the four context reasoning branches are concatenated in the channel dimension, and feature fusion is performed through a 3×3 convolution to obtain the interference map E′, and its calculation formula is:
[0140] E i_1 = w i_1 (E) + b i_1
[0141]
[0142] E i_3 = w i_3 (E i_2 ) + b i_3
[0143] E′ = w4(Concat(E 1_3 , E 2_3 , E 3_3 , E 4_3 )) + b4
[0144] Among them, E i_1 represents the feature output after passing through the 3×3 convolutional layer for channel reduction in the i-th branch, w i_1 and b i_1 correspond to its weights and biases; E i_2 represents the feature output after passing through the K i ×K i convolutional layer for local feature extraction in the i-th branch, w i_2 and b i_2 correspond to its weights and biases; E i_3 represents the feature output after passing through the dilated convolution for fusing context information in the i-th branch, w i_3 and b i_3 correspond to its weights and biases; Concat(·) represents the concatenation of features in a new dimension, w4 and b4 correspond to the weights and biases of the convolutional layer for feature fusion, and E′ represents the obtained interference map.
[0145] In this embodiment, the step S4 specifically includes the following steps:
[0146] Step S41: Taking ResNet50 as the basic network architecture, use its feature extraction network to extract features from the tampered image after preprocessing in Step S1, and obtain four feature maps X1, X2, X3, and X4 with the number of channels being 2048, 1024, 512, and 256 dimensions respectively. Reduce the number of channels of the four-level feature maps to 1 / 4 of the original through the convolutional layer, batch normalization layer, and ReLU layer to obtain multi-scale feature maps F1, F2, F3, and F4, whose numbers of channels are 512, 256, 128, and 64 dimensions respectively. Take the feature map F4 as the input of the hybrid attention module to obtain the semantically enhanced feature map F m and the initial prediction map of the tampered area Take the semantically enhanced feature map F m and the multi-scale feature maps F1, F2, F3 and input them into three refinement modules in a top-down manner to obtain three refined feature maps F r 1 、F r 2 、F r 3 and the tampered area prediction map Select the prediction map output by the third refinement module as the final prediction result;
[0147] Step S42: The loss function Loss of the hybrid attention module m is obtained by calculating the element-wise binary cross-entropy loss function loss between its output prediction map and the corresponding label value y bce and the element-wise intersection over union loss function loss iou ;
[0148] Step S43: The loss function Loss of the refinement module f is obtained by calculating the weighted binary cross-entropy loss function loss between its output prediction map and the corresponding label value y wbce and the weighted intersection over union loss function loss wiou ;
[0149] Step S44: The total loss function Loss of the multi-scale image tampering detection network model based on the hybrid attention mechanism all The formula is as follows:
[0150]
[0151] Among them, Loss m represents the loss output by the hybrid attention module, represents the loss output by the i-th refinement module;
[0152] According to the total loss function of the multi-scale image forgery detection network model based on the hybrid attention mechanism, calculate the gradients of the parameters in the image forgery detection network model by using the backpropagation method, and update the parameters by using the Adam optimization algorithm;
[0153] Step S45: Repeat the above steps S41 to S44 in batches until the total loss value calculated in step S44 converges and stabilizes. A total of 60 epochs are trained. Save the network parameters to complete the training process of the multi-scale image forgery detection network model based on the hybrid attention mechanism.
[0154] Furthermore, the loss function Loss of the hybrid attention module m is as follows:
[0155] Loss m = loss bce + loss iou
[0156]
[0157]
[0158] The loss function Loss of the refinement module f is as follows:
[0159] Loss f = loss wbce + loss wiou
[0160]
[0161]
[0162]
[0163] where represents the predicted map output by the k-th refinement module, α ij ∈ [0, 1] represents the weight assigned to each pixel, and A ij represents the pixel points around the pixel point (i, j).
[0164] In this embodiment, the step S5 specifically includes the following steps:
[0165] Step S51: Input the image in the test set into the trained multi-scale image forgery detection model based on the hybrid attention mechanism, and adjust its size to the size of the original forged image to obtain the corresponding forged region mask map.
[0166] The above are only the preferred embodiments of the present invention, and all equivalent changes and modifications made according to the scope of the patent application of the present invention shall fall within the scope of the present invention.
Claims
1. A multi-scale image forgery detection method based on a hybrid attention mechanism, characterized in that It includes the following steps: Step S1: Obtain the tampered dataset and divide it into a training set and a test set, and then perform data preprocessing on the tampered images and labels in the training set; Step S2: Construct a hybrid attention module that fuses channel attention and spatial attention to enhance the semantic information of the tampered images and obtain an initial prediction map of the tampered regions; Step S3: Construct a refinement module that fuses context information and refine the initial prediction map of the tampered regions using multi-scale features; Step S4: Construct and train a multi-scale image tampering detection model based on the hybrid attention mechanism; Step S5: Input the tampered images into the trained multi-scale image tampering detection model based on the hybrid attention mechanism and output the corresponding tampered region mask map; The specific steps of Step S2 are as follows: Step S21: Use the feature map F4 with dimensions C×H×W from the previous module as the input to the channel attention module, where C, H, and W represent the number of channels, height, and width of the feature map respectively; obtain the query Q, key K, and value V by changing the dimensions of the feature map F4, and the dimensions of Q, K, and V are all C×N, where N = H×W represents the number of pixels in the image; Step S22: Calculate the similarity by performing matrix multiplication on the transposed matrices of Q and K to obtain the attention weights, and use the Softmax layer to normalize the weights to obtain the channel attention weight matrix X. The specific expression is: Among them, represents matrix multiplication, and Softmax(·) represents the Softmax activation function; the dimension of X is C×C, where each element x ij represents the influence of the j-th channel on the i-th channel. Then, multiply the obtained channel attention weight matrix X by the value V and adjust its dimension to C×H×W; connect the obtained result with the input feature map F4 through a residual structure to enhance semantic features and obtain the channel attention feature map F c , and the specific expression is: where γ is a learnable scale parameter, and the dimension of F c is C×H×W; Step S23: Take the channel attention feature map F with dimensions C×H×W c as the input of the spatial attention module. First, for the input feature F c apply three 1×1 convolutional layers respectively to change its number of channels and obtain three new queries Q′, keys K′, and values V′ by changing its dimensions, where the dimensions of Q′ and K′ are C / 8×N, and the dimension of V′ is C×N; the specific expressions are: Q′ = w1(F c ) + b1 K′ = w2(F c ) + b2 V′ = w3(F c ) + b3 where w1 and b1, w2 and b2, w3 and b3 are the weights and biases of three different 1×1 convolutional layers respectively; Step S24: By taking the transpose Q' of Q T and performing matrix multiplication with K' to calculate the similarity to obtain the attention weights, and using a Softmax layer to normalize the weights to obtain the spatial attention weight matrix X', with the specific expression being: Among them, represents matrix multiplication, and Softmax(·) represents the Softmax activation function; the dimension of X′ is N×N, where each element x′ ij represents the influence of the j-th pixel on the i-th pixel; Then, take the transpose X' of the obtained spatial attention weight matrix X' T and perform matrix multiplication with the value V', and adjust its dimension to C×H×W; connect the obtained result with the channel attention feature map F c through a residual structure to enhance semantic features and obtain the spatial attention feature map F p , and the specific expression is: where γ′ is a learnable ratio parameter, initialized to 1, and the dimension of F p is C×H×W; Step S25: Concatenate the channel attention map F obtained in step S22 c and the spatial attention map F obtained in step S24 p with the input feature map F4. The dimension of the concatenated feature map is 3C×H×W, and then change its dimension to C×H×W through a 1×1 convolutional layer. The specific expression is: F′ = Concat(F4, F p , F c ) F m = w3(F') + b3 Among them, Concat(·) represents the concatenation of features in a new dimension, w3 and b3 are the weights and biases of the 1×1 convolutional layer, and F m is the feature map with enhanced semantics after passing through the hybrid attention module; Finally, for F m Apply a convolutional layer with a convolutional kernel size of 7×7 and a padding of 3 to obtain an initial prediction map of the tampered area The specific steps of Step S4 are as follows: Step S41: Taking ResNet50 as the basic network architecture, using its feature extraction network to extract features from the tampered image after preprocessing in Step S1, obtaining four feature maps X1, X2, X3, X4 with different numbers of channels; reducing the number of channels of the four-level feature maps to 1 / 4 of the original through convolutional layers, batch normalization layers, and ReLU layers to obtain multi-scale feature maps F1, F2, F3, F4; using the feature map F4 as the input of the hybrid attention module to obtain the semantically enhanced feature map F m and the initial prediction map of the tampered area Input the semantically enhanced feature map F m and the multi-scale feature maps F1, F2, F3 into three refinement modules in a top-down manner, respectively obtaining three refined feature maps F r 1 、F r 2 、F r 3 and the tampered area prediction map Select the prediction map output by the third refinement module as the final prediction result; Step S42: Loss function Loss of the hybrid attention module m By taking the predicted map output by it Calculate the per-element binary cross-entropy loss function loss with the corresponding label value y bce And the per-element intersection over union loss function loss iou Obtained; Step S43: Loss function Loss of the refinement module f By using the predicted map output by it to calculate the weighted binary cross-entropy loss function loss with the corresponding label value y wbce and the weighted intersection over union loss function loss wiou obtained; Step S44: Total loss function Loss of the multi-scale image tampering detection network model based on the hybrid attention mechanism all The formula is as follows: Among them, Loss m represents the loss output by the hybrid attention module, and represents the loss output by the i-th refinement module; According to the total loss function of the multi-scale image tampering detection network model based on the hybrid attention mechanism, use the backpropagation method to calculate the gradients of the parameters in the image tampering detection network model; Step S45: Repeat the above steps S41 to S44 in batches until the total loss value calculated in Step S44 converges and stabilizes, save the network parameters, and complete the training process of the multi-scale image tampering detection network model based on the hybrid attention mechanism; The loss function Loss of the hybrid attention module m is as follows: Loss m = loss bce + loss iou Loss function of the refinement module f As follows: Loss f = loss wbce + loss wiou Among them represents the predicted map output by the k-th refinement module, and α ij ∈ [0,1] represents the weight assigned to each pixel, and A ij represents the pixels around the pixel point (i,j).
2. The multi-scale image forgery detection method based on the hybrid attention mechanism according to claim 1, characterized in that The specific steps of Step S1 are as follows: Step S11: Obtain the tampered dataset and divide it into a training set and a test set according to a preset ratio; Step S12: Perform data augmentation on the images and labels in the training set to increase the number of samples in the dataset; Step S13: Perform preprocessing on the images and labels after data augmentation, including resizing to a fixed size, normalization operation, and converting the data into a standard normal distribution.
3. The multi-scale image forgery detection method based on a hybrid attention mechanism according to claim 1, wherein The specific steps of Step S3 are as follows: Step S31: The refinement module for integrating context information has three inputs, including the feature map F at the previous level h , the feature map F at the current level l , and the prediction map at the previous level The prediction map of the previous level Upsample it to make its resolution size the same as that of the feature map F of the current level l and normalize it using a Sigmoid layer; Then, the normalized result and its inverted result are multiplied element-wise with the feature map F at the current level l to generate the foreground attention feature map F fa and the background attention feature map F ba , respectively. The calculation formula is as follows: F fa = F l × y up F ba = F l × (1 - y up ) where U represents bilinear upsampling, Sigmoid(·) represents the Sigmoid activation function, and × represents element-wise multiplication; Step S32: The refinement module for fusing context information includes two context reasoning modules, and the foreground attention feature map F obtained in step S31 fa and the background attention feature map F ba are fed into the above context reasoning modules in parallel to obtain false positive interference F fpd and false negative interference F fnd ; Step S33: Adjust the dimension of the feature map F at the upper level to be the same as that of the feature map F at the current level, and perform element-wise subtraction on the false positive interference F obtained in step S32 and the false negative interference F to eliminate the false positive interference, and perform element-wise addition to eliminate the false negative interference, so as to correct the feature map F at the upper level and obtain a more refined feature map F, and its calculation formula is: h l fpd fnd h r F up = U(CBR(F h )) F r = BR(F up - αF fpd ) F r = BR(F r + βF fnd ) where α and β are learnable scale parameters, initialized to 1, CBR represents the combination of convolution, batch normalization layer, and ReLU activation function, BR represents the combination of batch normalization layer and ReLU function, and U represents bilinear upsampling; Finally, for F r Apply a convolutional layer with a convolutional kernel size of 7×7 and a padding of 3 to obtain a refined prediction map; The context reasoning module consists of four context reasoning branches. Each branch sequentially passes the input feature map E through a 3×3 convolutional layer to reduce the number of channels to 1 / 4 of the original, a K i ×K i convolutional layer for local feature extraction, and a dilated convolution with a convolutional kernel size of 3×3 and a dilation rate of r i to fuse context information. Among them, the K i and r i of the i-th (i = 1, 2, 3, 4) branch are {1, 3, 5, 7} and {1, 2, 4, 8} respectively; the output of the i-th (i = 1, 2, 3) branch will be fed into the (i + 1)-th branch, so that the feature map is further processed under a larger receptive field; finally, the output results of the four context reasoning branches are concatenated in the channel dimension and passed through a 3×3 convolution for feature fusion to obtain the interference map E′, and its calculation formula is: E i_1 = w i_1 (E) + b i_1 E i_3 = w i_3 (E i_2 ) + b i_3 E′ = w4(Concat(E 1_3 , E 2_3 , E 3_3 , E 4_3 )) + b4 Among them, E i_1 represents the feature output after passing through the 3×3 convolutional layer for channel reduction in the i-th branch, where w i_1 and b i_1 correspond to its weights and biases; E i_2 represents the feature output after passing through the K i ×K i convolutional layer for local feature extraction in the i-th branch, where w i_2 and b i_2 correspond to its weights and biases; E i_3 represents the feature output after passing through the dilated convolution for fusing context information in the i-th branch, where w i_3 and b i_3 correspond to its weights and biases; Concat(·) represents concatenating features in a new dimension, w4 and b4 correspond to the weights and biases of the convolutional layer for feature fusion, and E′ represents the resulting interference map.
4. The multi-scale image forgery detection method based on the hybrid attention mechanism according to claim 1, characterized in that The specific steps of Step S5 are as follows: Input the images in the test set into the trained multi-scale image tampering detection model based on the hybrid attention mechanism and adjust their sizes to the size of the original tampered images, then the corresponding tampered region mask map can be obtained.
Citation Information
Patent Citations
Image stitching tampering detection method
CN111080629A
Face tampering detection method and system based on multi-source clues and mixed attention
CN112818862A