A Multi-scale Image Copy-paste Tampering Detection Method Based on FPN

By using the multi-scale feature extraction method of FPN and residual neural networks in image copy-paste tamper detection, combined with attention mechanism and autocorrelation processing, the problems of single vision and poor robustness in the existing methods are solved, and image tamper detection with high accuracy and robustness are achieved.

CN117152595BActive Publication Date: 2025-06-17NANCHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311046652.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-19
Publication Date
2025-06-17
Estimated Expiration
2043-08-19

AI Technical Summary

Technical Problem

The existing image copy-paste tamper detection methods have problems such as single field of view and poor robustness, which leads to inaccurate detection and vulnerability to attacks.

Method used

A multi-scale image copy-paste tamper detection method based on feature pyramid network (FPN) and residual neural network is used, and combined with attention mechanism, autocorrelation module and mask decoder, high-rootability feature extraction and image tampering area positioning are performed.

Benefits of technology

High-accuracy binary classification detection of images is realized, and it can effectively handle image tampering areas and various attacks of different sizes, such as noise, Gaussian transformation and geometric scaling, improving the robustness and efficiency of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117152595B_ABST
    Figure CN117152595B_ABST
Patent Text Reader

Abstract

The present invention provides a multi-scale image copy-paste forgery detection method based on FPN, which combines the use of a Feature Pyramid Network and a Residual Neural Network to perform multi-scale and high-robustness feature extraction on images, uses an attention mechanism for feature selection, calculates similar regions in the self-correlation calculation area, and performs deconvolution to decode the image into a mask of the same size as the detected image, successfully locating the forgery regions in the image and classifying whether the image has been forged. The method of the present invention has strong robustness in multi-scale feature extraction, high detection accuracy and high detection efficiency. It detects whether there is copy-paste forgery in the image and effectively identifies the copy-paste forgery regions existing in the image. It protects the authenticity and usability of the image from being damaged.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image tampering detection, and particularly relates to a multi-scale image copy-paste tampering detection method based on FPN. Background Art

[0002] Traditional CMFD methods can be roughly divided into two categories: block division-based methods and key feature-based methods. For block division-based methods, the input image is first divided into overlapping or non-overlapping blocks. The shape of the block can be rectangular or circular. In the case of overlapping blocks, assuming an input image of size M×N pixels, the block size is set to b×b pixels, and then it is moved one pixel at a time across the entire image from left to right and top to bottom to form a set of blocks. The total number of overlapping blocks in the image is {(M - b + 1)×(N - b + 1)}. Then, robust feature information such as the texture, hue, and edges of the local image is extracted from each block. Finally, the extracted feature information is sorted or arranged, and by comparing the similarity of adjacent feature pairs, it is determined whether the input image has been tampered with. Classic methods for extracting robust features include those based on discrete wavelet transform, discrete cosine transform coefficients, perceptual error metrics, etc. In key-point-based techniques, the basic idea is that during copy-paste operations, since the copy and paste operations change the local features of the blocks in the image, by detecting changes in the key points in the image, it is possible to detect whether the image has been tampered with. The process of this method is to perform key point detection on the input image using a specific algorithm to obtain a set of key point coordinates, and to describe the features of the extracted key points to obtain the feature vectors of the key points. All key point feature vectors are matched, and by calculating features such as the distance and angle of the matching points, it is determined whether there is a copy-paste operation. Commonly used image key point features include methods based on SIFT feature point matching, SURF key point detection algorithms, and ORB-based key point tampering detection methods. Traditional methods have some common deficiencies: they require manual adjustment of various parameters, each module needs to be optimized separately, and the workload is large, etc. Therefore, in traditional image copy-paste tampering detection methods, block-based and key-point-based detection techniques are often combined to improve the detection effect, but at the same time, the workload and computational complexity are increased.

[0003] Image copy - paste forgery detection method based on deep learning. Deep learning techniques are used to extract image features and perform distance measurement on these features. With the ability of neural networks to automatically learn feature representations, the input image can be subjected to multi - layer non - linear transformations, thereby obtaining a more advanced feature expression ability. Convolutional neural networks are good at processing images, have the function of image feature extraction, and can extract image features and calculate the distance measurement between them. Specifically used techniques include generative adversarial networks, extracting image features based on CovLSTM, and methods based on deep convolutional neural networks. Existing deep - learning methods have problems such as inaccurate detection due to a single field of view, being vulnerable to targeted attacks (for example, adding some noise or interference information can cause the model to output incorrect results, resulting in poor robustness), and high computational resource consumption and insufficient performance. Summary of the Invention

[0004] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a multi - scale image copy - paste forgery detection method based on FPN to solve the problems of single field of view and poor robustness in existing methods. The present invention has high accuracy in binary classification detection of images. On different data sets, for image forgery regions of different sizes, in the face of attacks such as Gaussian noise and geometric scaling, various performance indicators still remain at a high level compared with existing methods, and it can be better applied in actual scenarios.

[0005] To achieve the above - mentioned purpose, the present invention provides the following technical solution: A multi - scale image copy - paste forgery detection method based on FPN, comprising the following steps:

[0006] Step 1: Divide the detection image set into a training set, a validation set, and a test set, and pre - process the images so that the pixel value size is 224 * 224 * 3;

[0007] Step 2: Construct an image multi - scale copy - paste forgery detection model, including an FPN feature extraction module, an attention mechanism module, an autocorrelation module, and a mask decoder module;

[0008] Step 3: Input the detected image into the FPN feature extraction module for feature extraction. FPN includes two processes. The first part is the bottom-up process, which uses the ResNet50 network as the convolutional feature extraction network. ResNet50 learns and extracts features at different levels during training, and the feature pyramid fusion is performed on the different-scale features extracted by each feature extraction layer. The second part of FPN is the fusion process of top-down and lateral connections. In the top-down process, FPN uses a pyramidal structure to construct feature maps of different scales. In each layer, the small-size feature map of the previous layer is upsampled to the same size as the current layer and fused with the feature map of the current layer. The lateral connections in FPN improve the semantic information of the feature map by directly adding the feature map of the previous layer and the feature map of the current layer.

[0009] Step 4: The attention mechanism processes the high-level image features obtained in Step 3, including the channel attention mechanism and the spatial attention mechanism. The channel attention mechanism adaptively adjusts the importance of each channel, and the spatial attention mechanism enables the neural network to more precisely focus on the important pixel regions in the image and ignore the unimportant regions.

[0010] Step 5: The self-correlation module passes the different-scale feature maps output by FPN through the attention mechanism and then enters the self-correlation process to calculate the pixel-level distance vector to obtain the feature similarity score. Through the percentage pooling layer, the similar feature matching statistics are performed to locate the similar regions in the image.

[0011] Step 6: The role of the mask decoder module is to fuse the different-sized feature maps obtained after the self-correlation process, decode and restore the image to the original image size through upsampling, and divide the image region into similar regions and original regions through the activation function to output the Ground Truth mask of the image.

[0012] Step 7: After obtaining the Ground Truth mask of the image, determine whether there are similar regions in the image, and for the image with tampered regions, locate the image tampering regions.

[0013] Furthermore, in Step 5, the pairwise similarity scores between any two feature pixels are calculated using the pixel distance vector. Specifically, for any two points n and m, their pixel values are (i, j) and (i', j') respectively, and their feature pixels are fx(i, j) and fx(i', j') respectively. The Pearson correlation coefficient ρ is used to calculate the feature similarity equation.

[0014]

[0015] Further, Step 6 includes: First, perform upsampling fusion processing on the self-correlation feature images of different scales output from the self-correlation module, and then use the BN-Inception convolutional neural network and BilinearUpPool2D upsampling in an alternating manner to restore the feature map to the original resolution.

[0016] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0017] The present invention combines the use of a Feature Pyramid Network and a Residual Neural Network to perform multi-scale and high-robustness feature extraction on images, uses an attention mechanism for feature selection, calculates the similarity of self-correlation calculation regions, and uses deconvolution to restore the image to a mask of the same size as the detected image for decoding, successfully locating the tampered region in the image and classifying whether the image has been tampered with. The method of the present invention has strong robustness in multi-scale feature extraction, high detection accuracy, and high detection efficiency. It can detect whether there is copy-paste tampering in the image and effectively identify the copy-paste tampered region in the image, protecting the authenticity and usability of the image from being damaged. Description of the Drawings

[0018] Figure 1 is the network structure diagram of the present invention;

[0019] Figure 2 is the schematic diagram of ResNet50+FPN feature extraction. Detailed Embodiments

[0020] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. The specific embodiments described herein are only used to explain the technical solutions of the present invention and are not limited to the present invention.

[0021] A multi-scale image copy-paste tampering detection method based on FPN extracts features of multiple levels at multiple scales, then performs self-correlation operations on these features in different image sizes at different layers to locate the tampered region of the image, and then connects the features of different sizes and finally performs feature decoding. First, the input image is preprocessed to the same image size of 224*224. Then it is mainly divided into four major modules, as Figure 1 shown. The first module is the FPN feature extraction module. First, image feature extraction is performed, as Figure 2As shown, the present invention selects ResNet50 as the convolutional feature extraction network, and then uses the feature maps generated during the ResNet50 process as the input of the FPN feature pyramid for multi-scale feature extraction to generate feature maps of different sizes. The second module introduces an attention mechanism. By assigning different weights, the attention mechanism can focus on key information while ignoring irrelevant information, and can flexibly adjust the weights to adapt to the needs in different situations, thereby improving the accuracy of the model. The third module is the self-correlation module. The different-scale feature maps output by the FPN enter the self-correlation processing after passing through the attention mechanism to calculate the pixel-level distance vector to obtain the feature similarity score. Through the percentage pooling layer, similar feature matching statistics are performed to locate the similar regions of the image. The fourth module is the mask decoder, which fuses the different-sized feature maps obtained after the self-correlation processing, decodes the image through upsampling to restore it to the original image size, and divides the image region into similar regions and original regions through the activation function to output the Ground Truth mask of the image. Finally, the image is classified as real and fake according to the mask. The specific implementation steps are as follows:

[0022] Step 1: Divide the detection image set into a training set, a validation set, and a test set. Preprocess the images so that the pixel value size is 224*224*3.

[0023] Step 2: Build an image multi-scale copy-paste forgery detection model, including an FPN feature extraction module, an attention mechanism module, a self-correlation module, and a mask decoder module.

[0024] Step 3: Input the detected image into the FPN feature extraction module for feature extraction. FPN mainly includes two processes. The first part is the bottom-up process, which uses CNN for feature extraction. In the present invention, the ResNet50 network is selected as the convolutional feature extraction network. ResNet50 has a deeper network structure and can extract more and more complex features compared with shallow networks. Secondly, the residual blocks in ResNet50 enable the network to have stronger feature representation ability. ResNet50 can learn and extract features at different levels during the training process, and the different-scale features extracted by each feature extraction layer can be perfectly fused into a feature pyramid (FPN). This multi-scale feature extraction method enables the model to have better detection ability and robustness for targets of different sizes and shapes. The second part of FPN is the fusion process of top-down and lateral connections. In the top-down process, FPN uses a pyramidal structure to construct feature maps of different scales. In each layer, the small-size feature map of the previous layer is upsampled to the same size as the current layer and fused with the feature map of the current layer. This fusion method makes full use of the stronger semantic features of the top layer and the high-resolution information of the bottom layer, and can improve the accuracy of classification and localization while maintaining high resolution. On the other hand, the lateral connections in FPN further improve the semantic information of the feature map by directly adding the feature map of the previous layer and the feature map of the current layer. This connection method can better capture the detailed information in the image, enhance the receptive field of the model, and improve the detection ability of the model for small targets.

[0025] Step 4: Process the high-level image features obtained in Step 3 using the attention mechanism. It mainly includes the channel attention mechanism and the spatial attention mechanism. The channel attention mechanism can adaptively adjust the importance of each channel, thereby improving the accuracy and generalization ability of the model, reducing redundant features, and thus reducing the number of parameters of the model and making the model more lightweight. The spatial attention can enable the neural network to more finely focus on the important pixel regions in the image and ignore the unimportant regions, thereby improving the accuracy and generalization ability of the model.

[0026] Step 5: The self-correlation module performs feature matching on the feature maps of different scales marked by the attention mechanism module to find the similar features in the image. The specific approach is to calculate the pairwise similarity scores between any two feature pixels using the pixel distance vector. Specifically, for any two points n and m, their pixel values are (i, j) and (i', j') respectively, and their feature pixels are fx(i, j) and fx(i', j'). The Pearson correlation coefficient ρ is used to calculate the feature similarity equation;

[0027]

[0028] Step 6: The role of the mask decoder is to convert the high-level feature map extracted by the previous layer into a mask image with the same size as the original image. First, it is necessary to perform upsampling fusion on the self-correlation feature images of different scales output by the previous module. Then, the BN-Inception convolutional neural network and BilinearUpPool2D upsampling are alternately applied to restore the feature map to the original resolution.

[0029] Step 7: After obtaining the Ground Truth mask of the image, it is judged whether there are similar regions in the image. For the image with tampered regions, the tampered regions of the image are located.

[0030] The present invention selects the traditional block- and key-point-based methods Two stages and HDBSCAN and the deep learning method Busternet for performance comparison experiments. The experimental results on the Microsoft COCO+USCISI-CMFD joint dataset are shown in Table 1. The binary classification accuracy, the accuracy of real pictures, and the accuracy of fake pictures are respectively selected for comparison. It can be seen that except that the real accuracy of Two stages is higher than that of the method of the present invention. The overall comprehensive performance of the method of the present invention is optimal.

[0031] Table 1 Comparison experimental results of image-level binary classification

[0032]

[0033] To test the generalization ability of the algorithm proposed by the present invention, the key-point algorithm Iterative based on iterative interest points, the algorithm Via hierarchical based on hierarchical features, the algorithm Coherency based on machine learning, and the deep learning method Busternet are selected for performance comparison experiments. Different datasets are used to facilitate the testing of the performance of different methods, and the test is carried out on the CASIA-CMFD dataset. This dataset contains a total of 1313 tampered images. The test results are shown in Table 2.

[0034] It can be seen that on the CASIA-CMFD dataset, both the precision and recall rate of the algorithm of the present invention are higher than those of the other four algorithms, and the performance is better. Compared with the traditional method Iterative, the F1 score of the algorithm proposed by the present invention has increased by 0.0568. Compared with another deep learning-based detection method Busternet, the performance of the algorithm of the present invention has increased by 0.097. And the performance of the algorithm of the present invention is close in terms of precision and recall rate, indicating that the algorithm of the present invention classifies the two types of pixel points more evenly.

[0035] Experimental comparison results on the CASIA-CMFD dataset in Table 2

[0036]

[0037] The above only expresses the preferred embodiments of the present invention, and its description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent for the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several variations, improvements and substitutions can be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent for the present invention shall be subject to the appended claims.

Claims

1. A multi-scale image copy-paste forgery detection method based on FPN, characterized in that: It includes the following steps: Step 1: Divide the detection image set into a training set, a validation set, and a test set, and preprocess the images so that the pixel value size is processed to 224*224*3; Step 2: Build an image multi-scale copy-paste forgery detection model, including an FPN feature extraction module, an attention mechanism module, a self-correlation module, and a mask decoder module; Step 3: Input the detection image into the FPN feature extraction module for feature extraction. FPN includes two processes. The first part is a bottom-up process, which uses the ResNet50 network as the convolutional feature extraction network. ResNet50 learns and extracts features at different levels during training, and the different-scale features extracted by each feature extraction layer are fused into a feature pyramid; The second part of FPN is a fusion process of top-down and lateral connections. In the top-down process, FPN uses a pyramid-shaped structure to build feature maps of different scales. In each layer, the small-size feature map of the previous layer is upsampled to the same size as the current layer and fused with the feature map of the current layer; The lateral connections in FPN improve the semantic information of the feature map by directly adding the feature map of the previous layer and the feature map of the current layer; Step 4: Use the attention mechanism to process the high-level image features obtained in Step 3; It includes a channel attention mechanism and a spatial attention mechanism; The channel attention mechanism adaptively adjusts the importance of each channel, and the spatial attention enables the neural network to more finely focus on the important pixel regions in the image and ignore the unimportant regions; Step 5: The self-correlation module passes the different-scale feature maps output by FPN through the attention mechanism and then enters the self-correlation process to calculate the pixel-level distance vector to obtain the feature similarity score, and performs similar feature matching statistics through the percentage pooling layer to locate the similar regions of the image; Step 6: The role of the mask decoder is to fuse the different-sized feature maps obtained after the self-correlation process, decode the image through upsampling to restore it to the original image size, and divide the image region into similar regions and original regions through the activation function to output the Ground Truth mask of the image; Step 7: After obtaining the Ground Truth mask of the image, determine whether there are similar regions in the image, and for the images with forged regions, locate the image forgery regions.

2. The multi-scale image copy-paste forgery detection method based on FPN according to claim 1, characterized in that: In Step 5, the pairwise similarity score between any two feature pixels is calculated using the pixel distance vector; specifically, for any two points n and m, their pixel values are (i,j) and (i‘,j’) respectively, and their feature pixels are fx(i,j) and fx(i‘,j’) respectively. The Pearson correlation coefficient ρ is used to calculate the feature similarity equation; 3. The multi-scale image copy-paste forgery detection method based on FPN according to claim 1, characterized in that: Step 6 includes: First, perform upsampling and fusion processing on the multi-scale and different-sized self-correlation feature images output by the module, and then alternately apply the BN-Inception convolutional neural network and BilinearUpPool2D upsampling to restore the feature map to the original resolution.