A method for detecting spliced and tampered images
By using a hybrid transformer neural network in image tamper detection, combining self-attention and cross-attention mechanisms, the problem of difficulty in positioning tampering areas at different scales in the prior art is solved, and the detection effect of high accuracy and robustness is achieved.
Patent Information
- Application Number
- CN202211715677.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-29
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-12-29
AI Technical Summary
The prior art is difficult to effectively locate image tampering areas of different scales, especially in case of large-scale tampering, where the positioning is incomplete or the error detection rate is high.
The image tamper detection method based on hybrid transformer neural network is adopted, combining feature extraction module, self-attention U-shaped block, cross-attention module and feature decoding module to integrate self-attention and cross-attention mechanisms to capture global and local features and achieve fine spatial recovery.
Accurate positioning of tampering areas at different scales is achieved, and the accuracy and robustness of detection is improved, especially on the CASIA and COLUMBIA datasets, which perform better than other methods.
Smart Images

Figure BDA0004027633150000101 
Figure FHA0000011264620000011 
Figure FHA0000011264620000031
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer technology, especially in the fields of computer vision and digital image processing technology. It relates to a method for detecting spliced and tampered images, specifically an image tampering detection method based on a hybrid transformer neural network, which can accurately locate the tampered area at different scales and has good localization performance. Background Art
[0002] With the rapid development of modern mobile devices, the generation and transmission of digital images have become very easy. At the same time, image editing software is easy to operate and convenient for anyone to modify images. Some spliced and tampered images may be maliciously misused, causing negative impacts on society and the country. Therefore, detecting spliced and tampered images has become increasingly important.
[0003] Currently, the detection techniques for image tampered areas are mainly divided into two types: detection methods based on traditional feature extraction and detection methods based on convolutional neural networks. Most traditional detection methods extract specific image fingerprints for detection, such as color filter array interpolation, sensor noise, and illumination inconsistency (PRNU). Traditional detection methods can only detect a specific type of image fingerprint, but when the specific fingerprint in the image does not exist or is not obvious, the detection will fail. In addition, specific image fingerprints are affected by post-processing effects such as image blur, JPEG compression, and downsampling, which will reduce the detection results of traditional detection methods. In recent years, many CNN-based tampering detection methods have been proposed and shown better performance than traditional methods. These CNN-based tampering detection methods can extract multiple image fingerprints simultaneously, making up for the defects of traditional methods that rely on a single image attribute, lack generalization ability, and have poor robustness. For example, Faster R-CNN uses RGB stream and noise stream to detect the tampered area of a given tampered image. They use the SRM filter layer to extract noise features from the tampered image, which allows the model to capture the noise inconsistency between the tampered area and the real area. However, the limitation of this method is that it can only achieve regional-level tampering localization results. The RRU-Net structure captures the differential features between the tampered area and the non-tampered area through the residual propagation and residual feedback modules, thus enhancing the learning mode of CNN. However, RRU-Net is not good at extracting global features, resulting in poor localization effect for large-scale tampered areas and a high false detection rate. MWC-Net can learn more comprehensive and representative features. However, this method still does not consider global features, which may lead to poor performance when locating large-scale tampered areas.
[0004] Most previous work has ignored the fact that the size of the tampered area varies and it is difficult to locate tampered areas of different sizes. Due to the inherent locality of convolutional operations, CNN-based methods are difficult to learn explicit global semantic information relationships and are difficult to jointly utilize local and global features. Therefore, most CNN-based detection methods can only handle limited scale changes. In addition, these methods may have problems such as incomplete localization or high false detection rates when locating large-scale tampered areas. Summary of the Invention
[0005] The object of the present invention is to provide an image tampering detection method based on a convolutional neural network in view of the deficiencies of the prior art.
[0006] The method of the present invention is specifically as follows:
[0007] Step (1) Collect and download public image tampering datasets, including the CASIA dataset and the COLUMBIA dataset; take 80-90% of them as the training set and the rest as the test set;
[0008] Step (2) Use the training set to train the model of the hybrid transformer neural network, including: a feature extraction module, a self-attention U-shaped block, a cross-attention module, and a feature decoding module;
[0009] Further, for the feature extraction module, use the encoder part of the U 2 -Net to extract the feature map X of the input image; first, send the input image with a size of 288×288×3 into the residual U-shaped block-7, perform two convolutions with a convolution kernel size of 3×3, a stride of 1, and a padding of 1, followed by a batch normalization data stream layer norm and an activation function Relu; then perform five convolutions with a convolution kernel size of 3×3, a stride of 1, and a padding of 1 plus max-pooling downsampling, followed by a batch normalization data stream layer norm and an activation function Relu; then perform one convolution with a convolution kernel size of 3×3, a stride of 1, a padding of 2, and a dilation rate of 2, followed by a batch normalization data stream layer norm and an activation function Relu; then perform one convolution with a convolution kernel size of 3×3, a stride of 1, and a padding of 1, followed by a batch normalization data stream layer norm and an activation function Relu; then perform five convolutions with a convolution kernel size of 3×3, a stride of 1, and a padding of 1 plus upsampling, followed by a batch normalization data stream layer norm and an activation function Relu; then add the output feature map after the first convolution and the feature map after the last convolution, and finally obtain 64 feature maps. Perform a max-pooling downsampling operation on the feature map output from the residual U-shaped block-7 to reduce the feature map resolution to half of the original; then send the feature map into the residual U-shaped block-6, residual U-shaped block-5, residual U-shaped block-4, and residual U-shaped block-4F in sequence, and finally output a feature map with a size of 9×9×512.
[0010] Furthermore, the self-attention U-shaped block uses self-attention to model the long-range dependencies of the input image and extract the global information of the input image; taking a feature map of size 9×9×512 as the input, two convolutions with a kernel size of 3×3, a stride of 1, and a padding of 1 are performed, followed by a batch normalization data stream layer norm and an activation function Relu; then, a convolution with a kernel size of 3×3, a stride of 1, a padding of 2, and a dilation rate of 2 is performed, followed by a batch normalization data stream layer norm and an activation function Relu; then, a convolution with a kernel size of 3×3, a stride of 1, a padding of 4, and a dilation rate of 4 is performed, followed by a batch normalization data stream layer norm and an activation function Relu; then, a convolution with a kernel size of 3×3, a stride of 1, a padding of 8, and a dilation rate of 8 is performed, followed by a batch normalization data stream layer norm and an activation function Relu;
[0011] Then, the feature map is fed into the self-attention module. A positional encoding is added to the input feature map X ∈ R d×H×W where H and W are the height and width of the feature map respectively, d is the number of channels, and R represents the real number field; then X is flattened and transposed into a sequence of size n×d, where the parameter n = H×W; three 1×1 convolutions are used to project X for query, key, and value embeddings: Q, K, V ∈ R d×n , where Q represents the query matrix, K represents the key matrix, and V represents the value matrix; the output of self-attention is a scaled dot product: A ∈ R n×n is the attention matrix, and the superscript T represents the transpose;
[0012] Then, a convolution with a kernel size of 3×3, a stride of 1, a padding of 4, and a dilation rate of 4 is performed, followed by a batch normalization data stream layer norm and an activation function Relu; then, a convolution with a kernel size of 3×3, a stride of 1, a padding of 2, and a dilation rate of 2 is performed, followed by a batch normalization data stream layer norm and an activation function Relu; then, a convolution with a kernel size of 3×3, a stride of 1, and a padding of 1 is performed, followed by a batch normalization data stream layer norm and an activation function Relu; finally, the output feature map after the first convolution and the feature map after the last convolution are added together, and a feature map of size 9×9×512 is finally output.
[0013] Furthermore, the cross-attention module uses cross-attention to filter out non-semantic features between the feature extraction module and the feature decoding module; first, a positional encoding is added to the low-level feature map U ∈ R d×2H×2W where H and W are the height and width of the feature map respectively, d is the number of channels, and R represents the real number field; then, the low-level feature map U is downsampled once and used as the value matrix of the cross-attention module; in the high-level feature map N ∈ R 2d×H×WAdd position encoding and use a 1×1 convolution to change the number of channels to d, which is then used as the query matrix and key matrix;
[0014] Then use matrix multiplication and the softmax function to calculate the attention matrix A∈R n×n , where n = H×W; rescale the calculated weight values through the Relu activation function to obtain the result S as a filter, where low-magnitude elements represent noise or irrelevant regions to be reduced;
[0015] Perform the Hadamard product of U and S to obtain the version of U that filters out non-semantic features;
[0016] Finally, connect the result of the filtering operation with the high-level feature map N.
[0017] Furthermore, the feature decoding module decodes the feature map with global and local information and predicts the tampered region at the pixel level; upsample the feature map obtained from the self-attention U-shaped block and concatenate it with the output feature map of the cross-attention module, and send the resulting feature map into the residual U-shaped block - 4F; upsample the output feature map of the residual U-shaped block - 4F and concatenate it with the output feature map of the cross-attention module, and send the resulting feature map into the residual U-shaped block - 4; upsample the output feature map of the residual U-shaped block - 4 and concatenate it with the output feature map of the cross-attention module, and send the resulting feature map into the residual U-shaped block - 5; upsample the output feature map of the residual U-shaped block - 5 and concatenate it with the output feature map of the cross-attention module, and send the resulting feature map into the residual U-shaped block - 6; upsample the output feature map of the residual U-shaped block - 6 and concatenate it with the output feature map of the cross-attention module, and send the resulting feature map into the residual U-shaped block - 7; for the feature map with a size of 288×288×64 output by the residual U-shaped block - 7, pass it through a 3×3 convolutional layer to output a feature map with a size of 288×288×1; then pass it through the activation function Sigmiod to obtain a single-channel tampering probability mask map with a size of 288×288×1; the entire neural network optimizes the prediction result by minimizing the cross-entropy loss function through the stochastic gradient descent optimization algorithm, and the cross-entropy loss function: where (r, c) are pixel coordinates, p (r,c) and y (r,c) represent the predicted pixel value of the input image and the pixel value of the ground truth respectively.
[0018] In step (3), use the trained network model to test the test set images to obtain the final detection effect.
[0019] The activation function Relu is defined as f(z i ) = max(0, z i ), z i is the result after the convolution operation. If zi ≤0, f(z i ) = 0, if z i > 0, f(z i ) = z i . The activation function Sigmiod is defined as z i is the result after the convolution operation.
[0020] The present invention integrates self-attention and cross-attention into the U 2 -Net, which can capture more text information and spatial correlations at different scales. The cross-attention module therein can filter out non-semantic features, enhance the low-level feature maps through skip connections under the guidance of high-level semantic information, and achieve fine spatial recovery in the decoder, thereby ultimately improving the correctness of the prediction results. Secondly, the last block of the encoder of the present invention applies self-attention to combine the advantages of the convolution and self-attention mechanisms. Therefore, the present invention can utilize the inductive bias of convolution to avoid large-scale pre-training, as well as the ability of Transformer to learn and model explicit global and long-range semantic information dependencies. In summary, the present invention combines convolution and Transformer together, can locate the splicing forgery regions at various scales, and thus achieves state-of-the-art performance on the two public image forgery datasets of CASIA and COLUMBIA. Specific embodiments
[0021] The technical solution of the present invention will be further described below through specific embodiments.
[0022] A method for detecting spliced forged images is as follows:
[0023] Step (1) Collect and download public image forgery datasets, including the CASIA dataset and the COLUMBIA dataset; take 80 - 90% of them as the training set, and the others as the test set. In this embodiment, 85% is taken as the training set.
[0024] Step (2) Use the training set to train the hybrid transformer neural network, including: a feature extraction module, a self-attention U-shaped block, a cross-attention module, and a feature decoding module;
[0025] The described feature extraction module uses the encoder part of the U 2 -Net to extract the feature map X of the input image; specifically:
[0026] An image with a size of 288×288×3 is fed into the Residual U-shaped block-7, and two convolutions with a kernel size of 3×3, a stride of 1, and a padding of 1 are performed, followed by a batch normalization data stream layer norm and an activation function Relu. Then, five convolutions with a kernel size of 3×3, a stride of 1, and a padding of 1 are performed, along with max-pooling downsampling, followed by a batch normalization data stream layer norm and an activation function Relu. Next, one convolution with a kernel size of 3×3, a stride of 1, a padding of 2, and a dilation rate of 2 is performed, followed by a batch normalization data stream layer norm and an activation function Relu. Then, one convolution with a kernel size of 3×3, a stride of 1, and a padding of 1 is performed, followed by a batch normalization data stream layer norm and an activation function Relu. Next, five convolutions with a kernel size of 3×3, a stride of 1, and a padding of 1 are performed, along with upsampling, followed by a batch normalization data stream layer norm and an activation function Relu. The output feature map after the first convolution and the feature map after the last convolution are added together to obtain a feature map with an output size of 288×288×64. After passing through the max-pooling data stream layer pool, the final output is a feature map with a size of 144×144×64;
[0027] The feature map with a size of 144×144×64 is continuously fed into the Residual U-shaped block-6, and two convolutions with a kernel size of 3×3, a stride of 1, and a padding of 1 are performed, followed by a batch normalization data stream layer norm and an activation function Relu. Then, four convolutions with a kernel size of 3×3, a stride of 1, and a padding of 1 are performed, along with max-pooling downsampling, followed by a batch normalization data stream layer norm and an activation function Relu. Next, one convolution with a kernel size of 3×3, a stride of 1, a padding of 2, and a dilation rate of 2 is performed, followed by a batch normalization data stream layer norm and an activation function Relu. Then, one convolution with a kernel size of 3×3, a stride of 1, and a padding of 1 is performed, followed by a batch normalization data stream layer norm and an activation function Relu. Next, four convolutions with a kernel size of 3×3, a stride of 1, and a padding of 1 are performed, along with upsampling, followed by a batch normalization data stream layer norm and an activation function Relu. The output feature map after the first convolution and the feature map after the last convolution are added together to obtain a feature map with an output size of 144×144×128. After passing through the max-pooling data stream layer pool, the final output is a feature map with a size of 72×72×128;
[0028] Continue to feed the feature map with a size of 72×72×128 into the residual U-shaped block - 5, perform two convolutions with a convolution kernel size of 3×3, a stride of 1, and a padding of 1, followed by batch normalization layer norm and activation function Relu; then perform three convolutions with a convolution kernel size of 3×3, a stride of 1, and a padding of 1, followed by max-pooling downsampling, batch normalization layer norm, and activation function Relu; then perform one convolution with a convolution kernel size of 3×3, a stride of 1, a padding of 2, and a dilation rate of 2, followed by batch normalization layer norm and activation function Relu; then perform one convolution with a convolution kernel size of 3×3, a stride of 1, and a padding of 1, followed by batch normalization layer norm and activation function Relu; then perform three convolutions with a convolution kernel size of 3×3, a stride of 1, and a padding of 1, followed by upsampling, batch normalization layer norm, and activation function Relu; add the output feature map after the first convolution and the feature map after the last convolution to obtain a feature map with an output size of 72×72×256, and then pass it through the max-pooling data stream layer pool to finally output a feature map with a size of 36×36×256;
[0029] Continue to feed the feature map with a size of 36×36×256 into the residual U-shaped block - 4, perform two convolutions with a convolution kernel size of 3×3, a stride of 1, and a padding of 1, followed by batch normalization layer norm and activation function Relu; then perform two convolutions with a convolution kernel size of 3×3, a stride of 1, and a padding of 1, followed by max-pooling downsampling, batch normalization layer norm, and activation function Relu; then perform one convolution with a convolution kernel size of 3×3, a stride of 1, a padding of 2, and a dilation rate of 2, followed by batch normalization layer norm and activation function Relu; then perform one convolution with a convolution kernel size of 3×3, a stride of 1, and a padding of 1, followed by batch normalization layer norm and activation function Relu; then perform two convolutions with a convolution kernel size of 3×3, a stride of 1, and a padding of 1, followed by upsampling, batch normalization layer norm, and activation function Relu; add the output feature map after the first convolution and the feature map after the last convolution to obtain a feature map with an output size of 36×36×512, and then pass it through the max-pooling data stream layer pool to finally output a feature map with a size of 18×18×512;
[0030] The feature map with a size of 36×36×256 is continuously fed into the residual U-shaped block - 4F, and two convolutions with a kernel size of 3×3, a stride of 1, and a padding of 1 are performed, followed by a batch normalization data stream layer norm and an activation function Relu; then a convolution with a kernel size of 3×3, a stride of 1, a padding of 2, and a dilation rate of 2 is performed, followed by a batch normalization data stream layer norm and an activation function Relu; then another convolution with a kernel size of 3×3, a stride of 1, a padding of 4, and a dilation rate of 4 is performed, followed by a batch normalization data stream layer norm and an activation function Relu; then another convolution with a kernel size of 3×3, a stride of 1, a padding of 8, and a dilation rate of 8 is performed, followed by a batch normalization data stream layer norm and an activation function Relu; then another convolution with a kernel size of 3×3, a stride of 1, a padding of 4, and a dilation rate of 4 is performed, followed by a batch normalization data stream layer norm and an activation function Relu; then another convolution with a kernel size of 3×3, a stride of 1, a padding of 2, and a dilation rate of 2 is performed, followed by a batch normalization data stream layer norm and an activation function Relu; then another convolution with a kernel size of 3×3, a stride of 1, and a padding of 1 is performed, followed by a batch normalization data stream layer norm and an activation function Relu; the output feature map after the first convolution and the feature map after the last convolution are added together to obtain a feature map with an output size of 18×18×512, and then through a max pooling data stream layer pool, the final output feature map with a size of 9×9×512 is obtained.
[0031] The batch normalization data stream layer norm, for each x in the data batch B = {x1,…,x i ,…,x n}, is transformed into y i The calculation update is defined as: i where the hyperparameters {γ,β} shuffle the data batch to prevent training drift. μ and B and respectively represent the mean and variance in the batch B, and ε is a small positive number used to avoid division by zero.
[0032] The self-attention U-shaped block uses self-attention to model the long-range dependencies of the input image and extract the global information of the input image; specifically:
[0033] Take the feature map with a size of 9×9×512 as the input and perform two convolutions with a kernel size of 3×3, a stride of 1, and a padding of 1, followed by a batch normalization data stream layer norm and an activation function Relu; then perform another convolution with a kernel size of 3×3, a stride of 1, a padding of 2, and a dilation rate of 2, followed by a batch normalization data stream layer norm and an activation function Relu; then perform another convolution with a kernel size of 3×3, a stride of 1, a padding of 4, and a dilation rate of 4, followed by a batch normalization data stream layer norm and an activation function Relu; then perform another convolution with a kernel size of 3×3, a stride of 1, a padding of 8, and a dilation rate of 8, followed by a batch normalization data stream layer norm and an activation function Relu;
[0034] Then send the feature map into the self-attention module. To consider the absolute semantic information, first add positional encoding to the input feature map X∈R d×H×W where H and W are the height and width of the feature map respectively, d is the number of channels, and R represents the real number field; positional encoding is particularly suitable for capturing the absolute and relative positions between tampered regions in self-attention; then flatten and transpose X into a sequence of size n×d, where the parameter n = H×W; use three 1×1 convolutions to project X for query, key, and value embeddings: Q, K, V∈R d×n where Q represents the query matrix, K represents the key matrix, and V represents the value matrix; the output of self-attention is a scaled dot product: A∈R n×n is the attention matrix, and the superscript T represents transpose; the attention matrix A is used as a weight to consider all interactions between the query and the key;
[0035] Perform another convolution with a kernel size of 3×3, a stride of 1, a padding of 4, and a dilation rate of 4, followed by a batch normalization data stream layer norm and an activation function Relu; then perform another convolution with a kernel size of 3×3, a stride of 1, a padding of 2, and a dilation rate of 2, followed by a batch normalization data stream layer norm and an activation function Relu; then perform another convolution with a kernel size of 3×3, a stride of 1, and a padding of 1, followed by a batch normalization data stream layer norm and an activation function Relu; then add the output feature map after the first convolution and the feature map after the last convolution, and finally output a feature map with a size of 9×9×512.
[0036] The described cross-attention module uses cross-attention to filter out non-semantic features between the feature extraction module and the feature decoding module, achieve more refined spatial recovery in the feature decoding module, and improve the correctness of the prediction results; specifically:
[0037] First, in the low-level feature map U∈R d×2H×2WAdd position encoding, where H and W are the height and width of the feature map respectively, d is the number of channels, and R represents the real number field; then, downsample the low-level feature map U once and use it as the value matrix of the cross-attention module; in the high-level feature map N ∈ R 2d×H×W After adding position encoding and changing the number of channels to d using a 1×1 convolution, use it as the query matrix and key matrix;
[0038] Then use matrix multiplication and the softmax function to calculate the attention matrix A ∈ R n×n , where n = H × W; rescale the calculated weight values through the Relu activation function to obtain the result S as a filter, where low-amplitude elements represent noise or irrelevant regions to be reduced;
[0039] Perform the Hadamard product of U and S to obtain the version of U that filters out non-semantic features;
[0040] Finally, connect the result of the filtering operation with the high-level feature map N. Through the cross-attention module, more detailed information than ordinary skip connections can be retained, thereby improving the detection performance.
[0041] The described feature decoding module decodes the feature map with global information and local information and predicts the tampered region at the pixel level. The feature decoding module has a structure similar to its symmetric feature extraction module; specifically:
[0042] The feature map with a size of 9×9×512 obtained from the self-attention U-shaped block is upsampled and concatenated with the output feature map passing through the cross-attention module to obtain a feature map of 18×18×1024. Then, it passes through the residual U-shaped block - 4F, and two convolutions with a kernel size of 3×3, a stride of 1, and a padding of 1 are performed, followed by batch normalization data stream layer norm and activation function Relu; then a convolution with a kernel size of 3×3, a stride of 1, a padding of 2, and a dilation rate of 2 is performed, followed by batch normalization data stream layer norm and activation function Relu; then a convolution with a kernel size of 3×3, a stride of 1, a padding of 4, and a dilation rate of 4 is performed, followed by batch normalization data stream layer norm and activation function Relu; then a convolution with a kernel size of 3×3, a stride of 1, a padding of 8, and a dilation rate of 8 is performed, followed by batch normalization data stream layer norm and activation function Relu; then a convolution with a kernel size of 3×3, a stride of 1, a padding of 4, and a dilation rate of 4 is performed, followed by batch normalization data stream layer norm and activation function Relu; then a convolution with a kernel size of 3×3, a stride of 1, a padding of 2, and a dilation rate of 2 is performed, followed by batch normalization data stream layer norm and activation function Relu; then a convolution with a kernel size of 3×3, a stride of 1, and a padding of 1 is performed, followed by batch normalization data stream layer norm and activation function Relu; then the output feature map after the first convolution and the feature map after the last convolution are added together, and finally a feature map with a size of 18×18×512 is output;
[0043] The feature map with a size of 18×18×512 output from the residual U-shaped block - 4F is upsampled and concatenated with the output feature map passing through the cross-attention module to obtain a feature map of 36×36×1024. Then, it passes through the residual U-shaped block - 4, and two convolutions with a kernel size of 3×3, a stride of 1, and a padding of 1 are performed, followed by batch normalization data stream layer norm and activation function Relu; then two convolutions with a kernel size of 3×3, a stride of 1, and a padding of 1 are performed along with max-pooling downsampling, followed by batch normalization data stream layer norm and activation function Relu; then a convolution with a kernel size of 3×3, a stride of 1, a padding of 2, and a dilation rate of 2 is performed, followed by batch normalization data stream layer norm and activation function Relu; then a convolution with a kernel size of 3×3, a stride of 1, and a padding of 1 is performed, followed by batch normalization data stream layer norm and activation function Relu; then two convolutions with a kernel size of 3×3, a stride of 1, and a padding of 1 are performed along with upsampling, followed by batch normalization data stream layer norm and activation function Relu; then the output feature map after the first convolution and the feature map after the last convolution are added together, and finally a feature map with a size of 36×36×256 is output;
[0044] The feature map with a size of 36×36×256 output by the residual U-shaped block - 4 is upsampled and concatenated with the output feature map passing through the cross-attention module to obtain a feature map of 72×72×512. Then, it passes through the residual U-shaped block - 5, and two convolutions with a kernel size of 3×3, a stride of 1, and a padding of 1 are performed, followed by the batch normalization data stream layer norm and the activation function Relu; then, three convolutions with a kernel size of 3×3, a stride of 1, and a padding of 1 are performed, followed by max-pooling downsampling, the batch normalization data stream layer norm, and the activation function Relu; then, one convolution with a kernel size of 3×3, a stride of 1, a padding of 2, and a dilation rate of 2 is performed, followed by the batch normalization data stream layer norm and the activation function Relu; then, one convolution with a kernel size of 3×3, a stride of 1, and a padding of 1 is performed, followed by the batch normalization data stream layer norm and the activation function Relu; then, three convolutions with a kernel size of 3×3, a stride of 1, and a padding of 1 are performed, followed by upsampling, the batch normalization data stream layer norm, and the activation function Relu; then, the output feature map after the first convolution and the feature map after the last convolution are added together, and finally, a feature map with a size of 72×72×128 is output;
[0045] The feature map with a size of 72×72×128 output by the residual U-shaped block - 5 is upsampled and concatenated with the output feature map passing through the cross-attention module to obtain a feature map of 144×144×256. Then, it passes through the residual U-shaped block - 6, and two convolutions with a kernel size of 3×3, a stride of 1, and a padding of 1 are performed, followed by the batch normalization data stream layer norm and the activation function Relu; then, four convolutions with a kernel size of 3×3, a stride of 1, and a padding of 1 are performed, followed by max-pooling downsampling, the batch normalization data stream layer norm, and the activation function Relu; then, one convolution with a kernel size of 3×3, a stride of 1, a padding of 2, and a dilation rate of 2 is performed, followed by the batch normalization data stream layer norm and the activation function Relu; then, one convolution with a kernel size of 3×3, a stride of 1, and a padding of 1 is performed, followed by the batch normalization data stream layer norm and the activation function Relu; then, four convolutions with a kernel size of 3×3, a stride of 1, and a padding of 1 are performed, followed by upsampling, the batch normalization data stream layer norm, and the activation function Relu; then, the output feature map after the first convolution and the feature map after the last convolution are added together, and finally, a feature map with a size of 144×144×64 is output;
[0046] The feature map with a size of 144×144×64 output by the residual U-shaped block - 6 is upsampled and concatenated with the output feature map passing through the cross-attention module to obtain a feature map of 288×288×128. Then, it passes through the residual U-shaped block - 7 and undergoes two convolutions with a kernel size of 3×3, a stride of 1, and a padding of 1, followed by a batch normalization data stream layer norm and an activation function Relu. Then, it undergoes five convolutions with a kernel size of 3×3, a stride of 1, and a padding of 1 plus a max-pooling downsampling, followed by a batch normalization data stream layer norm and an activation function Relu. Then, it undergoes one convolution with a kernel size of 3×3, a stride of 1, a padding of 2, and a dilation rate of 2, followed by a batch normalization data stream layer norm and an activation function Relu. Then, it undergoes one convolution with a kernel size of 3×3, a stride of 1, and a padding of 1, followed by a batch normalization data stream layer norm and an activation function Relu. Then, it undergoes five convolutions with a kernel size of 3×3, a stride of 1, and a padding of 1 plus an upsampling, followed by a batch normalization data stream layer norm and an activation function Relu. Then, the output feature map after the first convolution and the feature map after the last convolution are added together, and finally, a feature map with a size of 288×288×64 is output;
[0047] The feature map with a size of 288×288×64 output by the residual U-shaped block - 7 passes through a 3×3 convolutional layer to output a feature map with a size of 288×288×1. Then, it passes through the Sigmiod activation function to obtain a single-channel tampering probability mask map with a size of 288×288×1. The entire neural network optimizes the prediction result by minimizing the cross-entropy loss function through the stochastic gradient descent optimization algorithm. The cross-entropy loss function: where (r, c) are pixel coordinates, p (r,c) and y (r,c) represent the predicted pixel value of the input image and the calibrated true value pixel value respectively.
[0048] In step (3), the trained network model is used to test the test set images to obtain the final detection effect, and its accuracy can reach about 81%.
[0049] The test uses an I7 9700K CPU, 32GB of memory, a GTX 3060 GPU, Python version 3.8, and PyTorch version 1.8.1. During the test, the learning rate is 0.0001, the weight decay is set to 0, and after 150 iterations of training, the final trained model is obtained. We use 85% of the CASIA dataset and the COLUMBIA dataset as the training set to train the hybrid transformer-based neural network, and 15% of them as the test set to test the detection accuracy of the hybrid transformer neural network. The method of the present invention is compared with the existing methods ADQ, CFA, ELA, NOI, RRU-Net, U-Net and ManTra-Net, U 2 -Net for comparative analysis. The evaluation metrics used in the comparative analysis include: Precision is the precision rate, Recall is the recall rate, and F1 is the balance of the precision rate and the recall rate. All evaluation metrics are better when they are higher and closer to 1. The following table shows the comparison results of the tampering area localization of different methods on different datasets.
[0050]
[0051] It can be seen that the highest precision rate and F1 value are achieved on both the CASIA and COLUMBIA datasets. Further subjective comparison shows that the localization area of the present invention is more complete and the localization edge is more accurate, indicating the effectiveness of the present invention in dealing with splicing tampering types.
[0052] On the CASIA dataset, as the variance of Gaussian noise increases, the precision rate and F1 value of the present invention are far better than the other six detection methods. The recall rate of the present invention is slightly lower than that of RRU-Net, and the detection index of the present invention is the least affected. Experiments prove that the present invention is robust to the noise attack on the dataset. In addition, as the JPEG quality factor decreases from 100 to 50, the precision rate and F1 value of the present invention are still better than other methods, and it is robust to the compression attack on the dataset.
[0053] The above embodiments should be understood as only for illustrating the present invention and not for limiting the protection scope of the present invention. Any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present invention without departing from the content of the technical solution of the present invention shall still fall within the scope of the technical solution of the present invention.
Claims
1. A method for detecting spliced and tampered images, characterized in that: Step (1) Collect and download public image tampering datasets, including the CASIA dataset and the COLUMBIA dataset; take 80-90% of them as the training set, and the rest as the test set; Step (2) Use the training set to train the model of the hybrid transformer neural network, including: a feature extraction module, a self-attention U-shaped module, a cross-attention module, and a feature decoding module; The described feature extraction module uses the encoder part of U 2 -Net to extract the feature map X of the input image; The self-attention U-shaped module uses self-attention to model the long-range dependence of the input image and extract the global information of the input image; The cross-attention module uses cross-attention to filter out non-semantic features between the feature extraction module and the feature decoding module; The feature decoding module decodes the feature map with global information and local information, and predicts the tampered area at the pixel level; Step (3) Use the trained network model to test the test set images to obtain the final detection effect; The specific self-attention U-shaped module is: taking a feature map with a size of 9×9×512 as the input, performing two convolutions with a convolution kernel size of 3×3, a stride of 1, and a padding of 1, a batch normalization data stream layer norm, and an activation function Relu; then performing a convolution with a convolution kernel size of 3×3, a stride of 1, a padding of 2, and a dilation rate of 2, a batch normalization data stream layer norm, and an activation function Relu; then passing through a convolution with a convolution kernel size of 3×3, a stride of 1, a padding of 4, and a dilation rate of 4, a batch normalization data stream layer norm, and an activation function Relu; then performing a convolution with a convolution kernel size of 3×3, a stride of 1, a padding of 8, and a dilation rate of 8, a batch normalization data stream layer norm, and an activation function Relu; Then the feature map is fed into the self-attention module, and positional encoding is added to the input feature map X ∈ R d×H×W where H and W are the height and width of the feature map respectively, d is the number of channels, and R represents the real number field; then X is flattened and transposed into a sequence of size n × d, with the parameter n = H × W; three 1×1 convolutions are used to project X for query, key, and value embeddings: Q, K, V ∈ ℝ d×n , where Q represents the query matrix, K represents the key matrix, and V represents the value matrix; the output of self-attention is a scaled dot product: A ∈ ℝ n×n is the attention matrix, and the superscript T represents the transpose; Then perform a convolution with a convolution kernel size of 3×3, a stride of 1, a padding of 4, and a dilation rate of 4, a batch normalization data stream layer norm, and an activation function Relu; then perform a convolution with a convolution kernel size of 3×3, a stride of 1, a padding of 2, and a dilation rate of 2, a batch normalization data stream layer norm, and an activation function Relu; then perform a convolution with a convolution kernel size of 3×3, a stride of 1, and a padding of 1 convolution, a batch normalization data stream layer norm, and an activation function Relu; then add the output feature map after the first convolution and the feature map after the last convolution, and finally output a feature map with a size of 9×9×512.
2. The detection method for spliced and tampered images according to claim 1, characterized in that: The described feature extraction module first sends the input image with a size of 288×288×3 into the Residual U-shaped module-7, performs two convolutions with a kernel size of 3×3, a stride of 1, and a padding of 1, followed by a batch normalization data stream layer norm and an activation function Relu; then performs five convolutions with a kernel size of 3×3, a stride of 1, and a padding of 1 plus max-pooling downsampling, followed by a batch normalization data stream layer norm and an activation function Relu; then performs one convolution with a kernel size of 3×3, a stride of 1, a padding of 2, and a dilation rate of 2, followed by a batch normalization data stream layer norm and an activation function Relu; then performs one convolution with a kernel size of 3×3, a stride of 1, and a padding of 1, followed by a batch normalization data stream layer norm and an activation function Relu; then performs five convolutions with a kernel size of 3×3, a stride of 1, and a padding of 1 plus upsampling, followed by a batch normalization data stream layer norm and an activation function Relu; then adds the output feature map after the first convolution and the feature map after the last convolution, and finally obtains 64 feature maps. The feature map output from the Residual U-shaped module-7 undergoes a max-pooling downsampling operation to reduce the feature map resolution to half of the original; Then the feature map is successively sent into the Residual U-shaped module-6, the Residual U-shaped module-5, the Residual U-shaped module-4, and the Residual U-shaped module-4F, and finally outputs a feature map with a size of 9×9×512.
3. The detection method for spliced and tampered images according to claim 1, characterized in that, The specific cross-attention module is as follows: First, a positional encoding is added to the low-level feature map U ∈ R d×2H×2W . Here, H and W are the height and width of the feature map respectively, d is the number of channels, and R represents the real number field. Then, the low-level feature map U is downsampled once and used as the value matrix of the cross-attention module. For the high-level feature map N ∈ R 2d×H×W , a positional encoding is added, and after changing the number of channels to d using a 1×1 convolution, it is used as the query matrix and the key matrix. Then, the attention matrix A ∈ R is calculated using matrix multiplication and the softmax function n×n , where n = H × W; the calculated weight values are rescaled through the Relu activation function to obtain the result S as a filter, where the low-magnitude elements represent the noise or irrelevant regions to be reduced; Perform the Hadamard product on U and S to obtain a version of U that filters out non-semantic features; Finally, connect the result of the filtering operation with the high-level feature map N.
4. The detection method of spliced and tampered images according to claim 1, characterized in that, The described feature decoding module upsamples the feature map obtained from the self-attention U-shaped module and concatenates it with the output feature map passing through the cross-attention module. The resulting feature map is fed into the residual U-shaped module - 4F; the upsampled output feature map of the residual U-shaped module - 4F and the output feature map passing through the cross-attention module are concatenated, and the resulting feature map is fed into the residual U-shaped module - 4; the upsampled output feature map of the residual U-shaped module - 4 and the output feature map passing through the cross-attention module are concatenated, and the resulting feature map is fed into the residual U-shaped module - 5; the upsampled output feature map of the residual U-shaped module - 5 and the output feature map passing through the cross-attention module are concatenated, and the resulting feature map is fed into the residual U-shaped module - 6; the upsampled output feature map of the residual U-shaped module - 6 and the output feature map passing through the cross-attention module are concatenated, and the resulting feature map is fed into the residual U-shaped module - 7; the feature map with a size of 288×288×64 output from the residual U-shaped module - 7 passes through a 3×3 convolutional layer to output a feature map with a size of 288×288×1; and then passes through the Sigmiod activation function to obtain a single-channel tampering probability mask map with a size of 288×288×1; the entire neural network optimizes the prediction result by minimizing the cross-entropy loss function through the stochastic gradient descent optimization algorithm. The cross-entropy loss function: where (r, c) are pixel coordinates, p (r,c) and y (r,c) represent the predicted pixel value and the calibrated ground truth pixel value of the input image, respectively.