Digital Image Forgery Region Localization Method Based on Dual-Stream Deep Neural Network

By designing a dual-stream deep neural network in the digital image forgery area positioning method, combining the original image sampling unit and the Qualcomm image sampling subnet, effective positioning and recognition of various image forgery types is achieved, and the lack of feature extraction and information learning in the prior art is solved, and the accuracy and universality of positioning are improved.

CN115170933BActive Publication Date: 2025-07-01XIDIAN UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202210993852.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-18
Publication Date
2025-07-01
Estimated Expiration
2042-08-18

AI Technical Summary

Technical Problem

The prior art is difficult to effectively locate and identify other image forgery types except for stitching types, such as copy-movement and deletion, and there are shortcomings in the extraction of key feature of image forgery and the learning of feature information.

Method used

A digital image forgery area positioning method based on dual-stream deep neural network is designed. By introducing the original image sampling unit and the Qualcomm image sampling subnet in the encoder, and using convolutional operations, pooling operations and nonlinear operations in the feature fusion module, the effective fusion and learning of feature information is achieved.

Benefits of technology

This method can effectively learn the outline of the image forgery area, improve the accuracy and universality of the digital image forgery area positioning, and overcome the shortcomings of the existing methods in feature extraction and information learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115170933B_ABST
    Figure CN115170933B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for locating forged regions of digital images based on a two-stream deep neural network, mainly solving the problems of low universality of existing image forgery detection methods, lack of extraction of key features of image forgery, and omission of important feature information of forged images. The steps implemented by the present invention are as follows: constructing a two-stream deep neural network including an encoder, a feature fusion module, and a decoder; the encoder is implemented by two sub-feature extractors to extract features of the two streams; using a training set composed of various types of image forgery samples to train the two-stream deep neural network to locate the forged regions of digital images. Since the present invention uses double residual blocks, single residual blocks, and a feature fusion module to extract and learn the forgery features of images, it realizes the ability to locate the forged regions of digital images and improves the accuracy and efficiency of locating the forged regions of images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of digital information security technology, and is a method for locating forged regions of digital images based on a two-stream deep neural network. The present invention can be used to detect and judge the security of digital image content in the Internet environment, locate the forged regions of Internet digital images, and provide important basis for copyright protection, infringement tracing, and false news identification. Background Art

[0002] Image forgery, also known as tampering, is an attack mode on image content, usually including global tampering attacks and local tampering attacks. Global tampering is to perform blurring operations, flipping, adding noise, etc. on the image, which only affects the visual effect of the image without changing the context information of the image. Local tampering attacks can be divided into traditional forgery attacks and deep learning-based forgery attacks. Traditional types of image forgery attacks include splicing, copy-pasting, and deletion forgery. Deep learning-based forgery attacks generally refer to replacing specific targets in the image through a generative adversarial model, such as using artificial intelligence technology to change faces and using a generative adversarial model to retouch pictures. At present, the accuracy of locating forged regions of digital images is low, the universality is weak, and the ability to resist complex attacks is weak.

[0003] Hebei University of Technology proposed a method for detecting and locating spliced forged images using a two-channel convolutional neural network in its patented technology "A Detection Method for Spliced and Tampered Images" (application number: 201911325073.7, authorization announcement number: CN111062931B). This method uses two convolutional neural networks with the same structure to extract multi-stage features of the spliced and tampered image and its corresponding light source map respectively, combines multi-scale information, fuses and upsamples the two sets of multi-stage features to obtain a pyramid feature map, passes different layers of the pyramid feature map through a region generation network respectively to obtain tampering candidate regions, generates a fixed-size feature map through ROIAlign, classifies, regresses the bounding box, and predicts the mask for the fixed-size feature map, and finally obtains the bounding box and pixel-level localization of the tampering region, completing the detection of the spliced and tampered image. Although this method overcomes the defects of the existing technology that the extracted image tampering features are single and incomplete, it is easy to ignore small tampering targets, and it cannot achieve end-to-end pixel-level localization. However, the deficiencies of this method still exist: this method can only be used to locate spliced forged images, and this method cannot handle other forgery types including splicing, copy-moving, deletion, etc. known in current image forgery types.

[0004] Chongqing University of Posts and Telecommunications proposed a method for detecting image tampering based on a two-stream convolutional neural network in its patent document "A Method for Detecting Image Tampering Based on a Two-stream Convolutional Neural Network" (Application No.: 202110266702.4, Publication No.: CN 113129261 A). The implementation steps of this method are as follows: 1) Collect and organize publicly available tampered image samples; 2) Label the tampered image samples to obtain the labels of these tampered image samples to complete the construction of the image tampering dataset; 3) Use the collected image tampering dataset to train the two-stream convolutional neural network; 4) Use the trained model to test other tampered images to obtain the final effect. Although the model trained by this method using the two-stream convolutional neural network can detect tampered images in reality. However, the deficiencies of this method are still: only two convolutional neural networks are used to extract the tampering features of the image, and the learning of the edges of the forged areas of the image is lacking, resulting in the problem of missing extraction of key features of image forgery.

[0005] Zhou et al. designed a method for locating digital image forgery of a two-stream neural network in their published paper "Learning Rich Features for Image Manipulation Detection" (2018 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR) IEEE). The important contribution of this method is to design an SRM filter for feature preprocessing of digital images, effectively improving the ability of the model to locate image forgery. The deficiencies of this method are: when fusing features of the two-stream subnets, the traditional bilinear pooling method is adopted, which cannot be deployed for operation on the CPU, thus increasing the training time of the model. In addition, bilinear pooling uses methods such as principal component analysis for dimensionality reduction, so some key feature information extracted by the neural network will be lost. Summary of the Invention

[0006] The purpose of the present invention is to address the above deficiencies of the prior art and propose a method for locating the forged areas of digital images based on a two-channel deep neural network. It is used to solve the problems of being unable to handle other forgery types including splicing, copy-move, deletion, etc. among the currently known image forgery types, as well as the problems of missing extraction of key features of image forgery and omission of important feature information of forged images.

[0007] The technical idea for achieving the object of the present invention is that the present invention designs a two-stream deep neural network. The network includes an encoder, a feature fusion module, and a decoder. The encoder part includes two sub-networks, an original image sampling unit and a high-pass image sampling sub-network. The original image sampling unit includes four trainable iterative layers, which are specifically used to extract the essential attributes of the tampered image; the high-pass image sampling sub-network includes a high-pass filtering processing layer and four trainable iterative layers. Thus, the high-pass image sampling sub-network can learn the forged edge attributes of the image and effectively solve the problem of missing extraction of key features of image forgery. The feature fusion module of the present invention forms a trainable iterative layer through convolution operation, pooling operation, and non-linear operation, and performs information fusion on the feature maps extracted by the two sub-networks, effectively solving the problem of key information loss that occurs during feature fusion of the sub-networks. The decoder part of the present invention includes four trainable iterative layers, which are connected to each iterative layer in the original image sampling unit by short-circuit connections, stabilizing the image forgery feature information extracted at the current stage. Since the two-stream deep neural network proposed by the present invention is trained as a whole using a large-scale image forgery dataset, it is ensured that the model is fitted and a ".pkl" file is obtained. The trained model is used to locate three types of forged images, namely splicing, copy-move, and deletion, solving the problem that some image forgery localization methods can only process a specific type of forged image.

[0008] The implementation steps of the present invention are as follows:

[0009] Step 1, construct the encoder in the two-stream deep neural network:

[0010] Step 1.1, build an original image sampling unit composed of four serially connected double residual blocks with the same structure;

[0011] Step 1.2, the double residual block is composed of an input layer, a first convolutional layer, a first non-linear operation layer, a second convolutional layer, a first residual convolutional layer, a non-linear addition operation unit, a second non-linear operation layer, a third convolutional layer, a fourth convolutional layer, a gated convolutional layer, a gated non-linear operation layer, a third non-linear operation layer, a second residual convolutional layer, and an output layer;

[0012] Step 1.3, the first convolutional layer, the first non-linear operation layer, the second convolutional layer, the non-linear addition operation unit, and the first residual convolutional layer are sequentially connected in series to form a residual block for the first residual propagation and the output result of the residual block for the first residual propagation;

[0013] Step 1.4, the third convolutional layer, the second non-linear operation layer, the fourth convolutional layer, the third non-linear operation layer, and the second residual convolutional layer are sequentially connected in series to form a residual block for the second residual propagation;

[0014] Step 1.5, the input layer is connected to the first convolutional layer, the first residual convolutional layer, and the second residual convolutional layer respectively, and then is serially connected to the residual block of the first residual propagation, the gated convolutional layer, the gated non-linear operation layer, the residual block of the second residual propagation, and the output layer in sequence to form a double residual block as a whole;

[0015] Step 1.6, set the channel parameter of the input layer to 3 according to the size of the image resolution, set the number of input neurons to 384×256, set the input channels of the first and third convolutional layers, the first and second residual convolutional layers to 3, the output channels to 32, the kernel size to 3×3, the sliding step to 1, and the dilation coefficient to 1. Set the input and output channel numbers of the second and fourth convolutional layers to 32, the kernel size to 3×3, the sliding step to 1, and the dilation coefficient to 1. Set the input channel number of the gated convolutional layer to 32, the output channel size to 3, the kernel size to 3×3, the sliding step to 1, and the dilation coefficient to 1. The first to third non-linear operation layers and the non-linear addition operation unit are all implemented using the ReLU activation function, and the input and output channel numbers of the output layer are both set to 32;

[0016] Step 1.7, build a high-pass image sampling sub-network composed of a high-pass filter module and four single residual blocks with the same structure connected in series in sequence;

[0017] Step 1.8, the high-pass filter module consists of an input layer, a horizontal high-pass kernel, a vertical high-pass kernel, a superposition operation unit, and a high-pass filter output layer;

[0018] Step 1.9, after connecting the horizontal high-pass kernel, the vertical high-pass kernel, and the superposition operation unit in series, connect them to the input layer respectively to form a high-pass filter, and the high-pass filter is connected in series with the high-pass filter output layer to form a high-pass filter module;

[0019] Step 1.10, set the channel parameter of the high-pass filter module to 3 according to the size of the input sample image resolution, set the number of input neurons to 384×256, and set the convolution kernel size to 3×3;

[0020] Step 1.11, the single residual block includes an input layer, a residual block of single residual propagation, and an output layer. Among them, the residual block of single residual propagation is formed by the first convolutional layer, the first non-linear operation layer, the second convolutional layer, the second non-linear operation layer, and the first residual convolutional layer connected in series. The input layer is connected to the first convolutional layer and the first residual convolutional layer respectively, and then is connected in series with the output layer to form a single residual block;

[0021] Step 1.12: Set the number of input channels of the first convolutional layer and the first residual convolutional layer to 3, the number of output channels to 32, the convolutional kernel size to 3×3, the sliding step to 1, and the dilation coefficient to 1. Set the number of input and output channels of the second convolutional layer to 32, the convolutional kernel size to 3×3, the sliding step to 1, and the dilation coefficient to 1. Implement the first and second non-linear operation layers using the ReLU activation function, and set the number of input and output channels of the output layer to 32;

[0022] Step 1.13: Connect the original image sampling unit in series with the high-pass image sampling sub-network to form an encoder;

[0023] Step 2: Construct the feature fusion module of the two-stream deep neural network:

[0024] Step 2.1: The feature fusion module consists of a feature concatenation layer, a first convolutional layer, a first non-linear operation layer, and a first pooling layer connected in series in sequence;

[0025] Step 2.2: According to the results output by the original image sampling unit and the high-pass image sampling sub-network, set the number of input channels of the feature concatenation layer to 512, the number of input neurons to 24×16, set the number of input channels of the first convolutional layer to 512, the number of output channels to 256, the kernel size to 1×1, the step to 1, and the dilation coefficient to 1. Set the number of input channels of the first pooling layer to 256, the kernel size to 3×3, and the step to 1;

[0026] Step 3: Construct the decoder of the two-stream deep neural network:

[0027] Step 3.1: The decoder consists of a first input layer, a second input layer, a decoder sub-module group, and a network output sub-module connected in series in sequence;

[0028] Step 3.2: Connect four decoder sub-modules with the same structure in series to form a decoder sub-module group. The first input layer inputs the feature tensor of the current state, and the second input layer inputs the output feature tensor of the corresponding double residual block in the original image sampling unit. The decoder sub-module consists of a feature tensor concatenation layer, a first non-linear operation layer, and a double residual block connected in series in sequence;

[0029] Step 3.3: Connect the first convolutional layer and the second non-linear operation layer in series to form a network output sub-module;

[0030] Step 3.4, set the decoder parameters as follows. According to the output feature tensor of the feature fusion layer, set the number of input channels of the feature tensor concatenation layer to 512, the number of neurons to 24×16, set the first non-linear operation layer as the activation function ReLU, set the number of input channels of the double residual block to 512, the number of output channels to 128, set the number of input channels of the first convolutional layer to 32, the number of output channels to 1, the convolutional kernel to 1×1, the stride to 1, and the dilation coefficient to 1;

[0031] Step 4, construct a two-stream deep neural network:

[0032] Connect the encoder, the feature fusion module, and the decoder in series to form a two-stream deep neural network;

[0033] Step 5, generate a training set:

[0034] Step 5.1, form an image forgery sample set from at least 4500 forged image samples. The sample set should include at least three different types of forged images, and the proportion of each type of image is required to be 1:1:1;

[0035] Step 5.2, perform cropping and normalization operations on each sample in the image forgery sample set in sequence, and form a training set from all the normalized samples;

[0036] Step 6, train the two-stream deep neural network:

[0037] Input the training data set into the two-stream deep neural network batch by batch, and use the gradient optimization algorithm of Stochastic Gradient Descent (SGD) to optimize the network parameters, iteratively update the weight values and learning rate of the convolutional neural network until the cross-entropy loss function of the network converges, and obtain a trained two-stream deep neural network;

[0038] Step 7, locate the forged area of the digital image:

[0039] Step 7.1, perform preprocessing operations of cropping and normalization on the digital image to be located in sequence;

[0040] Step 7.2, input the preprocessed digital image into the trained two-stream deep neural network and output the predicted image after forgery localization.

[0041] Compared with the prior art, the present invention has the following advantages:

[0042] First, in the process of feature extraction of forged images, the present invention designs a high-pass filter module and introduces it into the high-pass filter sampling sub-network to extract the forged edge features of the image, and the extracted feature map is learned through the iterative neural unit in the high-pass filter sampling sub-network. This design idea overcomes the lack of key feature information of image forgery in the existing methods, enables the neural network designed by the present invention to effectively learn the contour of the image forgery area, and improves the accuracy of digital image forgery area positioning.

[0043] Second, the present invention designs a neuron module as a feature fusion module to fuse the two feature maps extracted by the high-pass filter sampling sub-network and the original image sampling unit in the encoder. This feature fusion module can realize the transformation of the feature map dimension and the fusion of feature information through convolution operation, pooling operation, and non-linear operation without losing the key information of the feature map. This design idea overcomes the defects of low training efficiency and feature loss after dimensionality reduction in the existing bilinear pooling method, enables the present invention to reduce the loss of network training time, reduce the loss of key information during feature fusion, and improve the efficiency of digital image forgery area positioning.

[0044] Third, the training set data generated by the present invention is large in scale and complete in image forgery types. The training set has forged image samples of three forgery types: splicing, copy-moving, and deletion. Therefore, the digital image forgery localization model obtained after training has universality, overcomes the limitation that the existing methods can only process a certain type of specific forged images, and enables the present invention to improve the universality of digital image forgery area positioning. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 is the flowchart of the present invention;

[0046] Figure 2 is the structural schematic diagram of the double residual block constructed by the present invention;

[0047] Figure 3 is the structural schematic diagram of the high-pass filter module constructed by the present invention;

[0048] Figure 4 is the structural schematic diagram of the single residual block constructed by the present invention;

[0049] Figure 5 is the structural schematic diagram of the decoder constructed by the present invention;

[0050] Figure 6 is the schematic diagram of digital image forgery area positioning of the present invention, where Figure 6 (a) is the forged image to be located, Figure 6 (b) is the schematic diagram of the true label of the forged image forgery positioning to be located, Figure 6(c) is the localization result diagram of the forged area of the forged image to be located. Detailed implementation manners

[0051] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments.

[0052] Refer to Figure 1 and embodiments, and the implementation steps of the digital image forgery area localization method of the present invention will be further described in detail.

[0053] Step 1: Construct an encoder in the two-stream deep neural network.

[0054] Step 1.1: Build an original image sampling unit composed of four serially connected double residual blocks with the same structure.

[0055] Refer to Figure 2 for a further description of the structure of the double residual block.

[0056] Step 1.1.1: The double residual block is composed of an input layer, a first convolutional layer, a first non-linear operation layer, a second convolutional layer, a first residual convolutional layer, a non-linear addition operation unit, a second non-linear operation layer, a third convolutional layer, a fourth convolutional layer, a gated convolutional layer, a gated non-linear operation layer, a third non-linear operation layer, a second residual convolutional layer, and an output layer.

[0057] The first convolutional layer, the first non-linear operation layer, the second convolutional layer, the non-linear addition operation unit, and the first residual convolutional layer are serially connected in sequence to form a residual block for the first residual propagation.

[0058] The third convolutional layer, the second non-linear operation layer, the fourth convolutional layer, the third non-linear operation layer, and the second residual convolutional layer are serially connected in sequence to form a residual block for the second residual propagation.

[0059] The input layer is respectively connected to the first convolutional layer, the first residual convolutional layer, and the second residual convolutional layer, and then serially connected to the residual block for the first residual propagation, the gated convolutional layer, the gated non-linear operation layer, the residual block for the second residual propagation, and the output layer in sequence to form the double residual block as a whole.

[0060] Step 1.1.2, set the parameters of the double residual block as follows. Set the channel parameter of the input layer to 3 according to the size of the image resolution, and set the number of input neurons to 384×256; set the input channels of the first, third convolutional layers, the first and second residual convolutional layers to 3, the output channels to 32, the kernel size to 3×3, the sliding step to 1, and the dilation coefficient to 1. Set the input and output channel numbers of the second and fourth convolutional layers to 32, the kernel size to 3×3, the sliding step to 1, and the dilation coefficient to 1; set the input channel number of the gated convolutional layer to 32, the output channel size to 3, the kernel size to 3×3, the sliding step to 1, and the dilation coefficient to 1; the first to third non-linear operation layers and the non-linear addition operation unit are all implemented using the activation function ReLU; the input and output channel numbers of the output layer are both set to 32.

[0061] Step 1.1.3, calculate the result output after the first residual propagation of each input image through the double residual block according to the following formula:

[0062] j = σ1(Conv2(σ2(Conv1(i))) + dConv1(i))

[0063] where j represents the result of the first residual propagation, σ1(.) represents performing ReLU operation using the non-linear addition operation unit, Conv1(i) represents the result of performing a two-dimensional convolution operation with a kernel size of 3×3 and a step of 1 on the i-th digital image input to the original image sampling unit (the digital image is a triple image), σ2(.) represents performing ReLU operation on Conv1(i) using the first non-linear operation layer, Conv2(.) represents the result of performing another two-dimensional convolution operation on the result of σ2(.) using the second convolutional layer, and dConv1(i) represents the result of performing a two-dimensional convolution operation with a kernel size of 3×3 and a step of 1 on i using the first residual convolutional layer.

[0064] The gated convolutional layer and the gated non-linear operation layer constitute the gating mechanism. Before the second residual mapping of the double residual block, the feature information of the input i-th digital image and the result j of the first residual propagation of the double residual block are superimposed through the gating mechanism, and the dimensions are adjusted to meet the input conditions of the residual block for the second residual propagation of the double residual block.

[0065] Calculate the result output after passing the input each image and the result j of the first residual propagation through the gating mechanism according to the following formula.

[0066]

[0067] Among them, Sigmoid(.) is an activation function, and its operation result is distributed in the interval (0, 1). j represents the result of the first residual propagation, and Conv(j) represents the result of performing a two-dimensional convolution operation with a kernel size of 3×3 and a stride of 1 on j.

[0068] According to the following formula, calculate the result output after the second residual propagation of the double residual block for the i-th input image and the result j of the first residual propagation:

[0069] o = σ3(Conv4(σ4(Conv3(g)))) + dConv2(i)

[0070] Among them, o represents the result of the second residual propagation of the double residual block. σ3(.) represents performing ReLU operation through the third non-linear operation layer. Conv3(g) represents the result of performing a two-dimensional convolution operation with a kernel size of 3×3 and a stride of 1 on the output result g of the gating mechanism using the third convolution layer. σ4(.) represents performing ReLU operation on Conv3(g). Conv4(.) represents performing another two-dimensional operation on the operation result of σ4(.) using the fourth convolution layer. dConv2(i) represents the result of performing a two-dimensional convolution operation with a kernel size of 3×3 and a stride of 1 on i using the second residual convolution layer.

[0071] According to the following formula, calculate the result output after the result o of the second residual propagation of the double residual block passes through the output layer:

[0072] p = Blurpool(Maxpool(o))

[0073] Among them, p represents the operation result of the output layer. Maxpool(o) represents performing a max pooling operation with a kernel size of 3×3 and a stride of 1 on o. Blurpool(.) represents performing a blur pooling operation with a kernel size of 3×3 and a stride of 1 on Maxpool(o).

[0074] The reason for extracting the essential features of the forged image through the double residual block in the original sampling sub-network is that the double residual block enhances the ability to extract image forgery features by allowing the output feature map to perform two residual propagations. In the double residual block, the result obtained from the first residual propagation serves as the input feature for the second residual propagation, and the second residual propagation can integrate these features and stabilize the extracted feature information.

[0075] Step 1.2, build a high-pass image sampling sub-network composed of a high-pass filtering module and four single residual blocks with the same structure connected in series in sequence.

[0076] Refer to Figure 3 and Figure 4, a further description of the structures of the high-pass filtering module and the single residual block is given.

[0077] Step 1.2.1, the high-pass filtering module consists of an input layer, a horizontal high-pass kernel, a vertical high-pass kernel, a superposition operation unit, and a high-pass filtering output layer.

[0078] After successively connecting the horizontal high-pass kernel, the vertical high-pass kernel, and the superposition operation unit in series, they are respectively connected to the input layer to form a high-pass filter. The high-pass filter is connected in series with the high-pass filtering output layer to form the high-pass filtering module.

[0079] Step 1.2.2, the single residual block includes an input layer, a residual block for single residual propagation, and an output layer. Among them, the residual block for single residual propagation is formed by successively connecting a first convolutional layer, a first non-linear operation layer, a second convolutional layer, a second non-linear operation layer, and a first residual convolutional layer. The input layer is respectively connected to the first convolutional layer and the first residual convolutional layer, and then connected in series with the output layer to form the single residual block.

[0080] Step 1.2.3, set the parameters of the high-pass filtering module as follows. Set the channel parameter of the high-pass filtering module to 3 according to the size of the resolution of the input sample image, set the number of input neurons to 384×256, and set the convolutional kernel size to 3×3; the parameters of the high-pass convolutional kernel are as follows:

[0081]

[0082] Among them, H x is the convolutional kernel in the horizontal direction, and H y is the convolutional kernel in the vertical direction.

[0083] Step 1.2.4, set the parameters of the single residual block as follows. Set the input channel numbers of the first convolutional layer and the first residual convolutional layer to 3, the output channel numbers to 32, the convolutional kernel sizes to 3×3, the sliding strides to 1, and the dilation coefficients to 1; set the input and output channel numbers of the second convolutional layer to 32, the convolutional kernel size to 3×3, the sliding stride to 1, and the dilation coefficients to 1; implement the first and second non-linear operation layers using the activation function ReLU. The input and output channel numbers of the output layer are both 32.

[0084] Step 1.2.5, calculate the result output after each input image passes through the high-pass filtering module according to the following formula:

[0085]

[0086] Among them, x represents the result of performing a convolution operation on the i-th input image sample using a high-pass kernel. The reason for using a high-pass filter in the high-pass image sampling sub-network is that the high-pass filter filters the input image to extract the edge features in the forged image, which is beneficial for the subsequent 4 single residual blocks to extract the forged edge information of the image and helps to determine the range of the image forgery area.

[0087] Step 1.2.6, according to the following formula, calculate the result output by the residual block after a single residual propagation of the result x of performing a convolution operation on each input image sample using a high-pass kernel through the single residual block:

[0088] y = σ(Conv6(σ5(Conv7(x))) + resConv(x))

[0089] Among them, y represents the result of a single residual propagation, σ(.) represents performing a ReLU operation through the second non-linear operation layer, Conv7(x) represents the result of performing a two-dimensional convolution operation with a kernel size of 3×3 and a stride of 1 on x using the first convolution layer, σ5(.) represents performing a ReLU operation on Conv7(x), Conv6(x) represents the result of performing a two-dimensional convolution operation with a kernel size of 3×3 and a stride of 1 on the result of the σ5 operation using the second convolution layer, and resConv(x) represents the result of performing a two-dimensional convolution operation with a kernel size of 3×3 and a stride of 1 on x using the first residual convolution layer.

[0090] According to the following formula, calculate the result output after the result y of a single residual propagation passes through the output layer:

[0091] z = Blurpool1(Maxpool1(y))

[0092] Among them, z represents the operation result of the output layer, Maxpool1(y) represents performing a max pooling operation with a kernel size of 3×3 and a stride of 1 on y, and Blurpool1(.) represents performing a blur pooling operation with a kernel size of 3×3 and a stride of 1 on Maxpool1(y).

[0093] The purpose of using a single residual block in the high-pass image processing sub-network is to reduce the impact of gradient disappearance in the neural network, accelerate the convergence of the entire neural network, and make the edge feature information of the image forgery area learned by the high-pass image processing sub-network more sufficient.

[0094] Step 1.3, connect the original image sampling unit and the high-pass image sampling sub-network in series to form an encoder.

[0095] Step 2, construct the feature fusion module of the two-stream deep neural network.

[0096] Step 2.1, the feature fusion module is composed of a feature concatenation layer, a first convolutional layer, a first non-linear operation layer, and a first pooling layer connected in series in sequence.

[0097] Step 2.2, the parameters of the feature fusion module are set as follows. According to the results output by the original image sampling unit and the high-pass image sampling sub-network, set the input channel number of the feature concatenation layer to 512, and set the number of input neurons to 24×16. Set the input channel number of the first convolutional layer to 512, the output channel number to 256, the kernel size to 1×1, the stride to 1, and the dilation coefficient to 1. Set the input channel number of the first pooling layer to 256, the kernel size to 3×3, and the stride to 1.

[0098] Step 2.3, according to the following formula, calculate the result output after the feature tensor s1 output by the original image sampling unit and the feature tensor s2 output by the high-pass image sampling sub-network pass through the feature fusion module:

[0099] f = Maxpool2(σ(Conv2(Concat(s1, s2))))

[0100] Where, f represents the operation result of the feature fusion module, s1 represents the feature tensor obtained after the input image x learns the forgery feature through the original image sampling unit, s2 represents the feature tensor obtained after the input image x learns the forgery area edge feature through the high-pass image sampling sub-network, Concat(s1, s2) represents the process of concatenating the two feature tensors (s1, s2), Conv2(.) represents a convolutional operation with a kernel size of 1×1, a stride of 1, and a dilation coefficient of 1. σ(.) represents the ReLU operation through the first non-linear operation layer, and Maxpool2(.) represents a max pooling operation with a kernel size of 3×3 and a stride of 1.

[0101] The reason for introducing the feature fusion module in the two-stream deep neural network is that the operations such as convolutional operation, pooling operation, and non-linear operation in the feature fusion module can all be executed on the graphics processing unit (GPU) of the computer, greatly improving the efficiency of network training; the convolutional operation and pooling operation in the feature fusion layer can change the dimension of the feature tensor without losing key features, avoiding the defect of key feature loss that may occur in dimensionality reduction operations such as principal component analysis and bilinear pooling.

[0102] Step 3, construct the decoder of the two-stream deep neural network.

[0103] Refer to Figure 5 Make a further description of the decoder of the two-stream deep neural network.

[0104] Step 3.1, the decoder is composed of a first input layer, a second input layer, a decoder sub-module group, and a network output sub-module connected in series in sequence. The decoder sub-module group is composed of four decoder sub-modules with the same structure. Among them, the first input layer inputs the feature tensor of the current state, and the second input layer inputs the output feature tensor of the corresponding double residual block in the original image sampling unit. For example, the input of the first decoder sub-module is the output feature tensor of the feature fusion module and the output feature tensor of the fourth double residual block in the original image sampling unit. The input layer of the second decoder sub-module is the result of concatenating the output feature tensor of the first decoder sub-module and the output feature tensor of the third double residual block in the original image sampling unit, and so on. The decoder sub-module is composed of a feature tensor concatenation layer, a first non-linear operation layer, and a double residual block connected in series in sequence. The network output sub-module is composed of a first convolutional layer and a second non-linear operation layer connected in series.

[0105] Step 3.2, set the decoder parameters as follows. According to the output feature tensor of the feature fusion layer, set the input channel number of the feature tensor concatenation layer to 512, and the number of neurons to 24×16; set the first non-linear operation layer as the activation function ReLU. Set the input channel number of the double residual block to 512, and the output channel number to 128. Set the input channel number of the first convolutional layer to 32, the output channel number to 1, the convolutional kernel to 1×1, the stride to 1, and the dilation coefficient to 1.

[0106] Step 3.3, according to the following formula, calculate the result output after the output feature tensor d1 of the feature fusion module and the output feature tensor e4 of the fourth double residual block in the original image sampling unit pass through the decoder sub-module:

[0107] c = dRes(σ8(Concat1(d1, e4)))

[0108] Where, c represents the output result of the decoder sub-module, d1 represents the output feature tensor of the feature fusion module, e4 represents the output feature tensor of the fourth double residual block in the original image sampling unit, Concat1(d1, e4) represents the process of concatenating the above content; σ8(.) represents performing ReLU operation on Concat1(d1, e4) through the first non-linear operation layer; dRes(.) represents the double residual block.

[0109] The function of the decoder sub-module is: the input feature tensor of each sub-module of the decoder is fused with the output feature tensor of the corresponding double residual block in the original image sampling unit in the encoder module, which can effectively repair the feature information and dimension of the feature tensor input to the decoder sub-module and enhance the stability of the network.

[0110] According to the following formula, calculate the result output after the output feature tensor c4 of the last decoder sub-module passes through the network output sub-module:

[0111] u = Sigmoid1(Conv3(c4))

[0112] Among them, u represents the output of the encoder part and the output feature tensor of the last decoder sub-module; Conv3(c4) represents the result of performing a convolution operation on c4 with a convolution kernel of 1×1, a stride of 1, and a dilation coefficient of 1; Sigmoid1(.) represents performing a non-linear operation on Conv3(c4).

[0113] Step 4, construct a two-stream deep neural network.

[0114] Connect the encoder, the feature fusion module, and the decoder in series to form a two-stream deep neural network.

[0115] Step 5, generate a training set.

[0116] Step 5.1, the embodiments of the present invention collect and organize a total of 4542 forged image samples from four different publicly available digital image forgery datasets to form an image forgery sample set. That is, 2000 spliced-type forged images and 2000 copy-move-type forged images are extracted from the 'Casia' CASIA dataset; 70 copy-move-type forged images are extracted from the 'Coverage' COVERAGE (abbreviated as COVER) dataset; 180 spliced-type forged images are extracted from the 'Realistic' REALISTIC (abbreviated as RT) dataset; 146 spliced-type forged images and 146 deletion-type forged images are extracted from the 'Nimble Challenge 2016' Nimble Challenge2016 (abbreviated as NC16) dataset.

[0117] Step 5.2, the embodiments of the present invention extract 15000 images from the 'Defacto' DEFACTO dataset, which include spliced forgery type, copy-move forgery type, and deletion forgery type images, with 5000 images of each forgery type to form an image forgery extended sample set.

[0118] Step 5.3, the embodiments of the present invention combine the image forgery sample set and the image forgery extended sample set to form a sample set for implementing the digital image forgery area localization method. This sample set includes a total of 17542 forged image samples. Each sample is cropped according to the 384×256 specification and normalized using OpenCV-python. All the normalized samples are combined to form a training set.

[0119] Step 6, train the two-stream deep neural network.

[0120] Step 6.1: Input the training dataset into the two-stream deep neural network batch by batch. Calculate the loss value of the two-stream deep neural network after inputting the selected forged images using the binary cross-entropy loss function. Optimize the network parameters using the gradient optimization algorithm of stochastic gradient descent (SGD), and iteratively update the weight values and learning rate of the convolutional neural network until the binary cross-entropy loss function of the network converges, obtaining a trained two-stream deep neural network. Store the weight values of each network layer of the two-stream deep neural network at the time of convergence in a '.pth' file.

[0121] The binary cross-entropy loss function is as follows:

[0122]

[0123] where loss represents the forged image to be localized input into the two-stream deep neural network. After one round of iteration, it outputs the loss value of the network. ∑[.] represents the summation operation, log represents the logarithm operation with base 10, N represents the total number of image samples in the training set, i represents the index of the image sample in the training set, P i represents the predicted result of the forged localization output by the trained two-stream deep neural network for the i-th image in the training set, and G i represents the true label of the forged localization of the i-th image in the training set.

[0124] Step 7: Locate the forged region of the digital image.

[0125] Step 7.1: Crop the digital image to be localized according to the specification of 384×256 and perform normalization operations using OpenCV-python to implement the preprocessing operation of the digital image to be localized.

[0126] Step 7.2: Input the preprocessed digital image into the trained two-stream deep neural network to localize the forged region of the input image. The output of the network is the result of localizing the forged region of the input image. Among them, the pixel points with a pixel value of 255 in the output digital image represent the forged region, and the pixel points with a pixel value of 0 represent the real region.

[0127] Refer to Figure 6 for further explanation of localizing the forged region of the digital image, where Figure 6 (a) is the forged image to be localized; Figure 6 (b) is a schematic diagram of the true label of the forged localization of the forged image to be localized. The white area in the figure represents the forged region; Figure 6 (c) is the localization result diagram of the forged region of the forged image to be localized. The white area in the figure represents the forged region predicted by the two-stream deep neural network for the input image. Figure 6(c) It can be seen that the trained two-stream deep neural network of the present invention can accurately locate the forged regions in images.

[0128] The effects of the present invention will be further described below with reference to the simulation diagrams.

[0129] 1.1, Simulation experiment conditions:

[0130] The simulation experiment platform of the present invention is: a cloud server with a graphics card RTX-3090 and a memory of 24G, and the two-stream deep neural network is built and trained using Pytorch (version 1.8).

[0131] 1.2, Simulation experiment test set:

[0132] The present invention collected and sorted a total of 2370 forged image samples from five different publicly available digital image forgery datasets to form a digital image forgery test set. That is, 1000 forged images of the splicing type and 1000 forged images of the copy-move type were extracted from the 'CASIA' dataset; 30 forged images of the copy-move type were extracted from the 'COVER' dataset; 40 forged images of the splicing type were extracted from the 'RT' dataset; 100 forged images of the splicing type were extracted from the 'In-wild' dataset; 100 forged images of the splicing type and 100 forged images of the deletion type were extracted from the 'NimbleChallenge2016' (abbreviated as 'NC16') dataset. Each image in the test set was cropped according to the specification of 384×256, and each sample was normalized using OpenCV-python.

[0133] 2.1, Simulation content and result analysis.

[0134] The simulation experiment of the present invention uses browser plugin development technology to solidify the digital image forgery area localization method into a Google Chrome security plugin, realizing the localization of the forged areas of Internet digital images in the Internet environment.

[0135] To verify the accuracy of the digital image forgery area localization method proposed by the present invention, the average accuracy rate and average F value of all prediction results of this method on the test set were tested. As shown in Table 1 and Table 2.

[0136] The calculation formula for the average accuracy rate is as follows:

[0137]

[0138] Among them, mAcc represents the calculation result of the average accuracy rate, M represents the number of forged image samples in the test set, and Rp i represents the number of all correctly predicted pixel points in the prediction result of the i-th image; Np i represents the total number of pixel points in the prediction result of the i-th image, and i represents the index of the image sample in the training set.

[0139] The calculation formula for the average F value is as follows:

[0140]

[0141] Among them, mF represents the calculation result of the average F value, M represents the number of forged image samples in the test set, and Tn i represents the number of all correctly predicted forged pixel points in the prediction result of the i-th image, Tp i represents the number of all forged pixel points in the prediction result of the i-th image, Ap i represents the number of all true forged pixel points in the prediction result of the i-th image, and i represents the index of the image sample in the training set.

[0142] Table 1 List of average accuracy rates of all prediction results on the test set of the two-stream deep neural network

[0143]

[0144] Table 2 List of average F values of all prediction results of the two-stream deep neural network on the test set

[0145]

[0146] To verify the robustness of the digital image forgery area localization method proposed by the present invention, under the external interference conditions of Gaussian noise, Gaussian blur, JPEG compression, rotation, and scaling, the average F value of all prediction results of the digital image forgery area localization method on the test set was tested. As shown in Tables 3 to 7.

[0147] Table 3 List of average F values of all prediction results on the test set under Gaussian noise interference

[0148]

[0149] Table 4 List of average F values of all prediction results on the test set under Gaussian blur interference

[0150]

[0151] Table 5 List of average F values of all prediction results on the test set under JPEG compression interference

[0152]

[0153] Table 6 List of average F-values of all prediction results on the test set under scaling interference

[0154]

[0155] Table 7 List of average F-values of all prediction results on the test set under rotation interference

[0156]

Claims

1. A method for locating forged regions in digital images based on a two-stream deep neural network, characterized in that, By using the double residual block, single residual block, and feature fusion module in the constructed dual-stream deep neural network to extract and learn the forgery features of the image, the ability to locate the forged area of the digital image is achieved; the steps of this location method are as follows: Step 1, construct the encoder in the dual-stream deep neural network: Step 1.1, build an original image sampling unit composed of four serially connected double residual blocks with the same structure; Step 1.2, the double residual block consists of an input layer, a first convolutional layer, a first non-linear operation layer, a second convolutional layer, a first residual convolutional layer, a non-linear addition operation unit, a second non-linear operation layer, a third convolutional layer, a fourth convolutional layer, a gated convolutional layer, a gated non-linear operation layer, a third non-linear operation layer, a second residual convolutional layer, and an output layer; Step 1.3, the first convolutional layer, the first non-linear operation layer, the second convolutional layer, the non-linear addition operation unit, and the first residual convolutional layer are serially connected in turn to form the residual block of the first residual propagation and the output result of the residual block of the first residual propagation; Step 1.4, the third convolutional layer, the second non-linear operation layer, the fourth convolutional layer, the third non-linear operation layer, and the second residual convolutional layer are serially connected in turn to form the residual block of the second residual propagation; Step 1.5, the input layer is respectively connected to the first convolutional layer, the first residual convolutional layer, and the second residual convolutional layer, and then serially connected to the residual block of the first residual propagation, the gated convolutional layer, the gated non-linear operation layer, the residual block of the second residual propagation, and the output layer in turn to form the double residual block as a whole; Step 1.6, according to the size of the image resolution, set the channel parameter of the input layer to 3, the number of input neurons to 384×256, set the input channels of the first and third convolutional layers, the first and second residual convolutional layers to 3, the output channels to 32, the kernel size to 3×3, the sliding step to 1, and the dilation coefficient to 1. Set the input and output channel numbers of the second and fourth convolutional layers to 32, the kernel size to 3×3, the sliding step to 1, and the dilation coefficient to 1. Set the input channel number of the gated convolutional layer to 32, the output channel size to 3, the kernel size to 3×3, the sliding step to 1, and the dilation coefficient to 1. The first to third non-linear operation layers and the non-linear addition operation unit are all implemented using the activation function ReLU, and the input and output channel numbers of the output layer are both set to 32; Step 1.7, build a high-pass image sampling sub-network composed of a high-pass filter module and four serially connected single residual blocks with the same structure; Step 1.8, the high-pass filter module consists of an input layer, a horizontal high-pass kernel, a vertical high-pass kernel, a superposition operation unit, and a high-pass filter output layer; The horizontal high-pass kernel is as follows: Among them, H x represents the convolutional kernel in the horizontal direction; The vertical high-pass kernel is as follows: Among them, H y represents the convolution kernel in the vertical direction; Step 1.9, after serially connecting the horizontal high-pass kernel, the vertical high-pass kernel, and the superposition operation unit in turn, connect them to the input layer respectively to form a high-pass filter, and the high-pass filter is serially connected to the high-pass filter output layer to form a high-pass filter module; Step 1.10: Set the channel parameter of the high-pass filtering module to 3 according to the size of the input sample image resolution, set the number of input neurons to 384×256, and set the convolution kernel size to 3×3; Step 1.11: The single residual block includes an input layer, a residual block for single residual propagation, and an output layer. Among them, the residual block for single residual propagation is formed by a first convolutional layer, a first non-linear operation layer, a second convolutional layer, a second non-linear operation layer, and a first residual convolutional layer connected in series in sequence. The input layer is connected to the first convolutional layer and the first residual convolutional layer respectively, and then connected in series with the output layer to form a single residual block; Step 1.12: Set the input channel numbers of the first convolutional layer and the first residual convolutional layer to 3, the output channel numbers to 32, the convolution kernel sizes to 3×3, the sliding strides to 1, and the dilation coefficients to 1. Set the input and output channel numbers of the second convolutional layer to 32, the convolution kernel size to 3×3, the sliding stride to 1, and the dilation coefficients to 1. The first and second non-linear operation layers are implemented using the activation function ReLU, and the input and output channel numbers of the output layer are both 32; Step 1.13: Connect the original image sampling unit in series with the high-pass image sampling sub-network to form an encoder; Step 2: Construct the feature fusion module of the two-stream deep neural network: Step 2.1: The feature fusion module is composed of a feature splicing layer, a first convolutional layer, a first non-linear operation layer, and a first pooling layer connected in series in sequence; Step 2.2: According to the results output by the original image sampling unit and the high-pass image sampling sub-network, set the input channel number of the feature splicing layer to 512, set the number of input neurons to 24×16, set the input channel number of the first convolutional layer to 512, the output channel number to 256, the kernel size to 1×1, the stride to 1, and the dilation coefficient to 1. Set the input channel number of the first pooling layer to 256, the kernel size to 3×3, and the stride to 1; Step 3: Construct the decoder of the two-stream deep neural network: Step 3.1: The decoder is composed of a first input layer, a second input layer, a decoder sub-module group, and a network output sub-module connected in series in sequence; Step 3.2: Connect four decoder sub-modules with the same structure in series to form a decoder sub-module group. The first input layer inputs the feature tensor of the current state, and the second input layer inputs the output feature tensor of the corresponding double residual block in the original image sampling unit. The decoder sub-module is composed of a feature tensor splicing layer, a first non-linear operation layer, and a double residual block connected in series in sequence; Step 3.3: Connect the first convolutional layer and the second non-linear operation layer in series to form a network output sub-module; Step 3.4: Set the decoder parameters as follows. According to the output feature tensor of the feature fusion layer, set the input channel number of the feature tensor splicing layer to 512, the number of neurons to 24×16, set the first non-linear operation layer as the activation function ReLU, set the input channel number of the double residual block to 512, the output channel number to 128, set the input channel number of the first convolutional layer to 32, the output channel number to 1, the convolution kernel to 1×1, the stride to 1, and the dilation coefficient to 1; Step 4: Construct the two-stream deep neural network: Connect an encoder, a feature fusion module, and a decoder in series to form a two-stream deep neural network; Step 5, generate a training set: Step 5.1, compose at least 4500 forged image samples into an image forgery sample set, which includes at least three different types of forged images, and the proportion requirement for each type of image is 1:1:1; Step 5.2, perform cropping and normalization operations on each sample in the image forgery sample set in sequence, and compose all the normalized samples into a training set; Step 6, train the two-stream deep neural network: Input the training data set into the two-stream deep neural network batch by batch, use the gradient optimization algorithm of stochastic gradient descent (SGD) to optimize the network parameters, iteratively update the weight values and learning rate of the convolutional neural network until the cross-entropy loss function of the network converges, and obtain the trained two-stream deep neural network; Step 7, locate the forged area of the digital image: Step 7.1, perform preprocessing operations of cropping and normalization on the digital image to be located in sequence; Step 7.2, input the preprocessed digital image into the trained two-stream deep neural network, and output the predicted image after forgery localization.

2. The method for locating forged regions of digital images based on a two-stream deep neural network according to claim 1, wherein The cross-entropy loss function described in Step 6 is as follows: Among them, loss represents the forged image to be located input into the dual-stream deep neural network. After one round of iteration, it is the loss value output by the network. ∑[.] represents the summation operation, log represents the logarithm operation with base 10, N represents the total number of image samples in the training set, i represents the index of the image sample in the training set, and P i represents the predicted result of forged localization output by the trained dual-stream deep neural network for the i-th image in the training set, and G i represents the true label of the forged localization of the i-th image in the training set.

Citation Information

Patent Citations

  • Method for detecting spliced and tampered images

    CN111062931A

  • A method for detecting spliced ​​and tampered images

    CN111062931B

  • Image tampering detection method based on double-flow convolutional neural network

    CN113129261A

  • Tampered image detection method based on deep learning

    CN110349136A

  • Video image identification method based on residual error-capsule network

    CN111241958A