Image forgery localization method based on edge-aware regional message passing control
By adopting an edge-aware area message delivery control method in image forgery positioning, the boundaries between the forgery area and the real area are explicitly modeled, and the misjudgment problem caused by feature coupling in the prior art is solved, and higher positioning accuracy and lower misjudgment rate are achieved.
Patent Information
- Application Number
- CN202310346607.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-03
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2043-04-03
AI Technical Summary
The existing image forgery positioning methods fail to effectively decouple the features of the forged area and the real area, resulting in misjudgment and reduced accuracy.
Adopting an edge-aware regional messaging control method, the boundary between the forged area and the real area is explicitly modeled through Bayar convolutional layer, backbone network, edge reconstruction module, area messaging controller and fusion module to control message delivery in the feature map.
Effectively capture tamper artifacts, improve the accuracy of image forgery positioning, reduce the rate of error judgment, and perform better than previous methods on multiple benchmarks.
Smart Images

Figure CN116452528B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image forgery localization, and specifically to an image forgery localization method based on edge-aware control region message passing. Background Art
[0002] With the development of deep learning technology, image forgery technology has become increasingly mature. Image tampering poses risks in various fields, such as deleting copyright watermarks, creating fake news, and even perjuring in court. These image tamperings can trigger a trust crisis and affect social order. Therefore, image tampering detection and localization are of great significance.
[0003] The key to image tampering detection is to model the inconsistency between the forged area and the real area, and locate the forged area on the suspicious image. However, with the wide application of GAN and VAE technologies, image tampering is visually difficult to detect. Therefore, in order to accurately locate the image forged area, it is particularly important to decouple the features between the forged area and the real area.
[0004] In recent years, with the development of deep learning, significant results have been achieved in image forgery localization. However, some existing image forgery localization methods do not explicitly decouple the features of the forged area and the real area, resulting in misjudgment and affecting the accuracy of image forgery localization. Therefore, the problem of how to better utilize the differences between the forged area and the real area for tampering localization still needs to be explored. Summary of the Invention
[0005] The present invention is to avoid the problems existing in the above-mentioned prior art, and proposes an image forgery localization method based on edge-aware region message passing control, in order to solve the problem of severe coupling of the features of the forged area and the real area, effectively determine the boundary between the forged area and the real area, and thus improve the accuracy of image forgery localization.
[0006] The present invention adopts the following technical solutions to achieve the above invention purpose:
[0007] The characteristics of an image forgery localization method based on edge-aware region message passing control of the present invention are as follows: the method is carried out according to the following steps:
[0008] Step 1: Obtain a forged image to be detected and perform preprocessing to obtain the preprocessed forged image C represents the number of channels of the forged image, and H and W respectively represent the height and width of the forged image;
[0009] Step 2: Construct an image forgery localization network, including: Bayar convolutional layer, backbone network, edge reconstruction module, region message passing controller, and fusion module;
[0010] Step 2.1: The Bayar convolution extracts the noise of the preprocessed image X to obtain a noise map X'.
[0011] Step 2.2: The backbone network is composed of an RGB branch and a noise branch, and each branch is composed of a pre-trained ResNet50 network and an Atrous Spatial Pyramid Pooling (ASPP);
[0012] The RGB branch processes the preprocessed forged image X to obtain RGB coarse features
[0013] The noise branch processes the noise map X' to obtain noise coarse features where C s represents the number of channels of the coarse features, and H s and W s represent the height and width of the features respectively;
[0014] Step 2.3: The edge reconstruction module includes a Sobel layer, a Contextual Enhancement Graph (CEG), and a differentiable binarization module, which are used to convert the RGB coarse features into edge features G e ;
[0015] Step 2.4: A regional message passing controller is constructed, and under the guidance of the edge features between the forged region and the real region, the RGB coarse features are constructed into RGB features Z after graph reasoning r ;
[0016] Step 2.5: The fusion module uses Equation (9) to obtain the fusion features and then performs bilinear upsampling on G o to obtain the predicted mask of the forged image X
[0017] G o = DA(G z , G n ) (9)
[0018] In Equation (9), DA represents the Dual Attention Network, is the reshaped graph representation of Z r ;
[0019] Step 3: Use Equation (10) to construct a Dice loss function L, and use the ADAM optimizer to train the image forgery localization network until the loss function L converges, so as to obtain a trained image forgery localization model for realizing the recognition and localization of forged images;
[0020]
[0021] In formula (10), represents the Dice loss, is the true mask label of the forged image X, is the edge feature of the true mask label, is the edge feature downsampled by E, and λ1, λ2, and λ3 are three weight factors, and λ1 + λ2 + λ3 = 1.
[0022] The feature of the image forgery localization method based on edge-aware regional message passing control according to the present invention also lies in that the step 2.3 includes:
[0023] Step 2.3.1, the Sobel layer extracts the RGB coarse feature G r of the edge-related feature
[0024] G c = G r ⊙σ(Norm(Sobel(G r )) (1)
[0025] In formula (1), Norm is the L2 normalization operation, σ is the Sigmoid function, Sobel is the Sobel convolution operation, and ⊙ represents the Hadamard product;
[0026] Step 2.3.2, the context information enhanced graph CEG reshapes the edge-related feature G c into the edge-related graph representation N represents the number of nodes of the edge-related graph representation, and N = H s ×W s ;
[0027] Step 2.3.3, the context information enhanced graph CEG obtains the adjacency matrix
[0028] A c = Norm(Conv(Conv(G′ c ))) (2)
[0029] In formula (2), Conv represents the convolution operation with a convolution kernel of 1×1;
[0030] Step 2.3.4, the context information enhanced graph CEG obtains the global feature map Global(G′ c ):
[0031] Global(G′ c ) = A c G′ c W c (3)
[0032] In Equation (3), represents the parameter to be learned in the context information enhanced graph CEG;
[0033] Step 2.3.5, the context information enhanced graph CEG extracts edge-related feature G c using two convolutional layers with a kernel size of 1×1 for the local feature Local(G c );
[0034] Step 2.3.6, the context information enhanced graph CEG obtains the edge probability map using Equation (4)
[0035] G p = σ(Conv(Cat(Local(G c ), Global′(G′ c ))) (4)
[0036] In Equation (4), Cat represents the concatenation operation;
[0037] Step 2.3.7, the differentiable binarization module obtains the binarized edge feature using Equation (5) where H e and W e represent the height and width of the edge feature respectively;
[0038]
[0039] In Equation (5), τ is the transformation function to be learned, and k is the scaling factor.
[0040] The said Step 2.4 includes:
[0041] Step 2.4.1, determine the relationship between the i-th feature point P r and the j-th feature point P i in the RGB rough feature G j respectively and the edge feature G e . If the two feature points are respectively inside and outside the edge feature G e , let the relationship value XN(P i , P j ) be 0. If the two feature points are both inside or both outside the edge feature G e , let the relationship value XN(P i , P j ) be 1, so as to obtain the relationship values of all feature points in the RGB rough feature G r and form a relationship matrix N′ represents the dimension of the relationship matrix, and N′ = He ×W e ;
[0042] Step 2.4.2: Calculate the attention coefficient α of the i-th feature point P i and the j-th feature point P j to obtain the attention matrix; i,j
[0043] α i,j = ψ(x i ) T ψ′(x j ) (6)
[0044] In formula (6), ψ and ψ′ are linear transformation operations, and ψ = Wx i and ψ′ = W′x j , where W and W′ represent two parameter matrices to be learned, and T represents transpose;
[0045] Step 2.4.3: Normalize the attention matrix using the softmax function to obtain a preliminary adjacency matrix
[0046] Step 2.4.4: Dynamically adjust the adjacency matrix A r to obtain an adjacency matrix guided by edge features
[0047] A′ r = A r ⊙A e (7)
[0048] Step 2.4.5: Obtain the RGB features after graph reasoning using formula (8)
[0049] Z r = ReLU(A′ r G′ r W z ) (8)
[0050] In formula (8), ReLU represents an activation function, is a parameter matrix to be learned, is the reshaped graph representation of G r .
[0051] A feature of an electronic device according to the present invention, including a memory and a processor, is that the memory is used to store a program supporting the processor to execute any one of the image forgery localization methods, and the processor is configured to execute the program stored in the memory.
[0052] A computer-readable storage medium of the present invention, characterized in that a computer program is stored on the computer-readable storage medium, and when the computer program is run by a processor, it executes the steps of any one of the image forgery localization methods.
[0053] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0054] 1. The present invention proposes a novel two-step image forgery localization method from coarse to fine, which uses edge information to explicitly model the inconsistency between the forged area and the real area. This explicit modeling method can effectively capture tampering artifacts, thereby improving the accuracy of image forgery localization.
[0055] 2. The present invention proposes an edge-aware dynamic graph to control the message passing between the forged area and the real area in the feature map. This is the first time the concept of controlling message passing is introduced in the field of image forgery localization. This strategy can effectively avoid the coupling between the features of the image forgery area and the real area, thereby reducing the misjudgment rate of the model.
[0056] 3. The present invention proposes an edge reconstruction module containing a context enhancement graph and a threshold adaptive differentiable binarization module to obtain the required edge information. Previous studies used CNN to extract the edges between the forged area and the real area, but sometimes they could not accurately obtain them. This is because local information is not sufficient to detect carefully processed forged areas. The context enhancement graph designed in the present invention captures global information to help obtain edges; and threshold adaptive differentiable binarization is used to adaptively extract fuzzy edges, so that the method of the present invention can obtain the edges between the forged area and the real area more accurately than the previous methods.
[0057] 4. The present invention has conducted extensive experiments on multiple benchmarks and proved that the method of the present invention is superior to the previous image forgery localization methods both qualitatively and quantitatively. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 is a flowchart of an image forgery localization method based on edge-aware control area message passing of the present invention;
[0059] Figure 2 is an effect diagram of comparison between the present invention and various methods. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0060] In this embodiment, an image forgery localization method based on edge-aware area message passing control is mainly realized through edge reconstruction and message passing control between graph structure nodes. First, the input image Extract noise using Bayar convolution, and then use two branches to process the RGB image and the noise. The coarse features extracted from the RGB branch are converted into edges through an edge reconstruction block. Meanwhile, guided by the reconstructed edge information, the coarse features are constructed into a graph structure to achieve message passing between the controlled forged area and the real area. Finally, the RGB features and the noise information after the graph convolutional network are fused through a dual attention network to output the predicted image forgery localization map. Specifically, as Figure 1 shown, the steps are as follows:
[0061] Step 1: Obtain a forged image to be detected and perform preprocessing to obtain the preprocessed forged image C represents the number of channels of the forged image, and H and W represent the height and width of the forged image respectively;
[0062] Step 2: Construct an image forgery localization network, including: Bayar convolutional layer, backbone network, edge reconstruction module, regional message passing controller, fusion module;
[0063] Step 2.1: Use Bayar convolution to extract the noise of the preprocessed image X to obtain the noise map X';
[0064] Step 2.2: The backbone network consists of an RGB branch and a noise branch, and each branch is composed of a ResNet50 network pre-trained on the ImageNet dataset and an Atrous Spatial Pyramid Pooling (ASPP);
[0065] The RGB branch processes the preprocessed forged image X to obtain RGB coarse features
[0066] The noise branch processes the noise map X' to obtain noise coarse features where C s represents the number of channels of the coarse features, and H s and W s represent the height and width of the features respectively;
[0067] Step 2.3: The edge reconstruction module includes: Sobel layer, Contextual Enhancement Graph (CEG), differentiable binarization module, and is used to convert RGB coarse features into edge features;
[0068] Step 2.3.1: The Sobel layer extracts the edge-related features of the RGB coarse feature G r using Equation (1)
[0069] G c = G r ⊙ σ(Norm(Sobel(G r ))) (1)
[0070] In Equation (1), Norm is the L2 normalization operation, σ is the Sigmoid function, Sobel is the Sobel convolution operation, and ⊙ represents the Hadamard product;
[0071] Step 2.3.2: The context information enhanced graph CEG reshapes the edge-related feature G c into an edge-related graph representation N represents the number of nodes in the edge-related graph representation, and N = H s ×W s ;
[0072] Step 2.3.3: The context information enhanced graph CEG obtains the adjacency matrix using Equation (2)
[0073] A c = Norm(Conv(Conv(G′ c ))) (2)
[0074] In Equation (2), Conv represents the convolution operation with a 1×1 convolution kernel;
[0075] Step 2.3.4: The context information enhanced graph CEG obtains the global feature map Global(G′ c ):
[0076] Global(G′ c ) = A c G′ c W c (3)
[0077] In Equation (3), represents the parameter to be learned in the context information enhanced graph CEG;
[0078] Step 2.3.5: The context information enhanced graph CEG extracts the local feature Local(G c ) of the edge-related feature G using two convolution layers with 1×1 convolution kernels; c
[0079] Step 2.3.6: The context information enhanced graph CEG obtains the edge probability map using Equation (4)
[0080] G p = σ(Conv(Cat(Local(G c ), Global′(G′ c ))) (4)
[0081] In Equation (4), Cat represents the concatenation operation;
[0082] Step 2.3.7: The differentiable binarization module obtains the binarized edge features using Equation (5). Where H e and W e represent the height and width of the edge features respectively;
[0083]
[0084] In Equation (5), τ is a transformation function to be learned, and k is a scaling factor. In this embodiment, k is set to 500;
[0085] Step 2.4: Construct a regional message passing controller, and under the guidance of the edge features between the forged region and the real region, construct the RGB coarse features into RGB features after graph reasoning;
[0086] Step 2.4.1: Determine the relationship between the i-th feature point P r in the RGB coarse feature G i and the j-th feature point P j respectively and the edge feature G e . If the two feature points are respectively inside and outside the edge feature G e , let the relationship value XN(P i , P j ) be 0. If the two feature points are both inside or both outside the edge feature G e , let the relationship value XN(P i , P j ) be 1, so as to obtain the relationship values of all feature points in the RGB coarse feature G r and form a relationship matrix N′ represents the dimension of the relationship matrix, and N′ = H e ×W e ;
[0087] Step 2.4.2: Calculate the attention coefficient α i of the i-th feature point P j and the j-th feature point P i,j using Equation (6), so as to obtain an attention matrix;
[0088] α i,j = ψ(x i ) T ψ′(x j ) (6)
[0089] In Equation (6), ψ and ψ′ are linear transformation operations, and ψ = Wx i and ψ′ = W′x j , where W and W′ represent two parameter matrices to be learned, and T represents the transpose;
[0090] Step 2.4.3: Use the softmax function to normalize the attention matrix to obtain a preliminary adjacency matrix
[0091] Step 2.4.4: Dynamically adjust the adjacency matrix A r to obtain the adjacency matrix guided by edge features
[0092] A′ r = A r ⊙ A e (7)
[0093] After Equation (7), if two nodes are located in the forged area and the real area respectively, their connection will be cut off. The present invention realizes the control of message passing between regions through this step.
[0094] Step 2.4.5: Use Equation (8) to obtain the RGB features after graph reasoning
[0095] Z r = ReLU(A′ r G′ r W z ) (8)
[0096] In Equation (8), ReLU represents the activation function, is the parameter matrix to be learned, is the reshaped graph representation of G r ;
[0097] Step 2.5: The fusion module uses Equation (9) to obtain the fusion features and then performs bilinear upsampling operation on G o to obtain the predicted mask of the forged image X
[0098] G o = DA(G z , G n ) (9)
[0099] In Equation (9), DA represents the dual attention network, is the reshaped graph representation of Z r ;
[0100] Step 3: Use Equation (10) to construct the Dice loss function L, and use the ADAM optimizer to train the image forgery localization network until the loss function L converges, so as to obtain the trained image forgery localization model for realizing the recognition and localization of forged images;
[0101]
[0102] In formula (10), represents the Dice loss, is the true mask label of the forged image X, is the edge feature of the true mask label, is the edge feature downsampled by E. λ1, λ2, and λ3 are three weight factors, and λ1 + λ2 + λ3 = 1.
[0103] In this embodiment, an electronic device includes a memory and a processor. The memory is used to store a program that supports the processor to execute the above method, and the processor is configured to execute the program stored in the memory.
[0104] In this embodiment, a computer-readable storage medium stores a computer program on the computer-readable storage medium. When the computer program is run by a processor, it executes the steps of the above method.
[0105] Embodiment:
[0106] To verify the effectiveness of the present method, the commonly used Columbia dataset, Coverage dataset, CASIA dataset, NIST16 dataset, and IMD20 dataset are selected in this paper.
[0107] In this paper, AUC and F1 are used as evaluation criteria.
[0108] In this embodiment, 4 methods using pre-trained models are selected for comparison with the pre-trained model method of the present invention. The selected methods are ManTraNet, SPAN, PSCCNet, and ObjectFormer as the inventive methods. According to the experimental results, the results are shown in Table 1 as follows:
[0109] Table 1 Comparison of AUC (%) of tampering localization of different pre-trained models
[0110]
[0111]
[0112] In this embodiment, 6 methods of fine-tuning after pre-training are selected for comparison with the method of fine-tuning after pre-training of the present invention. The selected methods are J-LSTM, H-LSTM, RGB-N, SPAN, PSCCNet, and ObjectFormer as the inventive methods. According to the experimental results, the results are shown in Table 2 as follows:
[0113] Table 2 Results of AUC (%) and F1 (%) of tampering localization using fine-tuned models
[0114]
[0115] The experimental results show that the method of the present invention has better effects compared with other 4 pre-training methods and other 6 fine-tuning methods, thus proving the feasibility of the method proposed by the present invention.
[0116] Shown in Figure 2 is the comparison between the present invention and the predicted forgery masks of other available-source-code methods. The results show that the present invention can not only more accurately locate the tampered areas, but also form clear boundaries. This benefits from the explicit modeling of the inconsistency between the two regions and the full utilization of the edges in the present invention.
Claims
1. An image forgery localization method based on edge-aware regional message passing control, characterized in that, The steps are as follows: Step 1: Obtain a forged image to be detected and perform preprocessing to obtain the preprocessed forged image C represents the number of channels of the forged image, and H and W represent the height and width of the forged image respectively; Step 2: Construct an image forgery localization network, including: Bayar convolutional layer, backbone network, edge reconstruction module, regional message passing controller, and fusion module; Step 2.1: The Bayar convolution extracts the noise of the preprocessed image X to obtain a noise map X'; Step 2.2: The backbone network is composed of an RGB branch and a noise branch, and each branch is composed of a pre-trained ResNet50 network and a spatial pyramid pooling ASPP; The RGB branch processes the preprocessed forged image X to obtain the RGB rough features The noise branch processes the noise map X' to obtain the coarse noise features where C s represents the number of channels of the coarse features, and H s and W s represent the height and width of the features, respectively; Step 2.3: The edge reconstruction module includes: a Sobel layer, a context information enhancement graph CEG, and a differentiable binarization module, which are used to convert the RGB coarse features into edge features G e ; Step 2.4: Construct a regional message passing controller, and under the guidance of the edge features between the forged region and the real region, construct the RGB coarse features into the RGB features Z after graph reasoning r ; Step 2.5, the fusion module obtains the fusion feature using Equation (9) so as to perform bilinear upsampling operation on G o and then obtain the predicted mask of the forged image X G o = DA(G z , G n ) (9) In formula (9), DA represents the dual attention network, is Z r the chart feature after shaping; Step 3: Use Equation (10) to construct a Dice loss function L, and use an ADAM optimizer to train the image forgery localization network until the loss function L converges, so as to obtain a trained image forgery localization model for realizing the recognition and localization of forged images; In Equation (10), represents the Dice loss, is the true mask label of the forged image X, is the edge feature of the true mask label, is the edge feature downsampled by E, and λ1, λ2, λ3 are three weight factors, and λ1 + λ2 + λ3 = 1.
2. The image forgery localization method based on edge-aware regional message passing control according to claim 1, characterized in that, Step 2.3 includes: Step 2.3.1: The Sobel layer extracts the RGB rough feature G using Equation (1) r Edge-related features G c = G r ⊙σ(Norm(Sobel(G r )) (1) In Equation (1), Norm is the L2 normalization operation, σ is the Sigmoid function, Sobel is the Sobel convolution operation, and ⊙ represents the Hadamard product; Step 2.3.2, the context information enhanced graph CEG reshapes the edge-related feature G c into an edge-related graph representation N represents the number of nodes of the edge-related graph representation, and N = H s × W s ; Step 2.3.
3. The context information enhanced graph CEG obtains the adjacency matrix by using Equation (2). A c = Norm(Conv(Conv(G′ c ))) (2) In Equation (2), Conv represents a convolution operation with a convolution kernel of 1×1; Step 2.3.4, the context information enhanced graph CEG obtains the global feature map Global(G' c ) by using Equation (3): Global(G′ c ) = A c G′ c W c (3) In formula (3), represents the parameter to be learned in the context information enhanced graph CEG; Step 2.3.5, the context information enhanced graph CEG extracts edge-related feature G using two convolutional layers with a convolution kernel of 1×1 c local feature Local(G c ); Step 2.3.
6. The context information enhanced graph CEG obtains the edge probability graph by using Equation (4). G p = σ(Conv(Cat(Local(G c ), Global′(G′ c ))) (4) In Equation (4), Cat represents a concatenation operation; Step 2.3.7, the differentiable binarization module obtains the binarized edge features using Equation (5) where H e and W e represent the height and width of the edge features, respectively; In Equation (5), τ is a transformation function to be learned, and k is a scaling factor.
3. The image forgery localization method based on edge-aware regional message passing control according to claim 2, characterized in that, Step 2.4 includes: Step 2.4.1, determine the RGB rough feature G r for the i-th feature point P i and the j-th feature point P j in it respectively, and their relationships with the edge feature G e . If the two feature points are respectively inside and outside the edge feature G e , set the relationship value XN(P i , P j ) to 0. If the two feature points are both inside or both outside the edge feature G e , set the relationship value XN(P i , P j ) to 1, so as to obtain the relationship values of all feature points in the RGB rough feature G r and form a relationship matrix Let N' denote the dimension of the relationship matrix, and N' = H e ×W e ; Step 2.4.2: Calculate the attention coefficient α of the i-th feature point P i and the j-th feature point P j using Equation (6), so as to obtain the attention matrix; i,j α i,j = ψ(x i ) T ψ'(x j ) (6) In Equation (6), ψ and ψ' are linear transformation operations, and ψ = Wx i and ψ' = W'x j , where W and W' represent two parameter matrices to be learned, and T represents the transpose; Step 2.4.3: Use the softmax function to normalize the attention matrix to obtain a preliminary adjacency matrix Step 2.4.
4. Dynamically adjust the adjacency matrix A using Equation (7) r to obtain the adjacency matrix guided by edge features A′ r = A r ⊙A e (7) Step 2.4.5: Obtain the RGB features after graph reasoning using Equation (8). Z r = ReLU(A' r G' r W z ) (8) In Equation (8), ReLU represents the activation function, is the parameter matrix to be learned, is the reshaped graph representation of G r after reshaping.
4. An electronic device, comprising a memory and a processor, characterized in that, The memory is used to store a program that supports the processor to execute any one of the image forgery localization methods in claims 1-3, and the processor is configured to execute the program stored in the memory.
5. A computer-readable storage medium, on which a computer program is stored, characterized in that, When the computer program is run by the processor, it executes the steps of any one of the image forgery localization methods in claims 1-3.