Image tampering localization method and system based on mask-guided query encoder framework
By adopting a mask-guided query encoder framework in image tampering positioning, the problems of inefficient training and difficulty in utilizing spatial location and shape details in the prior art are solved, and more efficient image tampering positioning performance is achieved.
Patent Information
- Application Number
- CN202510154585.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art is inefficient in training in image tampering positioning and it is difficult to effectively utilize the spatial position and shape details of the manipulated area.
Using a mask-guided query encoder framework, through truth mask guidance, the query token (LQT) can be learned to identify the forged area, and a mask-guided training strategy is designed to improve the convergence speed and positioning performance of the model.
The training efficiency and positioning performance of the image tampering positioning model are significantly improved, so that the model can more effectively focus on the spatial position and shape details of the tampered area.
Smart Images

Figure CN120071105A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical fields of digital security and multimedia digital forensics, and in particular, to an image tampering localization method and system based on a mask-guided query encoder framework. Background Art
[0002] Deepfake is a technology that uses deep learning algorithms such as generative adversarial networks to create fake images or videos. The tampered pictures generated by deepfake technology are visually indistinguishable from real content. The false information created by image tampering can violate personal privacy, damage finances, and have an adverse impact on society.
[0003] Currently, deep learning-based models have made great progress in image tampering localization. However, the training efficiency of these models is low. Existing image processing and localization networks have two drawbacks, resulting in poor performance. First, these methods use a convolutional neural network (CNN) to classify pixel-by-pixel features in the final decoder process, which limits access to global information in the image. Second, these image processing and localization networks only use the ground truth mask label through the cross-entropy loss. Since the cross-entropy loss operates at the pixel level to evaluate whether each position estimate is correct and emphasizes the accuracy of each pixel, it cannot utilize the spatial position and shape details of the manipulated area and the network training is inefficient. How to solve the above two problems is an urgent problem for those skilled in the art.
[0004] Therefore, there is an urgent need for a targeted image tampering localization method and system based on a mask-guided query encoder framework. Summary of the Invention
[0005] The purpose of the present invention is to provide an image tampering localization method based on a mask-guided query encoder framework. Using the ground truth mask to guide the learnable query token (LQT) to identify the forged area and the proposed mask-guided training strategy have a significant impact on the convergence speed and localization performance of MGQFormer training.
[0006] In a first aspect, this application provides an image tampering localization method based on a mask-guided query encoder framework, and the method includes:
[0007] Adding noise to the input image, using the BayarConv layer to transform the features of the input image, amplifying the forged traces or other details in the image, and obtaining the noisy image;
[0008] Using a multi-branch feature extractor to respectively extract the features of the input image and the noise features of the noisy image;
[0009] Process the features of the input image and the noise features through spatial attention and channel attention respectively, and then combine the two processing results to obtain the multi-modal feature Z f ;
[0010] Perform noise addition processing on the ground truth mask; use a convolutional network to convert the noisy mask into a guiding query token GQT;
[0011] Randomly initialize the learnable query token LQT;
[0012] Construct a mask-guided query encoder framework, and input the multi-modal feature Z f , the guiding query token GQT, and the learnable query token LQT into the query encoder for mask-guided training, including:
[0013] In the attention mechanism, the learnable query token LQT interacts with the multi-modal feature Z f to interactively extract rich forgery information, and then obtain the contextualized image feature Z f * and LQT*, and GQT is used as a guide for interacting with other queries to obtain GQT*;
[0014] Use the obtained image feature Z f * to perform scalar product calculation and bilinear upsampling operations with the updated LQT* and GQT* respectively to obtain a predicted mask and an auxiliary mask;
[0015] Calculate the loss functions required in mask-guided training, including auxiliary loss, mask-guided loss, and predicted mask localization loss, forcing GQT to guide LQT to focus on the spatial location and shape details of the tampered area;
[0016] Add the auxiliary loss, mask-guided loss, and predicted mask localization loss according to a ratio, and use the added result as the total loss to backpropagate and train the image tampering localization model;
[0017] After mask-guided training is completed, use the trained image tampering localization model to discriminate the tampered image to be detected and obtain the tampering localization of the tampered image.
[0018] In a second aspect, the present application provides an image tampering localization system based on a mask-guided query encoder framework, and the system includes:
[0019] A noise addition module for adding noise to the input image, using the BayarConv layer to transform the features of the input image, amplify the forgery traces or other details in the image, and obtain a noisy image;
[0020] A feature extraction module for respectively extracting the features of the input image and the noise features of the noisy image by using a multi-branch feature extractor;
[0021] An attention module, which is used to process the features of the input image and the noise features through spatial attention and channel attention respectively, and then merge the two processing results to obtain the multi-modal feature Z f ;
[0022] A token module, which is used to add noise to the ground truth mask, convert the masked mask into a guided query token GQT by using a convolutional network, and randomly initialize the learnable query token LQT;
[0023] A training module, which is used to build a mask-guided query encoder framework, and input the multi-modal feature Z f , the guided query token GQT, and the learnable query token LQT into the query encoder for mask-guided training, including:
[0024] In the attention mechanism, the learnable query token LQT interacts with the multi-modal feature Z f to extract rich forgery information, and then obtain the contextualized image feature Z f * and LQT*, and GQT is used as a guide for interacting with other queries to obtain GQT*;
[0025] Using the obtained image feature Z f * to perform scalar product calculation and bilinear upsampling operations with the updated LQT* and GQT* respectively to obtain a predicted mask and an auxiliary mask;
[0026] Calculate the loss functions required in the mask-guided training, including the auxiliary loss, the mask-guided loss, and the predicted mask localization loss, forcing GQT to guide LQT to focus on the spatial position and shape details of the tampered area;
[0027] Add the auxiliary loss, the mask-guided loss, and the predicted mask localization loss according to a ratio, and use the added result as the total loss to backpropagate and train the image tampering localization model;
[0028] A detection module, which is used to use the trained image tampering localization model to discriminate the tampered image to be detected and obtain the tampering localization of the tampered image after the mask-guided training is completed.
[0029] In a third aspect, the present application provides an image tampering localization system based on a mask-guided query encoder framework, and the system includes a processor and a memory:
[0030] The memory is used to store program codes and transmit the program codes to the processor;
[0031] The processor is used to execute the method according to any one of the four possible aspects in the first aspect according to the instructions in the program code.
[0032] In a fourth aspect, the present application provides a computer-readable storage medium for storing program code, which is used to be executed by a processor to implement the method described in any one of the four possibilities of the first aspect.
[0033] Beneficial effects
[0034] The present invention provides an image forgery localization method and system based on a mask-guided query encoder framework, which uses a ground-truth mask to guide a learnable query token (LQT) to identify forged regions, including: extracting feature embeddings of the ground-truth mask as guidance for query token (GQT) operations; constructing a mask-guided query encoder framework, and then inputting the GQT and LQT into the query encoder respectively to localize forged regions; designing a mask-guided loss algorithm to utilize the query encoder to learn the position and shape information in the ground-truth mask label, thereby reducing the feature distance between the GQT and LQT; finally, using the trained model to perform forgery localization on forged images, overcoming the problems that the prior art cannot utilize the spatial position and shape details of the manipulated region and the network training is inefficient.
[0035] The method and system of the present invention have the following advantages and effects:
[0036] (1) The present invention introduces a mask-guided query encoder framework, which includes a query-based encoder that uses learnable query tokens (LQT) to localize the manipulated region. It solves the problem that the prior method using a convolutional neural network (CNN) in the encoder process restricts access to global information in the image, making the network more interpretable and effectively utilizing the attention mechanism of the encoder.
[0037] (2) The present invention proposes a mask-guided training method that uses a guidance query token (GQT) extracted from the ground-truth mask as guidance to improve the LQT, guiding the network to focus on the forged region, thereby making the training process efficient.
[0038] (3) The present invention designs a mask-guided loss, expecting the LQT to be more similar to the GQT to make the prediction more accurate, forcing the GQT to guide the LQT to focus on the spatial position and shape details of the manipulated region. Brief description of the drawings
[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.
[0040] Figure 1 It is a flowchart of the method of the present invention;
[0041] Figure 2 This is the system architecture diagram of the present invention. Detailed implementation manners
[0042] The preferred embodiments of the present invention will be described in detail below in conjunction with the accompanying drawings, so that the advantages and features of the present invention can be more easily understood by those skilled in the art, thereby making a clearer and more definite definition of the protection scope of the present invention.
[0043] The image forgery localization method based on the mask-guided query encoder framework provided by this application, as Figure 1 shown, the method includes:
[0044] Add noise to the input image, and use the BayarConv layer to transform the features of the input image to amplify the forgery traces or other details in the image, obtaining the noisy image;
[0045] Use a multi-branch feature extractor to extract the features of the input image and the noise features of the noisy image respectively;
[0046] Process the features of the input image and the noise features through spatial attention and channel attention respectively, and then combine the two processing results to obtain the multi-modal feature Z f ;
[0047] Add noise to the ground truth mask; use a convolutional network to convert the noisy mask into a guiding query token GQT;
[0048] Randomly initialize the learnable query token LQT;
[0049] Construct a mask-guided query encoder framework, and input the multi-modal feature Z f , the guiding query token GQT, and the learnable query token LQT into the query encoder for mask-guided training, including:
[0050] In the attention mechanism, the learnable query token LQT interacts with the multi-modal feature Z f to extract rich forgery information, and then obtain the contextualized image feature Z f * and LQT*, and GQT is used as a guide for interacting with other queries to obtain GQT*;
[0051] Use the obtained image feature Z f * to perform scalar product calculation and bilinear upsampling operations with the updated LQT* and GQT* respectively to obtain the predicted mask and the auxiliary mask;
[0052] Calculate the loss functions required in the masked guidance training, including the auxiliary loss, the masked guidance loss, and the predicted mask localization loss, to force the GQT to guide the LQT to focus on the spatial location and shape details of the tampered region;
[0053] Add the auxiliary loss, the masked guidance loss, and the predicted mask localization loss in proportion, and use the sum as the total loss to backpropagate and train the image tampering localization model;
[0054] After the masked guidance training is completed, use the trained image tampering localization model to discriminate the tampered images to be detected and obtain the tampering localization of the tampered images.
[0055] In some preferred embodiments, adding noise to the input image and using a multi-branch feature extractor to extract the features of the input image and the noise features of the noise-added image specifically include:
[0056] First, use BayarConv to add noise to the input image X and extract its noise feature X n ∈R HxWx3 , where H represents the height of the input image X and W represents the width of the input image X; then use X and X n as two branches, divide them into patches of size P respectively, and reshape the patches into X p ∈R NxD , where N = HW / P 2 is the number of patches, and D is the dimension of the embedding; then embed the learnable embedding position pos ∈ R NxD into the image to generate the sequence token Z = X p + pos, and then process these tokens through the LTransformer layer; connect the outputs of the two branches to obtain Zc ∈ R Nx2D .
[0057] In some preferred embodiments, processing the features of the input image and the noise features through spatial attention and channel attention respectively, and then combining the two processing results to obtain the multi-modal feature Z f , specifically include:
[0058] First, reshape Z c ∈R Nx2D to Z m ∈R hxwxc , where h = H / P, w = W / P, and c = D; then transpose Z m to V = proj(Z m ) ∈ R hwxc , K = proj(Z m ) ∈ R hwxc , Q = transpose(proj(Zm )) ∈ R cxhw , where each proj is an independent projection layer, including a 1x1 convolutional layer and a reshaping operation; then the channel attention processing is performed as follows:
[0059] CAM(Z m ) = proj(V(softmax(QK)));
[0060] At the same time, the spatial attention processing is performed as follows:
[0061] SAM(Z m ) = proj(softmax(Q T K T )V);
[0062] Subsequently, the contextualized image feature Z f, , which is the multimodal feature Z f , is obtained as follows:
[0063] Z f = CAM(Z m ) + SAM(Z m ) + Z m .
[0064] In some preferred embodiments, the use of the convolutional network to convert the noise-added mask into a guiding query token GQT is specifically as follows:
[0065] Noise is added to the ground truth mask, and point noise is applied to the ground truth mask to increase robustness; points within the ground truth mask are randomly selected, and the original values are inverted to represent different regions. A parameter μ is used to represent the area noise percentage, and the number of noise points is μ·HW;
[0066] The convolutional network is used to convert the noise-added mask into GQT to preserve the spatial information in the mask, and the ground truth mask G ∈ R HxW is converted to GQT ∈ R 2xN .
[0067] In some preferred embodiments, the use of the obtained image feature Z f * is respectively subjected to scalar product calculation and bilinear upsampling operation with the updated LQT* and GQT* to obtain a predicted mask and an auxiliary mask, specifically as follows:
[0068] Using the obtained image feature Z f *, the predicted mask and the auxiliary mask are calculated respectively with LQT* and GQT*. The calculation process of the predicted mask is as follows:
[0069] M* = norm(proj(Z f*))*(norm(proj(LQT*)) T ;
[0070] where proj is a linear layer, norm represents L2 normalization, and for the image feature Z f *, perform a scalar product with LQT* to obtain M* ∈ R Nx2 ;
[0071] Further reshape the sequence M* into M** ∈ R hxwx2 , and calculate M** as follows:
[0072] M = upsample(norm(softmax(M**)));
[0073] where M ∈ R HxW is the predicted mask, upsample is a bilinear upsampling operation to adjust the mask to the same size as the input image;
[0074] For the image feature Z f *, perform the same calculation process as above with GQT* to obtain the auxiliary mask M aux .
[0075] Specifically, an embodiment can be introduced as follows.
[0076] An image forgery localization method based on a mask-guided query-based Transformer framework proposed in this embodiment includes the following steps:
[0077] S1. For the input image X ∈ R HxWx3 , add noise to it, and respectively extract the feature and noise feature of the original image and the noise-added image using the decoder in the Transformer framework, and then fuse the multi-modal features through the spatial and channel attention module (SCAM); adding noise to the input image means using the BayarConv layer to transform the features of the image to reveal or amplify the forgery traces or other details in the image; extracting the feature and noise feature using the decoder in the Transformer framework means inputting the original image and the noise-added image into a two-branch Transformer encoder to fully utilize the information from both domains;
[0078] In step S1, adding noise to the input image X ∈ R HxWx3 , extracting the feature and noise feature of the original image and the noise-added image respectively using the decoder in the Transformer framework, and then fusing the multi-modal features through the spatial and channel attention module (SCAM) is specifically as follows:
[0079] S11. BayarConv adds noise to the input image X to obtain X n ∈R HxWx3 ;
[0080] Preferably, H = W = 384;
[0081] S12. Send the input image X and the noise map X n to the dual-branch Transformer encoder. Specifically, divide X and X n into patches of size P, and reshape the patches into X p ∈R NxD ;
[0082] Preferably, the Transformer encoder is initialized with the weights of the ImageNet-pre-trained ViT model (Steiner et al 2021), and the model has 12 layers;
[0083] Preferably, N = HW / P 2 where P is the patch size and D is the dimension of the embedding;
[0084] Furthermore, add the learnable positional embedding pos ∈ R NxD to the image embedding to generate the sequence tokens Z = X p + pos, and then process these tokens through the LTransformer layer. The same above processing is also performed on the noise branch;
[0085] Furthermore, connect the outputs of the two branches to obtain Zc ∈ R Nx2D ;
[0086] S13. Fuse the multi-modal features from the output of the dual-branch Transformer encoder through the spatial and channel attention module (SCAM). Specifically
[0087] Reshape Z c ∈R Nx2D to Z m ∈R hxwxc ;
[0088] Preferably, h = H / P, w = W / P, c = D;
[0089] Furthermore, transpose Z m to V = proj(Z m ) ∈ R hwxc , K = proj(Z m ) ∈ R hwxc , Q = transpose(proj(Z m )) ∈ R cxhw ;
[0090] Preferably, each proj is an independent projection layer, including a 1x1 convolutional layer and a reshaping operation;
[0091] Furthermore, the channel attention module and the spatial attention module are executed as follows:
[0092] CAM(Z m ) = proj(V(softmax(QK)))
[0093] SAM(Z m ) = proj(softmax(Q T K T )V)
[0094] Furthermore, the contextualized image feature Z f , can be obtained as follows:
[0095] Z f = CAM(Z m ) + SAM(Z m ) + Z m
[0096] S2. Generate a guiding query token GQT; construct a mask-guided query-based Transformer framework, input the fused feature Zf, GQT, and LQT into the proposed Transformer decoder to obtain the contextualized image feature Zf*, the updated LQT, and the updated GQT; perform mask-guided training; the generation of the guiding query token GQT means using a convolutional network to convert the noisy ground truth mask into the guiding query token GQT; the Transformer decoder, preferably, is initialized with random weights from a truncated normal distribution and has 6 layers; the mask-guided training means using the guiding query token GQT to guide the update of LQT and forcing LQT to focus on the position and shape of the forgery area;
[0097] In step S2, the generation of the guiding query token GQT; constructing a mask-guided query-based Transformer framework, inputting the fused feature Zf, GQT, and LQT into the proposed Transformer decoder to obtain the contextualized image feature Zf*, the updated LQT, and the updated GQT; performing mask-guided training is specifically as follows:
[0098] S21. Add noise to the ground truth mask using point noise, randomly select points within the ground truth mask, and invert the original values to represent different regions;
[0099] Furthermore, use a convolutional network to convert the noisy mask into GQT;
[0100] Preferably, G ∈ R HxW Convert to GQT ∈ R 2xN ;
[0101] S22. Adopt randomly initialized learnable query token LQT;
[0102] Furthermore, GQT together with image feature Z f and LQT are input into the Transformer decoder for masked-guided training, where GQT serves as a guidance for interacting with other queries and helps the decoder update LQT;
[0103] Furthermore, in the attention mechanism, LQT interacts with image feature Z f to extract rich forgery information, and then the contextualized image feature Z f * and LQT* are obtained, and GQT serves as a guidance for interacting with other queries to obtain GQT*.
[0104] See Figure 2 It can be seen that step S2 completes the generation of the guiding query token GQT, the initialization of the learnable query token LQT, and the structure of the Transformer decoder in the process of constructing the masked-guided Transformer decoder. Mask prediction and loss calculation are completed in steps S3 and S4.
[0105] S3. Use the obtained image feature Zf* to perform a series of calculations such as scalar product with the updated LQT and GQT respectively to obtain the predicted mask and the auxiliary mask, specifically:
[0106] S31. Use the obtained image feature Z f * to calculate the predicted mask with LQT*, and the mask calculation is as follows:
[0107] M* = norm(proj(Z f *)) * (norm(proj(LQT*)) T
[0108] Preferably, proj is a linear layer, and norm represents L2 normalization. After performing the scalar product of the image feature Z f * and LQT*, M* ∈ R Nx2 .
[0109] Furthermore, reshape the sequence M* to M** ∈ R hxwx2 , and calculate M** as follows:
[0110] M = upsample(norm(softmax(M**)))
[0111] Preferably, M ∈ RHxW is the prediction mask, and upsample is a bilinear upsampling operation to adjust the mask to the same size as the input image.
[0112] S32. Image feature Z f * performs the same calculation as in step S31 with GQT* to obtain the auxiliary mask M aux .
[0113] S4. Calculate the loss functions required in mask-guided training, including auxiliary loss, mask-guided loss, and prediction mask localization loss;
[0114] Optionally, for the auxiliary loss, use pixel-level cross-entropy loss to make the auxiliary mask more accurate, specifically:
[0115]
[0116] where G ∈ R HxW is the ground truth mask, not the masked one after adding noise;
[0117] Optionally, for the mask-guided loss, use cosine similarity loss to reduce the distance between GQT and LQT, making LQT similar to GQT, specifically:
[0118] L guide = 1 - cos(LQT*, GQT*)
[0119] where cos represents calculating cosine similarity;
[0120] Optionally, for the prediction mask M localization loss L loc , it uses the same cross-entropy loss as the auxiliary loss;
[0121] Furthermore, calculate the total loss function, which consists of three parts: the auxiliary loss that makes M aux accurate, the mask-guided loss that makes LQT* and GQT* closer, and the prediction mask M localization loss, as follows:
[0122] L = L loc + L aux + λL guide
[0123] Preferably, λ is set to 0.5 during training.
[0124] S5. Use the trained localization model to locate the tampered position of the image to be tampered with and obtain the localization performance.
[0125] In a more specific embodiment, the steps are as follows:
[0126] S1: Preprocessing: Use the BayarConv layer to process the input image \(X\in\mathbb{R}\) HxWx3 Add noise and use the decoder of the dual-branch Transformer framework to extract the feature and noise feature of the original image and the noisy image, and then fuse the multi-modal features through the spatial and channel attention module (SCAM) to obtain the image feature token \(Z\). f .
[0127] S2: GQT generation and LQT initialization: Use the convolutional network to convert the noisy ground truth mask into the guiding query token GQT, and randomly initialize the learnable query token LQT.
[0128] S3: Generate the prediction mask and the auxiliary mask: Input the image feature token \(Z\) f , GQT and LQT into the Transformer decoder together. GQT serves as a guide for interacting with other queries and helps the decoder update LQT. In the attention mechanism, LQT interacts with the image feature \(Z\) f to extract rich forgery information, and then obtain the contextualized image feature \(Z\) f * and LQT*, and GQT serves as a guide for interacting with other queries to obtain GQT*. Calculate the scalar product and reshaping operations of the obtained image feature \(Z_f^*\) with the updated LQT and GQT respectively to obtain the prediction mask and the auxiliary mask.
[0129] S4: Mask-guided training and testing: Calculate the pixel-level cross-entropy loss between GQT and the auxiliary mask as the auxiliary loss, calculate the cosine similarity loss between GQT and LQT as the mask-guided loss, and calculate the pixel-level cross-entropy loss between GQT and the prediction mask as the localization loss. Add the auxiliary loss, the mask-guided loss and the localization loss proportionally as the total loss and backpropagate to train the model. After training, input the test set images into the trained model to localize the tampered parts of the images.
[0130] Figure 2 This is the architecture diagram of the image tampering localization system based on the mask-guided query encoder framework provided by this application. The system includes:
[0131] Noise addition module, used to add noise to the input image, use the BayarConv layer to transform the features of the input image, amplify the forgery traces or other details in the image, and obtain the noisy image;
[0132] Feature extraction module, used to extract the features of the input image and the noise features of the noisy image respectively by using the multi-branch feature extractor;
[0133] An attention module, which is used to process the features of the input image and the noise features through spatial attention and channel attention respectively, and then merge the two processing results to obtain multimodal feature Z f ;
[0134] A token module, which is used to add noise to the ground truth mask, convert the masked noise into a guiding query token GQT by using a convolutional network, and randomly initialize a learnable query token LQT;
[0135] A training module, which is used to build a mask-guided query encoder framework, and input the multimodal feature Z f , the guiding query token GQT, and the learnable query token LQT into the query encoder for mask-guided training, including:
[0136] In the attention mechanism, the learnable query token LQT interacts with the multimodal feature Z f to extract rich forgery information, and then obtain the contextualized image feature Z f * and LQT*, and GQT is used as a guide for interacting with other queries to obtain GQT*;
[0137] Using the obtained image feature Z f * to perform scalar product calculation and bilinear upsampling operations with the updated LQT* and GQT* respectively to obtain a predicted mask and an auxiliary mask;
[0138] Calculate the loss functions required in mask-guided training, including auxiliary loss, mask-guided loss, and predicted mask localization loss, forcing GQT to guide LQT to focus on the spatial location and shape details of the tampered area;
[0139] Add the auxiliary loss, mask-guided loss, and predicted mask localization loss according to a ratio, and use the added result as the total loss to backpropagate and train the image tampering localization model;
[0140] A detection module, which is used to use the trained image tampering localization model to discriminate the tampered image to be detected and obtain the tampering localization of the tampered image after the mask-guided training is completed.
[0141] The present application provides an image tampering localization system based on a mask-guided query encoder framework. The system includes: the system includes a processor and a memory:
[0142] The memory is used to store program code and transmit the program code to the processor;
[0143] The processor is used to execute the method described in any one of all the embodiments of the first aspect according to the instructions in the program code.
[0144] The present application provides a computer-readable storage medium for storing program codes, which are used to be executed by a processor to implement the method described in any one of all the embodiments of the first aspect.
[0145] In specific implementation, the present invention further provides a computer storage medium, wherein the computer storage medium can store a program, and when the program is executed, it may include some or all of the steps in various embodiments of the present invention. The storage medium may be a magnetic disk, an optical disk, a read-only memory (abbreviation: ROM) or a random access memory (abbreviation: RAM), etc.
[0146] Those skilled in the art can clearly understand that the technologies in the embodiments of the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions in the embodiments of the present invention, in essence or the part that makes contributions to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which may be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present invention.
[0147] For the same or similar parts among the various embodiments of this specification, reference can be made to each other. In particular, for the embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the description in the method embodiments.
[0148] The above-described embodiments of the present invention do not constitute a limitation on the protection scope of the present invention.
Claims
1. A method for locating image tampering based on a mask-guided query encoder framework, characterized in that: The method comprises: Add noise to the input image, use the BayarConv layer to transform the features of the input image, amplify the forgery traces or other details in the image, and obtain the noisy image; A multi-branch feature extractor is used to extract the features of the input image and the noise features of the noisy image respectively; The input image features and the noise features are processed respectively by spatial attention and channel attention, and then the two processing results are combined to obtain the multimodal feature Z f ; Noise the truth mask; use a convolutional network to convert the noisy mask into a guided query token GQT; Randomly initialize the learnable query token LQT; Construct a mask-guided query encoder framework to transform the multimodal feature Z f , a guided query token GQT, and a learnable query token LQT are input into the query encoder to perform mask guided training, including: In the attention mechanism, the query token LQT and the multimodal feature Z can be learned f Interactively extract rich forgery information and then obtain contextualized image features Z f * and LQT*, GQT is used as a guide to interact with other queries to obtain GQT*; Using the obtained image feature Z f * Perform scalar product calculation and bilinear upsampling operations with the updated LQT* and GQT* respectively to obtain the prediction mask and auxiliary mask; Calculate the loss functions required in mask-guided training, including auxiliary loss, mask-guided loss, and predicted mask positioning loss, forcing GQT to guide LQT to focus on the spatial position and shape details of the tampered area; The auxiliary loss, mask guidance loss and predicted mask positioning loss are added in proportion, and the result is used as the total loss to train the image tampering positioning model through back propagation. After the mask-guided training is completed, the trained image tampering localization model is used to identify the tampered image to be detected and obtain the tampering location of the tampered image.
2. The method according to claim 1, characterized in that: The step of adding noise to the input image and using a multi-branch feature extractor to respectively extract features of the input image and noise features of the image after adding noise is specifically as follows: First, BayarConv is used to perform noise processing on the input image X and extract its noise feature X n ∈R HxWx3 , where H represents the height of the input image X and W represents the width of the input image X; then X and X n As two branches, they are divided into patches of size P, and the patches are reshaped into X p ∈R NxD , where N = HW / P 2 is the number of patches, D is the dimension of embedding; then the learnable embedding position pos∈R NxD Embedded into the image, generating a sequence tag Z = X p +pos, and then process these tags through the LTransformer layer; concatenate the outputs of the two branches to obtain Zc∈R Nx2D .
3. The method according to claim 2, characterized in that: The input image feature and the noise feature are processed respectively by spatial attention and channel attention, and then the two processing results are combined to obtain a multimodal feature Z f , specifically: First reshape Z c ∈R Nx2D Z m ∈R hxwxc , where h = H / P, w = W / P, c = D; then Z m Transpose V = proj(Z m )∈R hwxc , K = proj(Z m )∈R hwxc , Q = transpose(proj(Z m ))∈R cxhw , where each proj is an independent projection layer, including a 1x 1 convolution layer and a reshape operation; then the channel attention processing is performed as follows: CAM(Z m )=proj(V(softmax(QK))); At the same time, spatial attention processing is performed as follows: SAM( Z m )=proj(softmax(Q T K T )V): Then we get the contextualized image feature Z f, , which is the multimodal feature Z f , as follows: WITH f =CAM(Z m )+MYSELF(WITH m )+Z m 。 4. The method according to claim 3, characterized in that: The convolutional network is used to convert the noisy mask into a guided query token GQT, specifically: Noise is added to the true value mask. Point noise is applied to the true value mask to increase robustness. Points in the true value mask are randomly selected, and the original values are inverted to represent different areas. A parameter μ is used to represent the area noise percentage, and the number of noise points is μ·HW. The noisy mask is converted to GQT using a convolutional network to preserve the spatial information in the mask and the true value mask G∈R HxW Transformed to GQT∈R 2xN .
5. The method according to claim 4, characterized in that: The image feature Z obtained by using f * respectively performs scalar product calculation and bilinear upsampling operation with the updated LQT* and GQT* to obtain the prediction mask and auxiliary mask, specifically: Using the obtained image feature Z f * is calculated with LQT* and GQT* to obtain the prediction mask and auxiliary mask. The calculation process of the prediction mask is as follows: M*=norm(proj(Z f *))*(norm(proj(LQT*)) T ; Where proj is a linear layer, norm represents L2 normalization, and f *Perform a scalar product with LQT* to obtain M*∈R Nx2 ; Further reshape the sequence M* into M**∈R hxwx2 , and M** is calculated as follows: M=upsample(norm(softmax(M**))); where M∈R HxW To predict the mask, upsample is a bilinear upsampling operation that adjusts the mask to the same size as the input image; Image feature Z f *Perform the same calculation process as GQT* to obtain the auxiliary mask M aux .
6. An image tampering localization system based on a mask-guided query encoder framework, characterized in that: The system comprises: The noise adding module is used to add noise to the input image, transform the features of the input image using the BayarConv layer, amplify the forgery traces or other details in the image, and obtain the noisy image; A feature extraction module, used to respectively extract features of the input image and noise features of the noisy image using a multi-branch feature extractor; An attention module is used to process the features of the input image and the noise features respectively through spatial attention and channel attention, and then merge the two processing results to obtain a multimodal feature Z f ; The token module is used to add noise to the truth mask and convert the noisy mask into a guided query token GQT using a convolutional network, and randomly initialize the learnable query token LQT; The training module is used to construct a mask-guided query encoder framework to transform the multimodal feature Z f , a guided query token GQT, and a learnable query token LQT are input into the query encoder to perform mask guided training, including: In the attention mechanism, the query token LQT and the multimodal feature Z can be learned f Interactively extract rich forgery information and then obtain contextualized image features Z f * and LQT*, GQT is used as a guide to interact with other queries to obtain GQT*; Using the obtained image feature Z f * Perform scalar product calculation and bilinear upsampling operations with the updated LQT* and GQT* respectively to obtain the prediction mask and auxiliary mask; Calculate the loss functions required in mask-guided training, including auxiliary loss, mask-guided loss, and predicted mask positioning loss, forcing GQT to guide LQT to focus on the spatial position and shape details of the tampered area; The auxiliary loss, mask guidance loss and predicted mask positioning loss are added in proportion, and the result is used as the total loss to train the image tampering positioning model through back propagation. The detection module is used to identify the tampered images to be detected after the mask-guided training is completed, using the trained image tampering localization model to obtain the tampering location of the tampered images.
7. An image tampering localization system based on a mask-guided query encoder framework, characterized in that: The system comprises a processor and a memory: The memory is used to store program code and transmit the program code to the processor; The processor is used to execute the method according to any one of claims 1 to 5 according to the instructions in the program code.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium is used to store program codes, and the program codes are used to be executed by a processor to implement the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Small tampering region positioning method based on multi-information guidance and progressive mask Transform
CN117876704A