A method for image forgery detection and localization based on noise-assisted cue learning

Through noise-assisted cue learning and forgery-enhanced noise adapter, combined with a contrastive language image pre-training model, the generalization and false alarm rate problems of image forgery detection and localization methods are solved, achieving higher accuracy and generalization.

CN119295383BActive Publication Date: 2025-09-30UNIV OF SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411253779.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-09
Publication Date
2025-09-30
Estimated Expiration
2044-09-09

AI Technical Summary

Technical Problem

Existing image forgery detection and localization methods are difficult to generalize to forgery types outside the training set and have a high false alarm rate.

Method used

A noise-assisted cue learning method is adopted, which uses a contrastive language image pre-training model and an instance-aware dual-path cue learning module, combined with a forgery-enhanced noise adapter. By constructing a forgery identification network, image forgery detection and localization are performed, reducing the false alarm rate and improving generalization.

Benefits of technology

It improves the accuracy and generalization of image forgery detection and localization, reduces the false alarm rate, and outperforms existing methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119295383B_ABST
    Figure CN119295383B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for detecting and locating image forgeries based on noise-assisted cue learning. The method introduces a comparative language image pre-training model CLIP and cue learning into the field of image forgery detection and positioning, and comprises the following steps: 1. using CLIP's image encoder and visual transformer to process the input image and its noise respectively to obtain dual-branch features, and enhancing the network's forgery perception capability through multi-domain fusion and memory mechanisms; 2. creating a pair of learnable cues as negative-positive samples to replace discrete cues, and fine-tuning these cues according to the features and categories of each image; 3. completing forgery detection by calculating the similarity between the cues and the image features, and simultaneously integrating the image and text features through feature linear modulation to obtain a forgery positioning result. The present invention can improve the accuracy and generalization of image forgery detection and positioning, and can reduce the problem of false alarms, thereby improving the performance of image forgery detection and positioning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image forgery positioning, and in particular to an image forgery detection and positioning method based on noise-assisted prompt learning. Background Art

[0002] Astonishing advances in media technology and editing tools have made it increasingly easy to manipulate images. These forged images pose risks in various fields, such as copyright watermark removal, fake news generation, and fabricated evidence in court. Therefore, the task of image forgery detection and localization has attracted widespread public attention. It aims to check whether an image has been modified and identify the modified areas in the image. With the development of technologies such as diffusion generation, image forgery technology is rapidly updating, which makes image forgery detection and localization constantly face new forgeries. At the same time, false positives for real images will affect media dissemination and have a negative impact. Therefore, it is crucial to develop accurate, generalizable, and false positive-free image forgery detection and localization methods.

[0003] Most previous methods directly learn the mapping from images to forgery masks in an end-to-end manner, which makes them difficult to generalize to forgery types outside the training set. Furthermore, most methods use the predicted forgery localization mask for forgery detection, which often leads to detection results that are biased towards forgeries, resulting in a high false alarm rate. Summary of the Invention

[0004] In order to overcome the shortcomings of the existing technology, the present invention proposes an image forgery detection and positioning method based on noise-assisted prompt learning, in order to improve the accuracy and generalization of image forgery detection and positioning, and reduce the false alarm problem, thereby improving the performance of image forgery detection and positioning.

[0005] In order to achieve the above-mentioned object, the present invention adopts the following technical solutions:

[0006] The image forgery detection and positioning method based on noise-assisted prompt learning of the present invention is characterized in that it is performed according to the following steps:

[0007] Step 1: Get the real image X a and the forged image X f , set X respectively a The vector to be learned V a and X f The vector to be learned V f ;Will and Any vector to be learned is denoted as V i ,i∈{a,f}; and Any image in is recorded as the input image X to be detected i ∈R H×W×C, where C represents the number of channels of the input image, H and W represent the height and width of the input image respectively;

[0008] Step 2: Construct a forgery detection network, including: a language image comparison pre-training model, a Bayar convolutional layer, a small visual transformer model, a cross attention layer, a memory mechanism layer, an instance-aware dual-path prompt learning module, and a prediction module. i and V i Processing is performed to obtain the image-text similarity ρ and the predicted forged positioning map G out ;

[0009] Step 3: Construct the total loss function of the authentication network :

[0010] (9)

[0011] In formula (9), are two equilibrium parameters, Forged images The true mask label of represents the Dice loss, represents the binary cross entropy loss; and:

[0012] (10)

[0013] In formula (10), is the input image X i If X i is a real image, then =1, otherwise, =0;

[0014] Step 4: Use the gradient descent method to train the authentication network and calculate the total loss function. To update the network parameters until the total loss function Until convergence, the optimal authentication network model is obtained.

[0015] The image forgery detection and positioning method based on noise-assisted cue learning according to the present invention is also characterized in that step 2 is performed as follows:

[0016] Step 2.1: Input image X i The N-layer image encoder frozen in the contrast language image pre-training model is processed, and the corresponding encoding features of each layer output are recorded as ;in, Represents the encoded features output by the n-th layer image encoder; n∈[1,N];

[0017] Step 2.2, Bayar convolution layer extracts the input image X i The noise G noise After that, it is input into a small N-layer visual transformer model for processing, and the noise characteristics of each layer output are recorded as ;in, Represents the noise characteristics of the n-th layer output in a small visual transformer model;

[0018] Step 2.3: The cross attention layer is fused using formula (1) and , get the n-th layer fusion feature :

[0019] (1)

[0020] In formula (1), 、 、 Represents three weight matrices to be learned; represents the activation function, T represents transposition;

[0021] Step 2.4: The memory mechanism layer uses formula (2) to Connect layers to get the enhanced features of the nth layer :

[0022] (2)

[0023] In formula (2), represents a linear layer, Indicates a concatenation operation; when season = ;

[0024] Step 2.5, The input is processed into the image encoder of the contrast language image pre-training model, so as to obtain the new encoding features of the nth layer using formula (3) and generate Category token , n∈[1,N];

[0025] (3)

[0026] In formula (3), Represents the parameter matrix to be learned in the nth layer;

[0027] Step 2.6: The instance-aware two-way prompt learning module processes the vectors to be learned Vi, i∈{a,f} of the real image and the forged image to obtain the image-text similarity ρ;

[0028] Step 2.7: The prediction module uses formula (7) to obtain the predicted forged positioning map :

[0029] (7)

[0030] In formula (7), Describes the decoder, Features linear modulation operation, and has:

[0031] (8)

[0032] In formula (8), and represents 2 different linear layers, It is Hadamard.

[0033] Furthermore, step 2.6 is performed as follows:

[0034] Step 2.6.1: Use formula (4) to Make a rough adjustment and get the prompt after the rough adjustment ,in, indicates the corresponding hint of the real image, Indicates the corresponding hint of the forged image;

[0035] (4)

[0036] In formula (4), PANet represents the prompt adjustment network, which is a bottleneck structure consisting of a linear layer, a ReLU activation function, and a linear layer;

[0037] Step 2.6.2: Use formula (5) to get the adjusted text embedding ,in, represents the text embedding corresponding to the real image, represents the text embedding corresponding to the forged image;

[0038] (5)

[0039] In formula (5), For frozen text encoders, is a parameter to be learned, and EANet represents the adjustment network based on the transformer decoder;

[0040] Step 2.6.3: Use formula (6) to get the image-text similarity ρ:

[0041] (6)

[0042] In formula (6), Represents cosine similarity.

[0043] The electronic device of the present invention includes a memory and a processor, wherein the memory is used to store a program that supports the processor to execute the image forgery detection and positioning method, and the processor is configured to execute the program stored in the memory.

[0044] The present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program executes the steps of the image forgery detection and positioning method when the computer program is executed by a processor.

[0045] Compared with the prior art, the present invention has the following beneficial effects:

[0046] 1. This paper designs a novel image forgery detection and localization method that leverages the intrinsic perception capabilities of a contrastive language-image pre-trained model. It not only uses the contrastive language-image pre-trained model to extract discriminative features, but also uses prompt learning to promote network optimization, thereby improving the generalization and accuracy of image forgery detection and localization.

[0047] 2. This paper proposes an instance-aware two-way cue learning to adaptively find accurate cues for each image to describe the real-forged attribute, thereby better utilizing the potential of the contrastive language-image pre-training model to improve the generalization of the model and reduce the false alarm rate.

[0048] 3. The present invention designs a forgery-enhancing noise adapter to enhance the network's learning ability for forged information, while avoiding the catastrophic forgetting of the CLIP prior caused by large-scale fine-tuning, thereby further ensuring the generalization of the forgery detection method.

[0049] 4. We conduct extensive experiments on several representative benchmarks and demonstrate that our method outperforms state-of-the-art image forgery detection and localization methods in terms of accuracy and generalization.

[0050] 5. The present invention applies prompt learning and contrastive language-image pre-training models for image forgery detection and localization for the first time. The text encoder and image encoder of the contrastive language-image pre-training model are enhanced through instance-aware two-way prompt learning and a forgery-enhanced noise adapter. This overcomes the difficulties of directly applying the contrastive language-image pre-training model in the field of image forgery detection and localization, reveals the potential of visual language priors in this field, and thus achieves accurate and well-generalized image forgery detection and localization, effectively avoiding false alarms for real images. It also solves the problem that image forgery detection and localization technologies are difficult to cope with out-of-distribution detection and are prone to false alarms. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 This is a flow chart of an image forgery detection and positioning method based on noise-assisted prompt learning according to the present invention;

[0052] Figure 2 The figure is a comparison of the present invention with various methods. DETAILED DESCRIPTION

[0053] In this embodiment, a method for detecting and locating image forgeries based on noise-assisted cue learning is mainly to enhance the text encoder and image encoder of the contrastive language-image pre-training model through instance-aware two-way cue learning and a forgery-enhanced noise adapter; for the text encoder, the present invention first establishes a learnable cue pair as a positive and negative sample pair to replace the discrete cue, and then performs a dual learnable adjustment on the cue according to the features and categories of each image; on this basis, the present invention updates the cue pair and simultaneously performs forgery detection by constraining the text-image similarity between the cue and the corresponding image in the latent space of the contrastive language-image pre-training model; for the image encoder, the present invention designs a forgery-enhanced noise adapter, which further enhances the perception ability of the contrastive language-image pre-training model for forgeries through multi-domain fusion, zero linear layer and memory mechanism; through the mutual promotion of the two, the present invention achieves accurate and generalized image forgery detection and positioning, and effectively avoids false alarms for real images; such as Figure 1 As shown, the method is carried out in the following steps:

[0054] Step 1: Get the real image X a and the forged image X f , set X respectively a The vector to be learned V a and X f The vector to be learned V f ;Will and Any vector to be learned is denoted as V i ,i∈{a,f}; and Any image in is recorded as the input image X to be detected i ∈R H×W×C , where C represents the number of channels of the input image, H and W represent the height and width of the input image respectively;

[0055] Step 2: Construct a forgery detection network, including: a language image comparison pre-training model, a Bayar convolutional layer, a small visual transformer model, a cross attention layer, a memory mechanism layer, an instance-aware dual-path prompt learning module, and a prediction module. i and V i Processing is performed to obtain the image-text similarity ρ and the predicted forged positioning map G out ;

[0056] Step 2.1: Input image X i The N-layer image encoder frozen in the contrast language image pre-training model is processed, and the corresponding encoding features of each layer output are recorded as ;in, Represents the encoded features output by the n-th layer image encoder; n∈[1,N]; usually, N can be set to 12;

[0057] Step 2.2, Bayar convolution layer extracts the input image X i The noise G noise After that, it is input into a small N-layer visual transformer model for processing, and the noise characteristics of each layer output are recorded as ;in, Represents the noise characteristics of the n-th layer output in a small visual transformer model; usually N can be set to 12;

[0058] Step 2.3: The cross attention layer is fused using formula (1) and , get the n-th layer fusion feature :

[0059] (1)

[0060] In formula (1), 、 、 Represents three weight matrices to be learned; represents the activation function, T represents the transpose, which can eliminate the domain difference between the noise and RGB branches, thereby facilitating the extraction of forged information;

[0061] Step 2.4: The memory mechanism layer uses formula (2) to Connect layers to get the enhanced features of the nth layer :

[0062] (2)

[0063] In formula (2), represents a linear layer, Indicates a concatenation operation; when season = ;

[0064] Step 2.5, The input is processed into the image encoder of the contrast language image pre-training model, so as to obtain the new encoding features of the nth layer using formula (3) and generate Category token , n∈[1,N];

[0065] (3)

[0066] In formula (3), Represents the parameter matrix to be learned in the nth layer. The present invention transforms The frozen contrastive language image pre-trained model image encoder is added to extract forged features while avoiding the catastrophic forgetting of the CLIP prior.

[0067] Step 2.6: The instance-aware two-way prompt learning module processes the vectors to be learned Vi, i∈{a,f} of the real image and the forged image to obtain the image-text similarity ρ;

[0068] Step 2.6.1: Use formula (4) to Make a rough adjustment and get the prompt after the rough adjustment ,in, indicates the corresponding hint of the real image, Indicates the corresponding hint of the forged image;

[0069] (4)

[0070] In formula (4), PANet represents the prompt adjustment network, which is a bottleneck structure consisting of a linear layer, a ReLU activation function, and a linear layer;

[0071] Step 2.6.2: Use formula (5) to get the adjusted text embedding ,in, represents the text embedding corresponding to the real image, represents the text embedding corresponding to the forged image;

[0072] (5)

[0073] In formula (5), For frozen text encoders, is a parameter to be learned, and EANet represents the adjustment network based on the transformer decoder;

[0074] Step 2.6.3: Use formula (6) to get the image-text similarity ρ:

[0075] (6)

[0076] In formula (6), represents cosine similarity;

[0077] Step 2.7: The prediction module uses formula (7) to obtain the predicted forged positioning map :

[0078] (7)

[0079] In formula (7), Describes the decoder, Features linear modulation operation, and has:

[0080] (8)

[0081] In formula (8), and represents 2 different linear layers, It is Hadamard;

[0082] Step 3: Construct the total loss function of the authentication network :

[0083] (9)

[0084] In formula (9), are two equilibrium parameters, Forged images The true mask label of represents the Dice loss, Represents the binary cross entropy loss. In the present invention, forgery localization and forgery detection are relatively independent, which can effectively reduce the false alarm rate of the forgery detection method; and:

[0085] (10)

[0086] In formula (10), is the input image X i If X i is a real image, then =1, otherwise, =0;

[0087] Step 4: Use the gradient descent method to train the authentication network and calculate the total loss function. To update the network parameters until the total loss function Until convergence, the optimal authentication network model is obtained.

[0088] In this embodiment, an electronic device includes a memory and a processor, wherein the memory is used to store a program that supports the processor to execute the above method, and the processor is configured to execute the program stored in the memory.

[0089] In this embodiment, a computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are executed.

[0090] Example:

[0091] In order to verify the effectiveness of this method, the commonly used Columbia dataset, Coverage dataset, CASIA dataset, NIST16 dataset and IMD20 dataset were selected in this example.

[0092] This experiment uses AUC and F1 as evaluation criteria;

[0093] In this embodiment, six methods using pre-trained models are selected to compare the forgery localization effects with the pre-trained model method of the present invention. The selected methods are ManTraNet, SPAN, PSCCNet, ObjectFormer, HiFi-Net, SAFL-Net and the invention method. ManTraNet comes from ManTra-Net: Manipulation tracing network for detection and localization of image forgeries with anomalous features, SPAN comes from SPAN: Spatial Pyramid Attention Network for Image Manipulation Localization, PSCC-Net comes from PSCC-Net: Progressive Spatio-Channel Correlation Network for Image Manipulation Detection and Localization, ObjectFormer comes from ObjectFormer for Image Manipulation Detection and Localization, HiFi-Net comes from Hierarchical Fine-Grained Image Forgery Detection and Localization, and SAFL-Net comes from SAFL-Net: Semantic-Agnostic Feature Learning Network with Auxiliary Plugins for Image Manipulation Detection. According to the experimental results, the results are shown in Table 1:

[0094] Table 1 Comparison of forged localization AUC (%) of different pre-trained models

[0095]

[0096] In this embodiment, six pre-training and fine-tuning methods and the pre-training and fine-tuning method of the present invention are selected to compare the forgery localization effect. The selected methods are RGB-N, SPAN, PSCCNet, ObjectFormer, HiFi-Net, SAFL-Net and the invention method. RGB-N comes from Learning Rich Features for Image Manipulation Detection, and the other comparison methods are from the same sources. According to the experimental results, the results are shown in Table 2:

[0097] Table 2 AUC (%) and F1 (%) results of forgery localization using fine-tuned models

[0098]

[0099] In this embodiment, six methods using pre-trained models are selected to compare the forgery detection effects with the pre-trained model method of the present invention. The selected methods are ManTraNet, SPAN, PSCCNet, ObjectFormer, HiFi-Net, SAFL-Net and the invention method. The comparison methods are derived from the same sources as above. The selected dataset is CASIA-D. According to the experimental results, the results are shown in Table 3:

[0100] Table 3 AUC (%) and F1 (%) results of forgery detection using pre-trained models

[0101]

[0102] The experimental results show that the proposed method performs better than other methods in both positioning and detection tasks, thus proving the feasibility of the proposed method.

[0103] In summary, the present invention proposes a novel image forgery localization framework, which introduces the contrastive language image pre-training model CLIP and prompt learning into the field of image forgery detection and localization; the framework includes instance-aware two-stream prompt learning and a forgery enhancement noise adapter; the present invention first creates a pair of learnable prompts as negative-positive samples to replace discrete prompts, and then fine-tunes these prompts according to the features and categories of each image; in addition, the present invention designs a forgery enhancement noise adapter to enhance the forgery perception ability of the image encoder through multi-domain fusion and zero linear layers; existing methods suffer from the problems of insufficient generalization and easy false alarms for real images; to overcome the above problems, the present invention introduces the contrastive language image pre-training model CLIP to utilize powerful visual language priors to improve the performance of image forgery detection and localization.

Claims

1. A method for detecting and locating image forgery based on noise-assisted cue learning, characterized in that: The steps are as follows: Step 1: Get the real image X a and the forged image X f , set X respectively a The vector to be learned V a and X f The vector to be learned V f ;Will and Any vector to be learned is denoted as V i ,i∈{a,f}; and Any image in is recorded as the input image X to be detected i ∈R H×W×C , where C represents the number of channels of the input image, H and W represent the height and width of the input image respectively; Step 2: Construct a forgery detection network, including: a language image comparison pre-training model, a Bayar convolutional layer, a small visual transformer model, a cross attention layer, a memory mechanism layer, an instance-aware dual-path prompt learning module, and a prediction module. i and V i Processing is performed to obtain the image-text similarity ρ and the predicted forged positioning map G out ; Step 2.1: Input image X i The N-layer image encoder frozen in the contrast language image pre-training model is processed, and the corresponding encoding features of each layer output are recorded as ;in, Represents the encoding features output by the n-th layer image encoder; n∈[1,N]; Step 2.2, Bayar convolution layer extracts the input image X i The noise G noise After that, it is input into a small N-layer visual transformer model for processing, and the noise characteristics of each layer output are recorded as ;in, Represents the noise characteristics of the n-th layer output in a small visual transformer model; Step 2.3: The cross attention layer is fused using formula (1) and , get the n-th layer fusion feature : (1) In formula (1), 、 、 Represents three weight matrices to be learned; represents the activation function, T represents transposition; Step 2.4: The memory mechanism layer uses formula (2) to Connect layers to get the enhanced features of the nth layer : (2) In formula (2), represents a linear layer, Indicates a concatenation operation; when season = ; Step 2.5, The input is processed into the image encoder of the contrast language image pre-training model, so as to obtain the new encoding features of the nth layer using formula (3) and generate Category token , n∈[1,N]; (3) In formula (3), Represents the parameter matrix to be learned in the nth layer; Step 2.6: The instance-aware two-way prompt learning module processes the vectors to be learned Vi, i∈{a,f} of the real image and the forged image to obtain the image-text similarity ρ; Step 2.6.1: Use formula (4) to Make a rough adjustment and get the prompt after the rough adjustment ,in, indicates the corresponding hint of the real image, Indicates the corresponding hint of the forged image; (4) In formula (4), PANet represents the prompt adjustment network, which is a bottleneck structure consisting of a linear layer, a ReLU activation function, and a linear layer; Step 2.6.2: Use formula (5) to get the adjusted text embedding ,in, represents the text embedding corresponding to the real image, represents the text embedding corresponding to the forged image; (5) In formula (5), For frozen text encoders, is a parameter to be learned, and EANet represents the adjustment network based on the transformer decoder; Step 2.6.3: Use formula (6) to get the image-text similarity ρ: (6) In formula (6), represents cosine similarity; Step 2.7: The prediction module uses formula (7) to obtain the predicted forged positioning map : (7) In formula (7), Describes the decoder, Features linear modulation operation, and has: (8) In formula (8), and represents 2 different linear layers, It is Hadamard; Step 3: Construct the total loss function of the authentication network : (9) In formula (9), are two equilibrium parameters, Forged images The true mask label of represents the Dice loss, represents the binary cross entropy loss; and: (10) In formula (10), is the input image X i If X i is a real image, then =1, otherwise, =0; Step 4: Use the gradient descent method to train the authentication network and calculate the total loss function. To update the network parameters until the total loss function Until convergence, the optimal authentication network model is obtained.

2. An electronic device comprising a memory and a processor, characterized in that: The memory is used to store a program that supports the processor to execute the image forgery detection and positioning method according to claim 1, and the processor is configured to execute the program stored in the memory.

3. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the image forgery detection and positioning method according to claim 1 are executed.

Citation Information

Patent Citations

  • Cross-modal cross-domain universal face forgery positioning method

    CN117292442A

  • Deep fake face detection positioning method capable of learning local difference

    CN117496583A