Image forgery detection and localization method based on lottery ticket hypothesis and masked autoencoder
By employing an image forgery detection and localization method based on the lottery hypothesis and mask autoencoder, forgery-sensitive parameters are screened and a multi-source forgery perception network is constructed. This solves the problems of insufficient accuracy and generalization in existing image forgery detection technologies, and achieves efficient forgery region identification and localization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- UNIV OF SCI & TECH OF CHINA
- Filing Date
- 2025-02-11
- Publication Date
- 2026-04-17
AI Technical Summary
Existing image forgery detection methods lack accuracy and generalization ability in handling complex forgery scenarios and out-of-distribution forgery detection tasks, and are unable to effectively utilize prior knowledge of natural images and forgery information.
We employ an image forgery detection and localization method based on the lottery hypothesis and mask autoencoder. By screening forgery-sensitive parameters, we construct a multi-source forgery perception network. Combining prior knowledge of natural images with multi-source forgery information, we utilize the multi-source forgery perception network framework and an 'expert hybrid' noise extractor to enhance the forgery feature perception capability.
It significantly improves the accuracy and generalization of image forgery detection, enabling efficient identification and precise localization of forged regions in complex forgery scenarios, thus enhancing the robustness and accuracy of the model.
Smart Images

Figure CN120032233B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image forgery detection and localization, specifically to an image forgery detection and localization method based on the lottery hypothesis and mask autoencoder. Background Technology
[0002] With the rapid development of image editing and generation technologies, image forgery has become increasingly easy and covert, posing a serious threat to social security, personal privacy, and legal fairness. For example, malicious users can use advanced image forgery techniques to manipulate images, create fake news, and even provide fabricated evidence in court. At the same time, the emergence of new technologies such as diffusion models has made image forgery methods increasingly complex, posing a significant challenge to media security. Therefore, developing efficient and widely applicable image forgery detection and localization methods has become particularly crucial.
[0003] Existing anti-forgery techniques have made some progress in combating image forgery, including detection using low-level features such as forgery traces. However, with the continuous advancement of forgery techniques, existing methods face significant challenges in handling complex forgery scenarios and out-of-distribution forgery detection tasks, and there is an urgent need for detection methods with greater accuracy and generalization ability. Summary of the Invention
[0004] To avoid the problems existing in the prior art, this invention proposes an image forgery detection and localization method based on the lottery hypothesis and mask autoencoder, aiming to solve the problems of insufficient utilization of prior knowledge of natural images and insufficient capture of forgery information in existing methods, thereby improving the accuracy and generalization of image forgery detection and localization.
[0005] To achieve the above-mentioned objectives, the present invention adopts the following technical solution:
[0006] The present invention provides an image forgery detection and localization method based on the lottery hypothesis and mask autoencoder, characterized by the following steps:
[0007] Step 1: Obtain an input image and perform preprocessing to obtain the preprocessed image. Make the image The true value of the location mask is The ground truth value of the detection label for image X is Where C represents the image The number of channels, H and W represent the image's channels, respectively. Height and width;
[0008] Step 2: Obtain the weight parameter set of the pre-trained masked autoencoder. And based on the lottery hypothesis, from the set of weight parameters After selecting and processing the parameters sensitive to image forgery detection and localization tasks, a gradient mask is obtained. Thus, based on the gradient mask Build an index set ;
[0009] Step 3: Construct a multi-source forgery perception network, including: an "expert hybrid" noise extractor, a noise stream network branch, an RGB stream network branch, a mask decoder, and a detector. and based on the index set right The image forgery location result is obtained through processing. Image forgery detection results ;
[0010] Step 4: Use equation (14) to establish the loss function of the multi-source forgery perception network. :
[0011] (14)
[0012] In equation (14), , For two binary cross-entropy losses, For Dice's loss, , There are two weighting factors;
[0013] Step 5: Generate the mask gradient of the loss function using equation (15). :
[0014] (15)
[0015] In equation (15), It is the product of Hadamah. These are parameters of the network encoder. yes The gradient;
[0016] Step 6: Train the multi-source forgery perception network using the ADAM optimizer, and adjust the loss function using equation (15). Update until the loss function is updated. The process continues until convergence, resulting in a trained multi-source forgery perception network model. This model is then used to process the image to be detected, yielding a predicted classification result for image forgery detection and a localization map of the suspected forgery region.
[0017] The present invention provides an image forgery detection and localization method based on the lottery hypothesis and mask autoencoder, characterized in that step 2 includes:
[0018] Step 2.1, the mask autoencoder is composed of... It consists of several Transformer-Base blocks, using weight parameters. The pre-trained mask autoencoder is initialized and self-reconstruction training is performed on a fake image dataset to obtain the fake weight parameter set of the trained mask autoencoder. And use equation (1) to calculate the set of spoofed change magnitudes of the weight parameters. :
[0019] (1)
[0020] In equation (1), This represents the operation of finding the K largest elements; Represents absolute value operation;
[0021] Step 2.2: Use weight parameters The pre-trained mask autoencoder is initialized and self-reconstruction training is performed on a real image dataset to obtain the true weight parameter set of the trained mask autoencoder. And use equation (2) to calculate the set of true change ranges of the weight parameters. :
[0022] (2)
[0023] Step 2.3: Use equation (3) to filter out the weight parameters that are sensitive to forgery. :
[0024] (3)
[0025] In equation (3), This represents the difference operation. Indicates the intersection operation;
[0026] Step 2.4: Generate using equation (4) The i-th weight parameter gradient mask Thus, the weight parameters are obtained. gradient mask :
[0027] (4)
[0028] Step 2.5: Based on the gradient mask Calculate the weight mask rate of the k-th Transformer-Base block in the masked autoencoder. Thus obtain The weight mask ratio of each Transformer-Base block is calculated, and the block with the highest weight mask ratio is selected. The index set consists of Transformer-Base blocks. ;in, This represents the i-th weight parameter in the k-th Transformer-Base block. The mask value, It is the total number of parameters in the k-th Transformer-Base block; k= .
[0029] Furthermore, step 3 includes:
[0030] The noise flow network branch in step 3 is composed of Transformer-Tiny blocks and their corresponding It consists of linear layers; the RGB network flow branch consists of a linear layer and It consists of several Transformer-Base blocks; The number of Transformer-Tiny blocks in the noise flow branch; The number of Transformer-Base blocks for the RGB stream branch;
[0031] Step 3.1, Image After processing by the "expert hybrid" noise enhancement device, multi-source forgery features are obtained. ;
[0032] Step 3.2, Noise Network Flow Branch Each Transformer-Tiny block addresses multi-source forgery features. Processing is performed to obtain Intermediate features of noise flow ,in, This represents an intermediate feature of the i-th noise stream. for The number of image patches in the image. for Feature dimensions;
[0033] Noisy network flow branches Linear layer Intermediate features of the noise flow Processing is performed to obtain Noise mapping features ,in, This represents the i-th linear layer. Represents the i-th noise mapping feature;
[0034] Step 3.3, Image After processing by the linear layer of the RGB network stream branch, the initial RGB stream intermediate features are obtained. Then input the cascaded branches of the RGB network stream. The intermediate features of the j-th RGB stream are processed in the Transformer-Base block and obtained using Equation (11). And thus obtain the first Intermediate features of RGB streams :
[0035] (11)
[0036] In equation (11), This represents the intermediate feature of the (j-1)th RGB stream. for The number of image patches in the image. for Feature dimensions; This represents the (j-1)th Transformer-Base block. This represents the k-th noise mapping feature;
[0037] Step 3.4: The mask decoder consists of several convolutional layers, batch normalization, ReLU function, and upsampling layers, and performs the following steps on the first... Intermediate features of RGB streams The image forgery location result is obtained through processing. ;
[0038] Step 3.5, Detector Using equation (12) to analyze the positioning results The image is converted to obtain the image forgery detection result. :
[0039] (12)
[0040] In equation (12), This represents a cascade operation of the first convolution, ReLU function, batch normalization, second convolution, and Sigmoid function. This represents a hyperparameter that decreases with the number of training iterations. Let represent the generalized average pooling operation, and we have:
[0041] (13)
[0042] In equation (13), p represents the parameters to be trained. for The pixel with coordinates (i,j) in the middle.
[0043] Furthermore, the "expert hybrid" noise denoising device in step 3.1 includes: an SRM filter, a Bayar convolution, a noise watermark extraction module, an expert weight generation module, and a second convolutional layer. The expert weight generation module consists of a first convolutional layer, a pooling layer, and a linear layer.
[0044] Step 3.1.1, Image After processing by an SRM filter, Bayar convolution, and a noise watermark extraction module, the SRM noise features are obtained. Bayar noise characteristics and noise watermark features ;
[0045] Step 3.1.2, Image The input is processed in the expert weight generation module to obtain the expert coefficient matrix. :
[0046] Step 3.1.2.1: Generate channel descriptors using equation (5). :
[0047] (5)
[0048] In equation (5), This represents the pixel value in the i-th row and j-th column of X;
[0049] Step 3.1.2.2: Obtain the first weight matrix using equation (6). ,in, Indicates the dimension of the first weight matrix:
[0050] (6)
[0051] In equation (6), It's a linear layer, while Pool is a pooling layer. It is the first convolutional layer;
[0052] Step 3.1.2.3: Obtain the expert coefficient matrix using equation (7). Where O is the number of experts:
[0053] (7)
[0054] In equation (7), Represents the ReLU function; The parameter matrix to be learned;
[0055] Step 3.1.3, Second convolutional layer Multi-source forgery features are obtained using equation (8). :
[0056] (8)
[0057] In equation (8), Indicates splicing.
[0058] The present invention provides an electronic device, including a memory and a processor, wherein the memory is used to store a program that supports the processor in executing the image forgery detection and localization method, and the processor is configured to execute the program stored in the memory.
[0059] The present invention discloses a computer-readable storage medium on which a computer program is stored, wherein the computer program, when executed by a processor, performs the steps of the image forgery detection and localization method.
[0060] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0061] 1. This invention proposes a multi-source forgery perception network framework. This framework, based on the feature extraction capabilities of a mask autoencoder and combined with prior knowledge of natural images, effectively improves the accuracy and generalization of image forgery detection and localization. By preserving the inherent characteristics of natural images and integrating multi-source forgery information, the multi-source forgery perception network performs excellently in handling complex forgery scenarios and out-of-distribution forgery detection tasks.
[0062] 2. This invention is the first to introduce the lottery hypothesis into the field of image forgery detection and localization, proposing a strategy based on the lottery hypothesis to identify forgery-sensitive parameters and perform sparse fine-tuning. By screening forgery-sensitive parameters during the pre-training stage, this invention can significantly enhance the model's ability to perceive forgery features while preserving natural image priors, thereby achieving more efficient forgery detection and localization.
[0063] 3. This invention develops an "expert hybrid" noise extractor to collect multi-source forgery information and inject it into the network layer containing forgery-sensitive parameters. During fine-tuning, this noise extractor enhances the model's ability to perceive forged images by fusing multiple forgery features, thereby further improving the robustness and accuracy of forgery detection. Attached Figure Description
[0064] Figure 1 This is a flowchart of an image forgery detection and localization method based on the lottery hypothesis and mask autoencoder according to the present invention;
[0065] Figure 2 The image shows the results compared to various other methods. Detailed Implementation
[0066] In this embodiment, an image forgery detection and localization method based on the lottery hypothesis and a masked autoencoder mainly preserves the inherent features of natural images while integrating multi-source forgery information. First, through pre-training with masked image self-reconstruction, parameters sensitive to forgery in the masked autoencoder are determined. Then, a multi-source forgery perception network is trained, and these parameters are fine-tuned using a lottery sparse fine-tuning strategy. Furthermore, a multi-source forgery perception fine-tuning strategy is employed to introduce multi-source forgery features into specific layers. According to the lottery hypothesis, the lottery sparse fine-tuning strategy can identify parameters related to forgery, freeze other parameters, and fine-tune only the relevant parameters. The multi-source forgery perception fine-tuning strategy enhances the lottery sparse fine-tuning strategy by injecting multi-source noise into selected layers, thereby improving the network's sensitivity to forgery. Specifically, as... Figure 1 As shown, the method is performed according to the following steps:
[0067] Step 1: Obtain an input image and perform preprocessing to obtain the preprocessed image. Make the image The true value of the location mask is The ground truth value of the detection label for image X is Where C represents the image The number of channels, H and W represent the image's channels, respectively. In this example, the height and width of H and W are both 512.
[0068] Step 2: Obtain the weight parameter set of the pre-trained masked autoencoder. And based on the lottery hypothesis, from the set of weight parameters After selecting and processing the parameters sensitive to image forgery detection and localization tasks, a gradient mask is obtained. The mask autoencoder is composed of It consists of several Transformer-Base blocks, in this example. The value is 12;
[0069] Step 2.1: Use weight parameters The pre-trained mask autoencoder is initialized and self-reconstruction training is performed on a fake image dataset to obtain the fake weight parameter set of the trained mask autoencoder. And use equation (1) to calculate the set of spoofed change magnitudes of the weight parameters. :
[0070] (1)
[0071] In equation (1), This represents the operation of finding the K largest elements; This represents absolute value operations.
[0072] Step 2.2: Use weight parameters The pre-trained mask autoencoder is initialized and self-reconstruction training is performed on a real image dataset to obtain the true weight parameter set of the trained mask autoencoder. And use equation (2) to calculate the set of true change ranges of the weight parameters. :
[0073] (2)
[0074] Step 2.3: Use equation (3) to filter out the weight parameters that are sensitive to forgery. :
[0075] (3)
[0076] In equation (3), This represents the difference operation. This indicates the intersection operation.
[0077] Step 2.4: Generate using equation (4) The i-th weight parameter gradient mask Thus, the weight parameters are obtained. gradient mask :
[0078] (4)
[0079] Step 2.5: Based on the gradient mask Calculate the weight mask rate of the k-th Transformer-Base block in the masked autoencoder. Thus obtain The weight mask rate of each Transformer-Base block, where... This represents the i-th weight parameter in the k-th Transformer-Base block. The mask value, It is the total number of parameters in the k-th Transformer-Base block; the range of k is... Select the one with the highest weight mask ratio. A set of indexes for Transformer-Base blocks .
[0080] Step 3: Construct a multi-source forgery perception network, including: an "expert hybrid" noise extractor, a noise stream network branch, an RGB stream network branch, a mask decoder, and a detector. Among them, the noise flow network branch is composed of Transformer-Tiny blocks and their corresponding It consists of linear layers; the RGB network flow branch consists of a linear layer and It consists of several Transformer-Base blocks; The number of Transformer-Tiny blocks in the noise flow branch, in this example, The value of is 4; The number of Transformer-Base blocks in the RGB stream branch.
[0081] Step 3.1, Image After processing by the "expert hybrid" noise enhancement device, multi-source forgery features are obtained. The "expert hybrid" noise denoising tool includes: an SRM filter, a Bayar convolution, a noise watermark extraction module, an expert weight generation module, and a second convolutional layer. The expert weight generation module consists of a first convolutional layer, a pooling layer, and a linear layer.
[0082] Step 3.1.1, Image After processing by an SRM filter, Bayar convolution, and a noise watermark extraction module, the SRM noise features are obtained. Bayar noise characteristics and noise watermark features ;
[0083] Step 3.1.2, Image The input is processed in the expert weight generation module to obtain the expert coefficient matrix. :
[0084] Step 3.1.2.1: Generate channel descriptors using equation (5). :
[0085] (5)
[0086] In equation (5), This represents the pixel value in the i-th row and j-th column of X.
[0087] Step 3.1.2.2: Obtain the first weight matrix using equation (6). ,in, Indicates the dimension of the first weight matrix:
[0088] (6)
[0089] In equation (6), It's a linear layer, while Pool is a pooling layer. It is the first convolutional layer.
[0090] Step 3.1.2.3: Obtain the expert coefficient matrix using equation (7). Where O is the number of experts:
[0091] (7)
[0092] In equation (7), Represents the ReLU function; The parameter matrix to be learned;
[0093] Step 3.1.3, Second convolutional layer Multi-source forgery features are obtained using equation (8). :
[0094] (8)
[0095] In equation (8), Indicates splicing.
[0096] Step 3.2, Noise Network Flow Branch Each Transformer-Tiny block addresses multi-source forgery features. Processing is performed to obtain Intermediate features of noise flow ,in, This represents an intermediate feature of the i-th noise stream. for The number of image patches in the image. for Feature dimensions; Noisy network flow branches Linear layer Intermediate features of the noise flow Processing according to equation (9), we obtain... Noise mapping features ,in, This represents the i-th linear layer. Represents the i-th noise mapping feature:
[0097] (9)
[0098] Step 3.3, Image After processing by the linear layer of the RGB network stream branch, the initial RGB stream intermediate features are obtained. Then input the cascaded branches of the RGB network stream. The intermediate features of the j-th RGB stream are processed in the Transformer-Base block and obtained using Equation (11). And thus obtain the first Intermediate features of RGB streams :
[0099] (10)
[0100] (11)
[0101] In equation (11), This represents the intermediate feature of the (j-1)th RGB stream. for The number of image patches in the image. for Feature dimensions; This represents the (j-1)th Transformer-Base block; Indicates the highest weight mask ratio A set of indices for Transformer-Base blocks. This represents the k-th noise mapping feature.
[0102] Step 3.4: The mask decoder consists of several convolutional layers, batch normalization, ReLU function, and upsampling layers, and performs the following steps on the first... Intermediate features of RGB streams The image forgery location result is obtained through processing. ;
[0103] Step 3.5, Detector Using equation (12) to analyze the positioning results The image is converted to obtain the image forgery detection result. :
[0104] (12)
[0105] In equation (12), This represents a cascade operation of the first convolution, ReLU function, batch normalization, second convolution, and Sigmoid function. This represents a hyperparameter that decreases with the number of training iterations; in this example, , Let represent the generalized average pooling operation, and we have:
[0106] (13)
[0107] In equation (13), p represents the parameters to be trained. for The pixel with coordinates (i,j) in the middle.
[0108] Step 4: Use equation (14) to establish the loss function of the multi-source forgery perception network. :
[0109] (14)
[0110] In equation (14), , For two binary cross-entropy losses, For Dice's loss, , There are two weighting factors; in this example, It was set to 0.15. It was set to 0.35.
[0111] Step 5: Generate the mask gradient of the loss function using equation (15). :
[0112] (15)
[0113] In equation (15), It is the product of Hadamah. These are parameters of the network encoder. yes The gradient.
[0114] Step 6: Train the multi-source forgery perception network using the ADAM optimizer, and adjust the loss function using equation (15). Update until the loss function is updated. The process continues until convergence, resulting in a trained multi-source forgery perception network model. This model is then used to process the image to be detected, yielding a predicted classification result for image forgery detection and a localization map of the suspected forgery region.
[0115] In this embodiment, an electronic device includes a memory and a processor. The memory stores a program that supports the processor in executing the methods described above, and the processor is configured to execute the program stored in the memory.
[0116] In this embodiment, a computer-readable storage medium stores a computer program, which is executed by a processor to perform the steps of the above method.
[0117] Example:
[0118] To verify the effectiveness of this method, the present invention selected commonly used datasets such as Casiav1.0, Columbia, NIST16, IMD2020, DSO-1, Korus, AutoSplice, and OpenForensics.
[0119] This invention uses F1, AUC and ACC as evaluation criteria.
[0120] In this embodiment, the localization methods of six models and the localization method of the present invention are compared. The selected methods are Multi-view Multi-scale Supervised Network (MVSS-Net), Compression Artifact Tracing Network (CAT-Net), Progressive Spatio-Channel Correlation Network (PSCCNet), Hierarchical Fine-grained Network (HiFi-Net), and DiffForensics based on diffusion model. The selected datasets are Casiav1.0, Columbia, NIST16, IMD2020, DSO-1, Korus, AutoSplice, and OpenForensics. The experimental results are shown in Table 1.
[0121] Table 1 Comparison of localization F1 and AUC for different models
[0122]
[0123] In this embodiment, the detection methods of six different models and the detection method of the present invention are compared. The selected methods are Multi-view Multi-scale Supervised Network (MVSS-Net), Compression Artifact Tracing Network (CAT-Net), Progressive Spatio-Channel Correlation Network (PSCCNet), Hierarchical Fine-grained Network (HiFi-Net), and DiffForensics based on diffusion model. The selected datasets are Casiav1.0, Columbia, IMD2020, and AutoSplice. The experimental results are shown in Table 2.
[0124] Table 2 Comparison of ACC and AUC for different models
[0125]
[0126] Experimental results show that the method of the present invention performs better than the detection and localization methods of the other six models, thus proving the feasibility of the method proposed in this invention.
[0127] Figure 2 The qualitative comparison between the positioning results of the present invention and those of other methods is shown. It can be seen that the positioning results of the present invention are accurate and the edges are clear, indicating that the method proposed in the present invention can effectively utilize prior information from real images and multi-source forgery information to achieve accurate positioning of forged areas.
[0128] In summary, this invention has undergone extensive experiments on multiple benchmarks, demonstrating that it outperforms state-of-the-art image forgery detection and localization methods both qualitatively and quantitatively. For example, the localization results achieve an F1 score of 0.612 and an AUC of 0.876 on the Casiav1.0 dataset, and an F1 score of 0.639 and an AUC of 0.950 on the AutoSplice dataset. The detection results achieve an ACC of 0.912 and an AUC of 0.989 on the Columbia dataset, all surpassing existing methods.
Claims
1. A method for detecting and locating image forgery based on the lottery hypothesis and mask autoencoder, characterized in that, The procedure is as follows: Step 1: Obtain an input image and perform preprocessing to obtain the preprocessed image. Make the image The true value of the location mask is The ground truth value of the detection label for image X is Where C represents the image The number of channels, H and W represent the image's channels, respectively. Height and width; Step 2: Obtain the weight parameter set of the pre-trained masked autoencoder. And based on the lottery hypothesis, from the set of weight parameters After selecting and processing the parameters sensitive to image forgery detection and localization tasks, a gradient mask is obtained. Thus, based on the gradient mask Build an index set ; Step 3: Construct a multi-source forgery perception network, including: an "expert hybrid" noise extractor, a noise stream network branch, an RGB stream network branch, a mask decoder, and a detector. and based on the index set right The image forgery location result is obtained through processing. Image forgery detection results Among them, the noise flow network branch is composed of Transformer-Tiny blocks and their corresponding It consists of linear layers; the RGB network flow branch consists of a linear layer and It consists of several Transformer-Base blocks; The number of Transformer-Tiny blocks in the noise flow branch; The number of Transformer-Base blocks for the RGB stream branch; Step 4: Use equation (14) to establish the loss function of the multi-source forgery perception network. : (14) In equation (14), , For two binary cross-entropy losses, For Dice's loss, , There are two weighting factors; Step 5: Generate the mask gradient of the loss function using equation (15). : (15) In equation (15), It is the product of Hadamah. These are parameters of the network encoder. yes The gradient; Step 6: Train the multi-source forgery perception network using the ADAM optimizer, and adjust the loss function using equation (15). Update until the loss function is updated. The process continues until convergence, resulting in a trained multi-source forgery perception network model. This model is then used to process the image to be detected, yielding a predicted classification result for image forgery detection and a localization map of the suspected forgery region.
2. The image forgery detection and localization method based on the lottery hypothesis and mask autoencoder according to claim 1, characterized in that, Step 2 includes: Step 2.1, the mask autoencoder is composed of... It consists of several Transformer-Base blocks, using weight parameters. The pre-trained mask autoencoder is initialized and self-reconstruction training is performed on a fake image dataset to obtain the fake weight parameter set of the trained mask autoencoder. And use equation (1) to calculate the set of spoofed change magnitudes of the weight parameters. : (1) In equation (1), This represents the operation of finding the K largest elements; Represents absolute value operation; Step 2.2: Use weight parameters The pre-trained mask autoencoder is initialized and self-reconstruction training is performed on a real image dataset to obtain the true weight parameter set of the trained mask autoencoder. And use equation (2) to calculate the set of true change ranges of the weight parameters. : (2) Step 2.3: Use equation (3) to filter out the weight parameters that are sensitive to forgery. : (3) In equation (3), This represents the difference operation. Indicates the intersection operation; Step 2.4: Generate using equation (4) The i-th weight parameter gradient mask Thus, the weight parameters are obtained. gradient mask : (4) Step 2.5: Based on the gradient mask Calculate the weight mask rate of the k-th Transformer-Base block in the masked autoencoder. Thus obtain The weight mask ratio of each Transformer-Base block is calculated, and the block with the highest weight mask ratio is selected. The index set consists of Transformer-Base blocks. ;in, This represents the i-th weight parameter in the k-th Transformer-Base block. The mask value, It is the total number of parameters in the k-th Transformer-Base block; k= .
3. The image forgery detection and localization method based on the lottery hypothesis and mask autoencoder according to claim 2, characterized in that, Step 3 includes: Step 3.1, Image After processing by the "expert hybrid" noise enhancement device, multi-source forgery features are obtained. ; Step 3.2, Noise Network Flow Branch Each Transformer-Tiny block addresses multi-source forgery features. Processing is performed to obtain Intermediate features of noise flow ,in, This represents an intermediate feature of the i-th noise stream. for The number of image patches in the image. for Feature dimensions; Noisy network flow branches Linear layer Intermediate features of the noise flow Processing is performed to obtain Noise mapping features ,in, This represents the i-th linear layer. Represents the i-th noise mapping feature; Step 3.3, Image After processing by the linear layer of the RGB network stream branch, the initial RGB stream intermediate features are obtained. Then input the cascaded branches of the RGB network stream. The intermediate features of the j-th RGB stream are processed in the Transformer-Base block and obtained using Equation (11). And thus obtain the first Intermediate features of RGB streams : (11) In equation (11), This represents the intermediate feature of the (j-1)th RGB stream. for The number of image patches in the image. for Feature dimensions; This represents the (j-1)th Transformer-Base block. This represents the k-th noise mapping feature; Step 3.4: The mask decoder consists of several convolutional layers, batch normalization, ReLU function, and upsampling layers, and performs the following steps on the first... Intermediate features of RGB streams The image forgery location result is obtained through processing. ; Step 3.5, Detector Using equation (12) to analyze the positioning results The image is converted to obtain the image forgery detection result. : (12) In equation (12), This represents a cascade operation of the first convolution, ReLU function, batch normalization, second convolution, and Sigmoid function. This represents a hyperparameter that decreases with the number of training iterations. Let represent the generalized average pooling operation, and we have: (13) In equation (13), p represents the parameters to be trained. for The pixel with coordinates (i,j) in the middle.
4. The image forgery detection and localization method based on the lottery hypothesis and mask autoencoder according to claim 3, characterized in that, The "expert hybrid" noise denoising device in step 3.1 includes: an SRM filter, a Bayar convolution, a noise watermark extraction module, an expert weight generation module, and a second convolutional layer. The expert weight generation module consists of a first convolutional layer, a pooling layer, and a linear layer. Step 3.1.1, Image After processing by an SRM filter, Bayar convolution, and a noise watermark extraction module, the SRM noise features are obtained. Bayar noise characteristics and noise watermark features ; Step 3.1.2, Image The input is processed in the expert weight generation module to obtain the expert coefficient matrix. : Step 3.1.2.1: Generate channel descriptors using equation (5). : (5) In equation (5), This represents the pixel value in the i-th row and j-th column of X; Step 3.1.2.2: Obtain the first weight matrix using equation (6). ,in, Indicates the dimension of the first weight matrix: (6) In equation (6), It's a linear layer, while Pool is a pooling layer. It is the first convolutional layer; Step 3.1.2.3: Obtain the expert coefficient matrix using equation (7). Where O is the number of experts: (7) In equation (7), Represents the ReLU function; The parameter matrix to be learned; Step 3.1.3, Second convolutional layer Multi-source forgery features are obtained using equation (8). : (8) In equation (8), Indicates splicing.
5. An electronic device, comprising a memory and a processor, characterized in that, The memory is used to store a program that supports the processor in executing the image forgery detection and localization method according to any one of claims 1-4, and the processor is configured to execute the program stored in the memory.
6. A computer-readable storage medium storing a computer program thereon, characterized in that, When the computer program is run by the processor, it performs the steps of the image forgery detection and localization method according to any one of claims 1-4.
Citation Information
Patent Citations
Building edge optimization method based on multi-task learning and dual lottery hypothesis
CN116052006A
Image forgery detecting and positioning method based on noise auxiliary prompt learning
CN119295383A