Image forgery detecting and positioning method based on lottery hypothesis and mask auto-encoder

By introducing lottery hypothesis and mask autoencoder in image forgery detection, combining multi-source forgery perception network and ‘expert hybrid’ noise extractor, the problem of insufficient accuracy and generalization capabilities of complex forgery scenarios and out-of-distribution forgery detection tasks in the prior art is solved, and more efficient image forgery detection and positioning effects are achieved.

CN120032233AActive Publication Date: 2025-05-23UNIV OF SCI & TECH OF CHINA

Patent Information

Application Number
CN202510147847.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-23
Estimated Expiration
2045-02-11

AI Technical Summary

Technical Problem

The prior art has problems of insufficient accuracy and generalization capabilities in dealing with complex forgery scenarios and out-of-distribution forgery detection tasks.

Method used

The image forgery detection and positioning method based on lottery hypothesis and mask autoencoder is adopted, and a multi-source forgery perception network and an ‘expert hybrid’ noise extractor are combined with the prior knowledge of natural images to identify forgery sensitive parameters and perform sparse fine-tuning.

Benefits of technology

It significantly improves the accuracy and generalization of image forgery detection and positioning, and can effectively deal with complex forgery scenarios and out-of-distribution forgery detection tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032233A_ABST
    Figure CN120032233A_ABST
Patent Text Reader

Abstract

The invention discloses an image forgery detecting and positioning method based on lottery hypotheses and a mask auto-encoder, which comprises the following steps of: 1, pre-training an image by using the mask auto-encoder, and extracting priori knowledge of a natural image; 2, identifying parameters sensitive to a forgery detection task in a pre-training stage based on the lottery hypothesis to obtain a gradient mask; and 3, constructing a multi-source counterfeit sensing network, realizing sparse fine tuning optimization based on gradient masks, and training the network so that the network can output accurate detection results and positioning results. According to the method, the feature extraction capability of the mask auto-encoder and the parameter screening strategy of the lottery hypothesis are combined, and the accuracy and generalization of counterfeit detection and positioning are remarkably improved through sparse fine tuning and multi-source feature fusion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the field of image forgery detection and positioning, and in particular to an image forgery detection and positioning method based on lottery ticket hypothesis and mask autoencoder. Background Art

[0002] With the rapid development of image editing and generation technology, the generation of image forgeries has become easier and more covert, posing a serious threat to social security, personal privacy, and legal justice. For example, malicious users can use advanced image forgery technology to manipulate images, create fake news, and even provide forged evidence in court. At the same time, the emergence of emerging technologies such as diffusion models has made image forgery methods more complicated, posing a huge challenge to media security protection. Therefore, it is particularly important to develop efficient and widely applicable image forgery detection and localization methods.

[0003] Existing forgery detection technologies have made some progress in dealing with image forgeries, including using underlying features (such as forgery traces) for detection. However, with the continuous advancement of forgery technology, existing methods face significant challenges in dealing with complex forgery scenarios and out-of-distribution forgery detection tasks, and there is an urgent need for more accurate and generalizable detection methods. Summary of the invention

[0004] In order to avoid the problems existing in the above-mentioned prior art, the present invention proposes an image forgery detection and positioning method based on the lottery ticket hypothesis and mask autoencoder, in order to solve the problems of insufficient utilization of prior knowledge of natural images and insufficient capture of forgery information in the existing methods, thereby improving the accuracy and generalization of image forgery detection and positioning.

[0005] In order to achieve the above-mentioned purpose, the present invention adopts the following technical scheme:

[0006] The image forgery detection and positioning method based on the lottery hypothesis and masked autoencoder of the present invention is characterized in that it is performed according to the following steps:

[0007] Step 1: Get an input image and preprocess it to get the preprocessed image , let the image The true value of the positioning mask is , the true value of the detection label of image X is , where C represents the image The number of channels, H and W represent the image height and width;

[0008] Step 2: Get the weight parameter set of the pre-trained mask autoencoder , and based on the lottery hypothesis from the weight parameter set After filtering out the parameters that are sensitive to the image forgery detection and positioning task and processing them, the gradient mask is obtained. , so that according to the gradient mask Build index collection ;

[0009] Step 3: Build a multi-source forgery-aware network, including: "expert mixture" noise extractor, noise stream network branch, RGB stream network branch, mask decoder and detector , and based on the index collection right Processing is performed to obtain the image forgery positioning result And image forgery detection results ;

[0010] Step 4: Use formula (14) to establish the loss function of the multi-source forgery perception network :

[0011] (14)

[0012] In formula (14), , are two binary cross entropy losses, is the Dice loss, , are two weight factors;

[0013] Step 5: Generate the mask gradient of the loss function using formula (15) :

[0014] (15)

[0015] In formula (15), is the Hadamard product, are the parameters of the network encoder, yes The gradient of

[0016] Step 6: Use the ADAM optimizer to train the multi-source forgery perception network, and use formula (15) to optimize the loss function Update until the loss function The trained multi-source forgery-aware network model is obtained until convergence, which is used to process the image to be detected to obtain the predicted classification results of image forgery detection and the positioning map of the suspected forgery area.

[0017] The image forgery detection and positioning method based on the lottery hypothesis and mask autoencoder of the present invention is also characterized in that the step 2 comprises:

[0018] Step 2.1, the masked self-encoder is composed of Transformer-Base blocks, using weight parameters Initialize the pre-trained mask autoencoder and perform self-reconstruction training on the forged image dataset to obtain the forged weight parameter set of the trained mask autoencoder , and use formula (1) to calculate the forged change range set of weight parameters :

[0019] (1)

[0020] In formula (1), Represents the operation of finding the first K largest elements; Represents absolute value operation;

[0021] Step 2.2: Use weight parameters Initialize the pre-trained mask autoencoder and perform self-reconstruction training on the real image dataset to obtain the real weight parameter set of the trained mask autoencoder , and use formula (2) to calculate the actual change range of the weight parameter :

[0022] (2)

[0023] Step 2.3: Use formula (3) to filter out weight parameters that are sensitive to forgery :

[0024] (3)

[0025] In formula (3), represents the difference operation, Represents the intersection operation;

[0026] Step 2.4: Generate using formula (4) The i-th weight parameter in Gradient mask of , thus obtaining the weight parameter Gradient mask of :

[0027] (4)

[0028] Step 2.5: Gradient mask Compute the weighted mask rate of the kth Transformer-Base block in the masked autoencoder , thus obtaining The weight mask rate of the Transformer-Base block, and select the one with the highest weight mask rate Transformer-Base blocks constitute an index set ;in, Represents the i-th weight parameter in the k-th Transformer-Base block The mask value of is the total number of parameters in the kth Transformer-Base block; k= .

[0029] Furthermore, the step 3 comprises:

[0030] The noise flow network branch in step 3 is composed of Transformer-Tiny blocks and their corresponding Linear layers; the RGB network stream branch consists of a linear layer and Transformer-Base blocks; is the number of Transformer-Tiny blocks in the noise stream branch; is the number of Transformer-Base blocks in the RGB stream branch;

[0031] Step 3.1: Image After being processed by the “Expert Mixture” noise denoiser, multi-source forged features are obtained ;

[0032] Step 3.2: Noisy network flow branch The Transformer-Tiny blocks forge features from multiple sources respectively. Process and obtain Noise flow intermediate features ,in, represents the intermediate feature of the i-th noise stream, for The number of image blocks in for The characteristic dimension of

[0033] Noisy network flow branching Linear layers The intermediate features of the noise flow Process and obtain Noise Mapping Features ,in, represents the i-th linear layer, represents the i-th noise mapping feature;

[0034] Step 3.3, Image After the linear layer processing of the RGB network flow branch, the initial RGB flow intermediate features are obtained , and then input the cascaded RGB network flow branch The jth RGB stream intermediate feature is obtained by using formula (11) , and then get the RGB stream intermediate features :

[0035] (11)

[0036] In formula (11), represents the intermediate feature of the j-1th RGB stream, for The number of image blocks in for The characteristic dimension of represents the j-1th Transformer-Base block, represents the kth noise mapping feature;

[0037] Step 3.4: The mask decoder consists of several convolutional layers, batch normalization, ReLU functions, and upsampling layers. RGB stream intermediate features Processing is performed to obtain the image forgery positioning result ;

[0038] Step 3.5: Detector Using formula (12) to calculate the positioning result Convert and get the image forgery detection result :

[0039] (12)

[0040] In formula (12), Represents the cascade operation of the first convolution, ReLU function, batch normalization, the second convolution, and the Sigmoid function. represents a hyperparameter that decreases with the number of training iterations, represents the generalized average pooling operation, and has:

[0041] (13)

[0042] In formula (13), p is the parameter to be trained, for The pixel point with coordinates (i, j) in the image.

[0043] Furthermore, the "expert hybrid" noise enhancer in step 3.1 includes: SRM filter, Bayar convolution, noise watermark extraction module, expert weight generation module and second convolution layer ; Among them, the expert weight generation module consists of the first convolutional layer, the pooling layer, and the linear layer;

[0044] Step 3.1.1, Image After being processed by the SRM filter, Bayar convolution and noise watermark extraction module, the SRM noise feature is obtained. , Bayar noise characteristics and noise watermark features ;

[0045] Step 3.1.2, Image Input into the expert weight generation module for processing to obtain the expert coefficient matrix :

[0046] Step 3.1.2.1: Generate channel descriptor using formula (5) :

[0047] (5)

[0048] In formula (5), represents the pixel value of the i-th row and j-th column in X;

[0049] Step 3.1.2.2: Use formula (6) to get the first weight matrix ,in, Represents the dimensions of the first weight matrix:

[0050] (6)

[0051] In formula (6), is a linear layer, Pool is a pooling layer, is the first convolutional layer;

[0052] Step 3.1.2.3: Use formula (7) to get the expert coefficient matrix , where O is the number of experts:

[0053] (7)

[0054] In formula (7), represents the ReLU function; is the parameter matrix to be learned;

[0055] Step 3.1.3, second convolutional layer Using formula (8) to obtain multi-source forgery features :

[0056] (8)

[0057] In formula (8), Indicates splicing.

[0058] An electronic device of the present invention includes a memory and a processor, wherein the memory is used to store a program that supports the processor to execute the image forgery detection and positioning method, and the processor is configured to execute the program stored in the memory.

[0059] The present invention provides a computer-readable storage medium, wherein a computer program is stored on the computer-readable storage medium, and the computer program executes the steps of the image forgery detection and positioning method when the computer program is executed by a processor.

[0060] Compared with the prior art, the present invention has the following beneficial effects:

[0061] 1. This paper proposes a multi-source forgery-aware network framework, which is based on the feature extraction capability of masked autoencoders and combines the prior knowledge of natural images to effectively improve the accuracy and generalization of image forgery detection and localization. By retaining the inherent characteristics of natural images and integrating multi-source forgery information, the multi-source forgery-aware network performs well in processing complex forgery scenes and out-of-distribution forgery detection tasks.

[0062] 2. This paper introduces the lottery ticket hypothesis into the field of image forgery detection and positioning for the first time, and proposes a strategy based on the lottery ticket hypothesis for identifying forgery-sensitive parameters and performing sparse fine-tuning. By screening out forgery-sensitive parameters in the pre-training stage, this paper can significantly enhance the model's ability to perceive forgery features while retaining the natural image prior, thereby achieving more efficient forgery detection and positioning.

[0063] 3. We developed a “mixture of experts” noise extractor to collect multi-source forgery information and inject it into the network layer where forgery-sensitive parameters are located. During fine-tuning, the noise extractor enhances the model’s perception of forged images by fusing multiple forgery features, thereby further improving the robustness and accuracy of forgery detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Figure 1 A flowchart of an image forgery detection and positioning method based on the lottery hypothesis and mask autoencoder of the present invention;

[0065] Figure 2 The following is a comparison chart of the effects of various methods. DETAILED DESCRIPTION

[0066] In the present embodiment, a method for detecting and locating image forgery based on the lottery ticket hypothesis and mask autoencoder mainly retains the inherent features of natural images while integrating multi-source forgery information. First, the parameters in the mask autoencoder that are sensitive to forgery are determined through pre-training of mask image self-reconstruction. Then, the multi-source forgery-aware network is trained, and the lottery ticket sparse fine-tuning strategy is applied to fine-tune these parameters, and the multi-source forgery-aware fine-tuning strategy is used to introduce multi-source forgery features in specific layers. According to the lottery ticket hypothesis, the lottery ticket sparse fine-tuning strategy can identify parameters related to forgery, freeze other parameters, and only fine-tune related parameters. The multi-source forgery-aware fine-tuning strategy enhances the lottery ticket sparse fine-tuning strategy by injecting multi-source noise into the selected layer, thereby increasing the network's sensitivity to forgery. Specifically, if Figure 1 As shown, the method is performed in the following steps:

[0067] Step 1: Get an input image and preprocess it to get the preprocessed image , let the image The true value of the positioning mask is , the true value of the detection label of image X is , where C represents the image The number of channels, H and W represent the image In this example, the values ​​of H and W are both 512.

[0068] Step 2: Get the weight parameter set of the pre-trained mask autoencoder , and based on the lottery hypothesis from the weight parameter set After filtering out the parameters that are sensitive to the image forgery detection and positioning task and processing them, the gradient mask is obtained. , the masked autoencoder is composed of Transformer-Base blocks. In this example, The value of is 12;

[0069] Step 2.1: Use weight parameters Initialize the pre-trained mask autoencoder and perform self-reconstruction training on the forged image dataset to obtain the forged weight parameter set of the trained mask autoencoder , and use formula (1) to calculate the forged change range set of weight parameters :

[0070] (1)

[0071] In formula (1), Represents the operation of finding the first K largest elements; Represents the absolute value operation.

[0072] Step 2.2: Use weight parameters Initialize the pre-trained mask autoencoder and perform self-reconstruction training on the real image dataset to obtain the real weight parameter set of the trained mask autoencoder , and use formula (2) to calculate the actual change range of the weight parameter :

[0073] (2)

[0074] Step 2.3: Use formula (3) to filter out weight parameters that are sensitive to forgery :

[0075] (3)

[0076] In formula (3), represents the difference operation, Represents an intersection operation.

[0077] Step 2.4: Generate using formula (4) The i-th weight parameter in Gradient mask of , thus obtaining the weight parameter Gradient mask of :

[0078] (4)

[0079] Step 2.5: Gradient mask Compute the weighted mask rate of the kth Transformer-Base block in the masked autoencoder , thus obtaining The weight mask rate of the Transformer-Base block, where Represents the i-th weight parameter in the k-th Transformer-Base block The mask value of is the total number of parameters in the kth Transformer-Base block; the range of k is , select the one with the highest weighted masking rate The index set of Transformer-Base blocks .

[0080] Step 3: Build a multi-source forgery-aware network, including: "expert mixture" noise extractor, noise stream network branch, RGB stream network branch, mask decoder and detector , where the noise flow network branch is composed of Transformer-Tiny blocks and their corresponding Linear layers; the RGB network stream branch consists of a linear layer and Transformer-Base blocks; is the number of Transformer-Tiny blocks in the noise stream branch. In this example, The value of is 4; is the number of Transformer-Base blocks in the RGB stream branch.

[0081] Step 3.1: Image After being processed by the "Expert Mixture" noise denoiser, multi-source forged features are obtained The "expert hybrid" noise denoiser includes: SRM filter, Bayar convolution, noise watermark extraction module, expert weight generation module and the second convolution layer ; Among them, the expert weight generation module consists of the first convolutional layer, pooling layer, and linear layer.

[0082] Step 3.1.1, Image After being processed by the SRM filter, Bayar convolution and noise watermark extraction module, the SRM noise feature is obtained. , Bayar noise characteristics and noise watermark features ;

[0083] Step 3.1.2, Image Input into the expert weight generation module for processing to obtain the expert coefficient matrix :

[0084] Step 3.1.2.1: Generate channel descriptor using formula (5) :

[0085] (5)

[0086] In formula (5), Represents the pixel value at the i-th row and j-th column in X.

[0087] Step 3.1.2.2: Use formula (6) to get the first weight matrix ,in, Represents the dimensions of the first weight matrix:

[0088] (6)

[0089] In formula (6), is a linear layer, Pool is a pooling layer, is the first convolutional layer.

[0090] Step 3.1.2.3: Use formula (7) to get the expert coefficient matrix , where O is the number of experts:

[0091] (7)

[0092] In formula (7), represents the ReLU function; is the parameter matrix to be learned;

[0093] Step 3.1.3, second convolutional layer Using formula (8) to obtain multi-source forgery features :

[0094] (8)

[0095] In formula (8), Indicates splicing.

[0096] Step 3.2: Noisy network flow branch The Transformer-Tiny blocks forge features from multiple sources respectively. Process and obtain Noise flow intermediate features ,in, represents the intermediate feature of the i-th noise stream, for The number of image blocks in for The characteristic dimension of the noisy network flow branch Linear layers The intermediate features of the noise flow According to formula (9), we can get Noise Mapping Features ,in, represents the i-th linear layer, Represents the i-th noise mapping feature:

[0097] (9)

[0098] Step 3.3, Image After the linear layer processing of the RGB network flow branch, the initial RGB flow intermediate features are obtained , and then input the cascaded RGB network flow branch The jth RGB stream intermediate feature is obtained by using formula (11) , and then get the RGB stream intermediate features :

[0099] (10)

[0100] (11)

[0101] In formula (11), represents the intermediate feature of the j-1th RGB stream, for The number of image blocks in for The characteristic dimension of represents the j-1th Transformer-Base block; Indicates the weighted mask with the highest rate The index set of Transformer-Base blocks, represents the kth noise mapping feature.

[0102] Step 3.4: The mask decoder consists of several convolutional layers, batch normalization, ReLU functions, and upsampling layers. RGB stream intermediate features Processing is performed to obtain the image forgery positioning result ;

[0103] Step 3.5: Detector Using formula (12) to calculate the positioning result Convert and get the image forgery detection result :

[0104] (12)

[0105] In formula (12), Represents the cascade operation of the first convolution, ReLU function, batch normalization, the second convolution, and the Sigmoid function. represents a hyperparameter that decreases with the number of training iterations. In this example, , represents the generalized average pooling operation, and has:

[0106] (13)

[0107] In formula (13), p is the parameter to be trained, for The pixel point with coordinates (i, j) in the image.

[0108] Step 4: Use formula (14) to establish the loss function of the multi-source forgery perception network :

[0109] (14)

[0110] In formula (14), , are two binary cross entropy losses, is the Dice loss, , are two weight factors; in this example, is set to 0.15, is set to 0.35.

[0111] Step 5: Generate the mask gradient of the loss function using formula (15) :

[0112] (15)

[0113] In formula (15), is the Hadamard product, are the parameters of the network encoder, yes gradient.

[0114] Step 6: Use the ADAM optimizer to train the multi-source forgery perception network, and use formula (15) to optimize the loss function Update until the loss function The trained multi-source forgery-aware network model is obtained until convergence, which is used to process the image to be detected to obtain the predicted classification results of image forgery detection and the positioning map of the suspected forgery area.

[0115] In this embodiment, an electronic device includes a memory and a processor, wherein the memory is used to store a program that supports the processor to execute the above method, and the processor is configured to execute the program stored in the memory.

[0116] In this embodiment, a computer-readable storage medium stores a computer program on the computer-readable storage medium, and the computer program executes the steps of the above method when executed by a processor.

[0117] Example:

[0118] In order to verify the effectiveness of this method, the present invention selected the commonly used Casiav1.0 dataset, Columbia dataset, NIST16 dataset, IMD2020 dataset, DSO-1 dataset, Korus dataset, AutoSplice dataset and OpenForensics dataset.

[0119] The present invention adopts F1, AUC and ACC as evaluation criteria.

[0120] In this embodiment, the positioning methods of 6 models and the positioning method of the model of the present invention are selected for effect comparison. The selected methods are the Multi-View Multi-Scale Supervised Network (MVSS-Net), the Compression Artifact Tracing Network (CAT-Net), the Progressive Spatio-Channel Correlation Network (PSCCNet), the Hierarchical Fine-grained Network (HiFi-Net), and forgery detection based on diffusion model (DiffForensics). The selected datasets are Casiav1.0, Columbia, NIST16, IMD2020, DSO-1, Korus, AutoSplice, and OpenForensics. According to the experimental results, the results are shown in Table 1 as follows:

[0121] Table 1 Comparison of positioning F1 and AUC of different models

[0122]

[0123] In this embodiment, the detection methods of 6 models and the detection method of the model of the present invention are selected for effect comparison. The selected methods are the Multi-View Multi-Scale Supervised Network (MVSS-Net), the Compression Artifact Tracing Network (CAT-Net), the Progressive Spatio-Channel Correlation Network (PSCCNet), the Hierarchical Fine-grained Network (HiFi-Net), and forgery detection based on diffusion model (DiffForensics). The selected datasets are Casiav1.0, Columbia, IMD2020, and AutoSplice. According to the experimental results, the results are shown in Table 2 as follows:

[0124] Table 2 Comparison of detection ACC and AUC of different models

[0125]

[0126] The experimental results show that the method of the present invention has better effects compared with the detection and positioning methods of the other 6 models, thus proving the feasibility of the method proposed by the present invention.

[0127] Figure 2 A qualitative comparison of the positioning results of the present invention and those of other methods is shown, and it can be seen that the positioning results of the present invention are accurate and have clear edges, indicating that the method proposed in the present invention can effectively utilize real image priors and multi-source forged information to accurately locate the forged area.

[0128] In summary, the present invention has been extensively experimented on multiple benchmarks and has proven that the present invention is superior to the most advanced image forgery detection and localization methods both qualitatively and quantitatively. For example, the localization results on the Casiav1.0 dataset have an F1 of 0.612 and an AUC of 0.876, on the AutoSplice dataset have an F1 of 0.639 and an AUC of 0.950, and the detection results on the Columbia dataset have an ACC of 0.912 and an AUC of 0.989, all of which surpass existing methods.

Claims

1. A method for detecting and locating image forgery based on the lottery hypothesis and masked autoencoder, characterized in that: The steps are as follows: Step 1: Get an input image and preprocess it to get the preprocessed image , let the image The true value of the positioning mask is , the true value of the detection label of image X is , where C represents the image The number of channels, H and W represent the image height and width; Step 2: Get the weight parameter set of the pre-trained mask autoencoder , and based on the lottery hypothesis from the weight parameter set After filtering out the parameters that are sensitive to the image forgery detection and positioning task and processing them, the gradient mask is obtained. , so that according to the gradient mask Build index collection ; Step 3: Build a multi-source forgery-aware network, including: "expert mixture" noise extractor, noise stream network branch, RGB stream network branch, mask decoder and detector , and based on the index collection right Processing is performed to obtain the image forgery positioning result And image forgery detection results ; Step 4: Use formula (14) to establish the loss function of the multi-source forgery perception network : (14) In formula (14), , are two binary cross entropy losses, is the Dice loss, , are two weight factors; Step 5: Generate the mask gradient of the loss function using formula (15) : (15) In formula (15), is the Hadamard product, are the parameters of the network encoder, yes The gradient of Step 6: Use the ADAM optimizer to train the multi-source forgery perception network, and use formula (15) to optimize the loss function Update until the loss function The trained multi-source forgery-aware network model is obtained until convergence, which is used to process the image to be detected to obtain the predicted classification results of image forgery detection and the positioning map of the suspected forgery area.

2. According to claim 1, an image forgery detection and positioning method based on lottery hypothesis and mask autoencoder, characterized in that: The step 2 comprises: Step 2.1, the masked autoencoder is composed of Transformer-Base blocks, using weight parameters Initialize the pre-trained mask autoencoder and perform self-reconstruction training on the forged image dataset to obtain the forged weight parameter set of the trained mask autoencoder , and use formula (1) to calculate the forged change range set of weight parameters : (1) In formula (1), Represents the operation of finding the first K largest elements; Represents absolute value operation; Step 2.2: Use weight parameters Initialize the pre-trained mask autoencoder and perform self-reconstruction training on the real image dataset to obtain the real weight parameter set of the trained mask autoencoder , and use formula (2) to calculate the actual change range of the weight parameter : (2) Step 2.3: Use formula (3) to filter out weight parameters that are sensitive to forgery : (3) In formula (3), represents the difference operation, Represents the intersection operation; Step 2.4: Generate using formula (4) The i-th weight parameter in Gradient mask of , thus obtaining the weight parameter Gradient mask of : (4) Step 2.5: Gradient mask Compute the weighted mask rate of the kth Transformer-Base block in the masked autoencoder , thus obtaining The weight mask rate of the Transformer-Base block, and select the one with the highest weight mask rate Transformer-Base blocks constitute an index set ;in, Represents the i-th weight parameter in the k-th Transformer-Base block The mask value of is the total number of parameters in the kth Transformer-Base block; k= .

3. According to claim 2, a method for detecting and locating image forgery based on the lottery hypothesis and masked autoencoder, characterized in that: The step 3 comprises: The noise flow network branch in step 3 is composed of Transformer-Tiny blocks and their corresponding Linear layers; the RGB network stream branch consists of a linear layer and Transformer-Base blocks; is the number of Transformer-Tiny blocks in the noise stream branch; is the number of Transformer-Base blocks in the RGB stream branch; Step 3.1: Image After being processed by the "Expert Mixture" noise denoiser, we get the multi-source forged features. ; Step 3.2: Noisy network flow branch The Transformer-Tiny blocks forge features from multiple sources respectively. Process and obtain Noise flow intermediate features ,in, represents the intermediate feature of the i-th noise stream, for The number of image blocks in for The characteristic dimension of Noisy network flow branching Linear layers The intermediate features of the noise flow Process and obtain Noise Mapping Features ,in, represents the i-th linear layer, represents the i-th noise mapping feature; Step 3.3: Image After the linear layer processing of the RGB network flow branch, the initial RGB flow intermediate features are obtained , and then input the cascaded RGB network flow branch The jth RGB stream intermediate feature is obtained by using formula (11) , and then get the RGB stream intermediate features : (11) In formula (11), represents the intermediate feature of the j-1th RGB stream, for The number of image blocks in for The characteristic dimension of represents the j-1th Transformer-Base block, represents the kth noise mapping feature; Step 3.4: The mask decoder consists of several convolutional layers, batch normalization, ReLU functions, and upsampling layers. RGB stream intermediate features Processing is performed to obtain the image forgery positioning result ; Step 3.5: Detector Using formula (12) to calculate the positioning result Convert and get the image forgery detection result : (12) In formula (12), Represents the cascade operation of the first convolution, ReLU function, batch normalization, the second convolution, and the Sigmoid function. represents a hyperparameter that decreases with the number of training iterations, represents the generalized average pooling operation, and has: (13) In formula (13), p is the parameter to be trained, for The pixel with coordinates (i, j) in the image.

4. A method for detecting and locating image forgery based on the lottery hypothesis and masked autoencoder according to claim 3, characterized in that: The "expert hybrid" noise enhancer in step 3.1 includes: SRM filter, Bayar convolution, noise watermark extraction module, expert weight generation module and second convolution layer ; Among them, the expert weight generation module consists of the first convolutional layer, the pooling layer, and the linear layer; Step 3.1.1, Image After being processed by the SRM filter, Bayar convolution and noise watermark extraction module, the SRM noise feature is obtained. , Bayar noise characteristics and noise watermark features ; Step 3.1.2, Image Input into the expert weight generation module for processing to obtain the expert coefficient matrix : Step 3.1.2.1: Generate channel descriptor using formula (5) : (5) In formula (5), represents the pixel value of the i-th row and j-th column in X; Step 3.1.2.2: Use formula (6) to get the first weight matrix ,in, Represents the dimensions of the first weight matrix: (6) In formula (6), is a linear layer, Pool is a pooling layer, is the first convolutional layer; Step 3.1.2.3: Use formula (7) to get the expert coefficient matrix , where O is the number of experts: (7) In formula (7), represents the ReLU function; is the parameter matrix to be learned; Step 3.1.3, second convolutional layer Using formula (8) to obtain multi-source forgery features : (8) In formula (8), Indicates splicing.

5. An electronic device, comprising a memory and a processor, characterized in that: The memory is used to store a program that supports the processor to execute the image forgery detection and positioning method according to any one of claims 1 to 4, and the processor is configured to execute the program stored in the memory.

6. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the image forgery detection and positioning method according to any one of claims 1 to 4 are executed.

Citation Information

Patent Citations

  • Building edge optimization method based on multi-task learning and dual lottery hypothesis

    CN116052006A

  • Image forgery detecting and positioning method based on noise auxiliary prompt learning

    CN119295383A

  • Method and apparatus with image preprocessing

    US20230030937A1

Cited By

  • Lottery recognizer self-adaptive adjustment method and system based on image recognition

    CN121747135A