Method for training forged image detection model based on key forgetting mechanism
By introducing key forgetting modules into the forged image detection model, the key pixels in the attribute feature map are weakened, and the problem of poor detection results in the existing technology when facing new forgery means is solved, achieving higher detection accuracy and generalization capabilities.
Patent Information
- Application Number
- CN202510147642.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-02-10
AI Technical Summary
When detecting forged images, the accuracy of the prior art depends on the richness of the types of forged technologies in the training set, and the detection effect is poor when facing new forgery methods.
The training method of the forged image detection model based on the key forgetting mechanism is adopted, and the key pixels in the attribute feature map are weakened through the key forgetting module, forcing the model to learn general forged trace features.
The model's detection effect of different types of forged images is improved, the model's generalization ability is improved, and new methods of forgery can be better faced.
Smart Images

Figure CN120125932A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification belong to the technical field of image processing, and particularly relate to a training method for a forged image detection model based on a key forgetting mechanism. Background Art
[0002] With the development of Artificial Intelligence (AI) technology, content generation technology based on AI technology has become increasingly mature, and the content generated by it has become increasingly realistic and difficult to distinguish its authenticity. For example, Deepfake technology is a technology that generates seemingly real fake images or fake videos through AI technology. Malicious parties may use Deepfake technology to generate forged images or forged videos for malicious acts such as forging others' identities, creating fake news, and fraud. Therefore, a method for detecting forged images is needed.
[0003] The accuracy of related technologies in detecting forged images often depends on the richness of the types of forgery technologies covered in the training set. When new forgery methods appear, the detection effect is not good. Summary of the Invention
[0004] The purpose of the present invention is to provide a training method for a forged image detection model based on a key forgetting mechanism and a forged image detection method, which can improve the detection effect of the model when facing different types of forged images by forcing the model to learn the characteristics of common forgery traces.
[0005] The first aspect of this specification provides a training method for a forged image detection model based on a key forgetting mechanism. The forged image detection model at least includes an attribute extractor and a key forgetting module, including:
[0006] Obtain a training sample set, where the training sample set contains multiple sample images, and the sample images have true / false labels, and the true / false labels are used to label whether the sample images are forged images;
[0007] The attribute extractor extracts attributes based on the sample images to obtain an attribute feature map, and the attribute feature map is used to represent information related to the forgery traces in the sample images;
[0008] The key forgetting module obtains the influence degree of each pixel in the attribute feature map on the prediction result, determines the key pixels in the attribute feature map based on the influence degree of each pixel, and weakens the influence degree of the key pixels;
[0009] The classifier in the key forgetting module performs true / false prediction on the attribute feature map after weakening processing to obtain a prediction label;
[0010] Train a forged image detection model based on the differences between the predicted labels and the true / false labels.
[0011] The second aspect of this specification provides a forged image detection method based on a key forgetting mechanism, including:
[0012] Obtain a target image to be detected and a forged image detection model trained by the method according to the first aspect. The forged image detection model includes at least an attribute extractor and a key forgetting module;
[0013] The attribute extractor extracts attributes based on the target image to obtain a target attribute feature map, which is used to represent information related to forged traces in the target image;
[0014] The classifier in the key forgetting module makes a true / false prediction on the target attribute feature map to obtain a detection result, which is used to indicate whether the target image is a forged image.
[0015] The third aspect of this specification provides a computing device, including a memory and a processor. Among them, executable code is stored in the memory, and when the processor executes the executable code, it implements the method described in the first aspect or the second aspect.
[0016] The training method of the forged image detection model based on the key forgetting mechanism provided by the technical solution of the embodiments of this specification, through the key forgetting module, weakens the key pixels in the attribute feature map based on the influence degree of each pixel in the attribute feature map on the prediction result, so as to force the model to forget the key parts related to forgery in the attribute feature map during the training stage. These key parts correspond to the types of forgery techniques used in the sample images in the training sample set, so as to be able to force the model to learn the features of more common and general forged traces, improve the generalization ability of the model, and the model can achieve a better detection effect when facing different types of forged images. Description of the Drawings
[0017] To more clearly illustrate the technical solutions of the embodiments of this specification, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0018] Figure 1 It is a schematic flowchart of a training method of a forged image detection model based on a key forgetting mechanism in an embodiment of this specification;
[0019] Figure 2 It is a schematic diagram of a forgery detection model in an embodiment of this specification;
[0020] Figure 3 It is a schematic diagram of a forgery detection model in an embodiment of this specification;
[0021] Figure 4 It is a schematic diagram of a forgery detection model in an embodiment of this specification;
[0022] Figure 5 It is a schematic diagram of a forgery detection model in an embodiment of this specification;
[0023] Figure 6 It is a schematic architecture diagram of a forgery detection model in an embodiment of this specification;
[0024] Figure 7 It is a schematic flowchart of a training method for another forgery image detection model based on a key forgetting mechanism in an embodiment of this specification;
[0025] Figure 8 It is a schematic flowchart of a forgery image detection method based on a key forgetting mechanism in an embodiment of this specification;
[0026] Figure 9 It is a schematic diagram of a training device for a forgery image detection model based on a key forgetting mechanism in an embodiment of this specification;
[0027] Figure 10 It is a schematic diagram of a forgery image detection device based on a key forgetting mechanism in an embodiment of this specification. Detailed implementation manners
[0028] In order to enable those skilled in the art of this technology to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the scope of protection of this specification.
[0029] As mentioned above, with the development of artificial intelligence, the images forged by Deepfake technology are becoming more and more realistic. The forgery detection models used in current forgery detection methods focus on learning the specific forgery types included in the training set, which makes the generalization performance of the trained models poor and the detection performance for other forgery types outside the training set unsatisfactory. However, it can be foreseen that with the continuous progress of Deepfake technology, more complex and diverse forgery means will be mixed in forged images. Therefore, it is urgent to improve the generalization performance of current forgery detection models.
[0030] Based on the above analysis, considering that the features extracted during the training of the current forgery detection model can easily capture the specific forgery traces in the training set, which makes the model tend to learn the specific forgery types included in the training set, the embodiments of this specification propose a training method for a forgery image detection model based on a key forgetting mechanism, forcing the classifier of the model to forget the significant features captured during training. These significant features are often related to specific forgery types in the training set, so that the model can expose the features of more common and general forgery traces that were previously ignored. These features are widely present in various types of forged images, greatly improving the generalization ability of the model.
[0031] The following combines Figure 1 , and describes the training process of the forgery image detection model in the embodiments of this specification. Among them, Figure 1 is the flowchart of the training method for the forgery image detection model based on the key forgetting mechanism in the embodiments of this specification. This training process can be executed by any device, platform or cluster of devices with computing and processing capabilities, including the following steps S101 - S105. As Figure 2 shown in the schematic diagram of the network structure of the forgery detection model, the forgery detection model in the embodiments of this specification at least includes an attribute extractor and a key forgetting module.
[0032] As Figure 1 shown, in step S101, a training sample set is obtained. The training sample set contains multiple sample images, and the sample images have true / false labels, and the true / false labels are used to label whether the sample images are forged images.
[0033] In practice, the multiple sample images in the training sample set include real images and forged images. Among them, the forged images can be images forged through image editing, deep forgery technologies such as deep forgery networks, etc. The true / false labels of real images can be labeled as "real", and the true / false labels of forged images can be labeled as "fake".
[0034] In this embodiment, there is no restriction on the objects forged in the forged images. For example, the forged images can contain forged or replaced faces, forged scenes, forged actions, or forged objects, etc.
[0035] Next, in step S102, the attribute extractor extracts attributes based on the sample images to obtain an attribute feature map, and the attribute feature map is used to represent information related to the forgery traces in the sample images.
[0036] The feature information in the sample images can be divided into information related to forgery traces and information unrelated to forgery traces. Among them, the Attribute Extractor is used to separate the information related to forgery traces in the sample images.
[0037] This embodiment does not limit the network structure of the attribute extractor. For example, it can adopt structures such as convolutional neural networks and deep autoencoders.
[0038] Next, in step S103, the key forgetting module obtains the influence degree of each pixel in the attribute feature map on the prediction result, determines the key pixels in the attribute feature map based on the influence degree of each pixel, and weakens the influence degree of the key pixels.
[0039] Among them, each pixel in the attribute feature map represents the forgery traces captured from the sample image. For example, when the image is spliced or tampered with, the edges at the splicing points may not be smooth, resulting in the pixel values of some pixels in the attribute feature map reflecting abnormal edges or inconsistent textures in that area. The prediction result refers to the result output by the classifier in the key forgetting module for authenticity prediction.
[0040] Since the specific forgery traces in the training sample set are easily captured, that is, compared with the hidden common forgery traces, the specific forgery traces on the sample images in the training sample set have a relatively greater impact on the prediction result. Therefore, the key pixels in the attribute feature map related to the specific forgery traces in the training sample set can be determined through the influence degree of each pixel in the attribute feature map on the prediction result during the training process. The influence degree on the prediction result specifically refers to the influence degree on the accuracy of the prediction result when the numerical value of the channel value corresponding to the pixel changes. For example, when the sample image is an RGB image, the channel value corresponding to the pixel can refer to the values of each pixel in the image on different color channels (Red, Green, Blue).
[0041] In practice, the influence degree of each pixel in the attribute feature map on the prediction result can be quantified as the gradient corresponding to each pixel in the attribute feature map during the backpropagation process. The magnitude of the influence degree is presented by the magnitude of the gradient. The larger the gradient, the greater the influence degree. Since the specific forgery traces in the training sample set are easily captured, they usually appear as some of the pixels with the largest gradients among the pixels corresponding to the attribute feature map during the backpropagation process.
[0042] In one implementation, this step can specifically be to determine the gradient corresponding to each pixel in the attribute feature map during the backpropagation process based on the difference between the prediction result and the authenticity label. For each pixel in the attribute feature map, if the gradient of the pixel is higher than the preset threshold, the pixel is determined as a key pixel, and the value of the channel corresponding to the key pixel in the attribute feature map is set to a weakening value.
[0043] Specifically, the prediction result refers to the prediction result output by the classifier of the key forgetting module during the forward propagation of the previous iteration of training. The loss function is calculated based on the difference between the prediction result and the true / false label. The gradients corresponding to each pixel used in this embodiment represent the derivatives of the loss function with respect to each neuron input (i.e., each pixel in the attribute feature map can be regarded as the input of each corresponding neuron in this layer of the network). Assume, by way of example, that the calculation formula for the gradients G corresponding to each pixel in the attribute feature map is as follows:
[0044]
[0045] where N is the number of sample images in a batch in the training sample set, y i is the true / false label corresponding to the i-th sample image, p i is the prediction result of the classifier of the key detection module corresponding to the i-th sample image, F a is the attribute feature map, and the loss function used is the BCE (Binary Cross-Entropy) loss function, that is, -[y i log(p i )+(1 - y i )log(1 - p i )]. This formula can obtain the gradients corresponding to each pixel by taking the partial derivative of each pixel in F a through the chain rule.
[0046] It should be noted that the gradient G used in this embodiment is not exactly the same as the gradient of the loss function with respect to each parameter (such as the weight w) used in the gradient descent during the backpropagation process. The relationship between the two can be regarded as in the chain rule For each pixel in the attribute feature map F
[0047] a can be sorted according to the magnitude of the gradient, and then, through a set preset threshold, the pixels in the attribute feature map are filtered, and the values of the channels corresponding to the key pixels are set to weakening values to forcefully forget some pixels with greater influence in the attribute feature map. By way of example, this weakening process can be achieved by using a mask to mask the attribute feature map. The process of masking is as shown in the following formula:
[0048]
[0049] where Sort represents sorting each pixel in F a according to the magnitude of the gradient G, represents the dot product operation, and F′ aRepresents the attribute feature map after weakening processing, also known as the forgotten feature map, M σ is a 0-1 matrix, which can be regarded as a mask. σ is a hyperparameter used to control the forgetting ratio. θ is a preset threshold selected from the sorted gradients according to the hyperparameter σ. i, j represent the positions of pixels in the attribute feature map. Represents the mask M σ The value at position i, j in it. For example, M σ=0.3 Indicates that in the attribute feature map, the pixels with the top 30% in terms of influence degree (gradient) are determined as key pixels, and the channel values corresponding to the key pixels are set to the weakening value 0, while the channel values corresponding to the remaining pixels in the attribute feature map remain unchanged. In this way, by masking the channels with gradients exceeding the preset threshold. In other examples, 0 in formula (3) can also be set to other values between 0 and 1 to weaken the influence degree of key pixels on the prediction result.
[0050] In other implementation manners, for the weakening processing of key pixels, it can also be an enhancement processing of other pixels except key pixels in the attribute feature map. For example, set 0 in formula (3) to 1 and set 1 to other values greater than 1 to enhance the influence degree of other pixels except key pixels on the prediction result.
[0051] Continuing, in step S104, the classifier in the key forgetting module performs authenticity prediction on the attribute feature map after weakening processing to obtain a prediction label.
[0052] In this step, the attribute feature map F a ' after weakening processing is input into the classifier, and the prediction label output by the classifier is obtained. The classifier is used to identify which label {fake, real} the attribute feature map after weakening processing belongs to. The prediction label represents the probability that the sample image belongs to a forged image or a real image. For example, the prediction label can be a value between 0 and 1. The closer the prediction label is to 0, the more likely the model thinks the sample image is a forged image. The closer the prediction label is to 1, the more likely the model thinks the sample image is a real image.
[0053] To avoid the classifier being insensitive to the forged types contained in the training sample set due to excessive forgetting in the key forgetting module, thus affecting the detection performance of the model for the forged types in the training sample set. As an implementation manner, before using the classifier for prediction, the attribute feature map and the attribute feature map after weakening processing can be concatenated, and authenticity prediction is performed on the concatenated attribute feature map to obtain a prediction label. Connect the attribute feature map F a with the forgotten feature map F a ' after weakening processing as the input of the classifier, instead of using the forgotten feature map F aPerform forgery classification so that the model can learn the characteristics of specific forgery types included in the training sample set while learning the common forgery traces that are overlooked. Additionally, this processing actually generates two attribute feature maps with different detection difficulties for each sample image, and this approach can also be regarded as a feature-level data augmentation strategy.
[0054] Finally, in step S105, train the forgery image detection model based on the difference between the predicted label and the true / false label.
[0055] Specifically, according to the difference between the predicted label and the true / false label, a loss function can be calculated, and the forgery image detection model is trained through the loss function.
[0056] This embodiment does not limit the specific loss function to be used. After determining the loss function, according to this loss function, the parameters of the forgery image detection model can be adjusted through backpropagation. When the network iteration reaches the training end condition, the network training ends, and this embodiment does not limit the training end condition. Among them, this condition can be that the iteration reaches a certain number of times, or the value of the loss function is less than a certain threshold.
[0057] Exemplarily, the loss function can use a classification loss function, such as the BCE loss function. When using the spliced attribute feature map F a and the forgetting feature map F a ' as the input of the classifier, the loss function can be expressed as follows:
[0058]
[0059] Among them, Con represents the splicing operation, BCE represents the BCE loss function, and y ∈ {fake, real} is the true / false label of the sample image.
[0060] The training method of the forgery image detection model based on the key forgetting mechanism provided by the embodiments of this specification weakens the key pixels in the attribute feature map based on the influence degree of each pixel in the attribute feature map on the prediction result through the key forgetting module, so as to force the model to forget the key parts related to forgery in the attribute feature map during the training phase. These key parts correspond to the forgery technique types used in the sample images in the training sample set, so as to be able to force the model to learn the characteristics of more common and general forgery traces, avoid the model overfitting the sample data in the training set, improve the generalization ability of the model, and the model can achieve better detection effects when facing different types of forgery images.
[0061] In the context of forged image detection, the purpose of decoupling is to separate the forgery-related attribute information and the forgery-unrelated content information from the original image, so as to effectively identify forged images through the attribute information. For example, in the current state-of-the-art forgery detection method based on decoupled representation, attribute features and content features are respectively extracted for a pair of real images and forged images, and then cross-reconstruction is performed using the content features of one image and the attribute features of the other image to ensure that the attribute features are related to forgery while the content features are unrelated to forgery. In addition to the poor generalization performance of the trained model, this method also has some other problems. For example, using information-intensive images as the decoupling target increases the difficulty of decoupling; this method assumes that the information that does not affect reconstruction is related to forgery, and the extracted attribute features are more inclined to be unrelated to image reconstruction rather than related to forgery information.
[0062] As an implementation, as Figure 3 shown, the forged image detection model further includes an Encoder. To reduce the difficulty of decoupling, the Encoder can be used to convert the sample image into a feature representation, and then decoupling is performed based on the feature representation instead of the sample image. In step 102, it can be the Encoder that extracts features from the sample image to obtain a feature representation; and the Attribute Extractor extracts attributes from the feature representation to obtain an attribute feature map. Compared with the dense and complex information volume in the sample image, the information volume contained in the feature representation is less, which is more conducive to the separation of the attribute feature map related to forgery.
[0063] As an implementation, as Figure 4 shown, the forged image detection model further includes a Content Extractor; this method may further include: the Content Extractor extracts content from the feature representation to obtain a content feature map, and the content feature map is used to represent the information unrelated to the forgery traces in the sample image. To make the separation of attribute features from the sample image more inclined to be related to forgery rather than just unrelated to image reconstruction and avoid the information unrelated to forgery from dominating the attribute features, step S105 can be to train the forged image detection model based on the difference between the prediction label and the authenticity label and the irrelevance between the content feature map and the attribute feature map.
[0064] In practice, the loss function can include not only the classification loss function calculated based on the difference between the prediction label and the authenticity label, but also the irrelevance loss function calculated based on the irrelevance between the content feature map and the attribute feature map. The stronger the irrelevance between the content feature map and the attribute feature map, the smaller the value of the irrelevance loss function, and the weaker the irrelevance between the two, the larger the value of the irrelevance loss function. By training, the value of the loss function is continuously reduced to make the content feature map and the attribute feature map as separated as possible.
[0065] For example, when calculating irrelevance, the Kullback-Leibler Divergence, Spearman's Rank Correlation, etc. can be used to calculate the irrelevance between the content feature map and the attribute feature map. This embodiment does not limit the specific calculation method adopted.
[0066] Exemplarily, in the decoupling stage, the Pearson correlation coefficient can be introduced and used to distinguish the attribute feature map related to forgery and the content feature map unrelated to forgery. This makes it easier for the decoupler to isolate the content feature map and the attribute feature map. The classification loss and the Pearson correlation coefficient loss will force the information related to forgery to dominate the attribute feature map, and the information unrelated to forgery will gradually transfer to the content feature map.
[0067] As an implementation, as Figure 5 shown, the forgery image detection model further includes a Decoder. To ensure that the information in the feature representation is completely and separately extracted into the attribute feature map and the content feature map, the method may further include: inputting the attribute feature map and the content feature map into the decoder for reconstruction to obtain a reconstructed image; step S105 may be to train the forgery image detection model based on the difference between the reconstructed image and the sample image and the difference between the predicted label and the true / false label.
[0068] By using the decoder to restore the attribute feature map and the content feature map to a reconstructed image, it can be ensured that the attribute feature map and the content feature map are a complete representation of the sample image, ensuring that the information in the sample image is fully utilized. Specifically, based on the difference between the reconstructed image and the sample image, a reconstruction loss function can be calculated. The smaller the difference between the reconstructed image and the sample image, the smaller the value of the reconstruction loss function. Conversely, the larger the difference between the reconstructed image and the sample image, the larger the value of the reconstruction loss function. Using the reconstruction loss function and the classification loss function as an overall loss function for model training makes the reconstructed image closer to the sample image during the training process.
[0069] Considering that the single-level decoupling method cannot capture enough discriminative information for separating the attribute feature map and the content feature map, the embodiments of this specification propose another training method for the forgery image detection model based on the key forgetting mechanism. This method combines the multi-level decoupling method to separate the attribute feature map and the content feature map, so that the extracted attribute feature map can better converge the information related to forgery in the image.
[0070] Next, combined with Figure 6 the framework diagram of the forgery image detection model, the complete process of model training will be introduced.Figure 7 It is another flowchart of the training method of the forged image detection model based on the key forgetting mechanism in the embodiments of this specification. This method can be executed by any device, platform or cluster of devices with computing and processing capabilities, including steps S701 - S708 shown below.
[0071] As Figure 7 shown, in step S701, a training sample set is obtained. The training sample set contains multiple sample images, and the sample images have authenticity labels, where the authenticity labels are used to label whether the sample images are forged images.
[0072] Different from the existing forged image detection methods based on decoupled representation that require inputting paired images, the method in the embodiments of this specification inputs a single sample image x into the model each time.
[0073] Next, in step S702, the encoder extracts features from the sample image to obtain feature representations at multiple levels, and the feature representations at different levels are used to reflect image feature information in different aspects.
[0074] Among them, the encoder E n can split the dense information in the sample image, and the extracted feature representations f i can independently focus on image feature information in different aspects such as facial contours, low - frequency information, textures, colors, or edges. Exemplarily, the output of the encoder E n can be as follows:
[0075] f 1 ,…,f n =E n (x) (5)
[0076] Among them, f i ,i∈[1,…,n] represents n feature representations output by different - level networks in the encoder. For example, the feature representation output by the first - layer network close to the input end in the encoder is f 1 , and the feature representation output by the second - layer network is f 2 .
[0077] Next, in step S703, for the feature representation of each level, the corresponding attribute extractor of this level extracts attributes from the feature representation to obtain the initial attribute feature map corresponding to this level. For the feature representation of each level, the corresponding content extractor of this level extracts content from the feature representation to obtain the initial content feature map corresponding to this level.
[0078] In this embodiment, the forgery representation decoupling is extended from single-level to multi-level. The forgery image detection model includes attribute extractors, content extractors, and decoders at multiple levels, and the number of attribute extractors, content extractors, and decoders is the same. As Figure 6 shown, the encoder extracts 3-layer feature representations, and the model includes 3 attribute extractors, content extractors, and decoders.
[0079] The decoupler at the i-th level includes an attribute extractor and a content extractor Exemplarily, they can have the same network structure but do not share parameters. Among them, the content extractor is used to extract the initial content feature map that has nothing to do with forgery The attribute extractor is used to separate the initial attribute feature map related to forgery As shown in the following formula:
[0080]
[0081] As an implementation, the Pearson correlation coefficient can be used as the separation constraint between the initial content feature map and the initial attribute feature map to calculate their irrelevance. The smaller the value of the Pearson correlation coefficient, the greater the difference in direction and amplitude between the two feature maps. A large enough difference can avoid the possible correlation between the attribute feature and the content feature. The Pearson correlation coefficient loss function for each level can be expressed as:
[0082]
[0083] Among them, Cov represents the covariance between the initial attribute feature map and the initial content feature map, and σ represents the variance. Through the constraint of the Pearson correlation coefficient loss function, through iterative training, the information related to forgery can be made to dominate in the attribute feature map, and the information unrelated to forgery can be made to dominate in the content feature map.
[0084] As an implementation, a corresponding classifier can be connected after each initial attribute feature map to ensure its relevance to forgery and further enhance the correlation between the attribute feature map and the forgery information. The forgery image detection model also includes classifiers corresponding to multiple levels; for the initial attribute feature map of each level, the corresponding classifier of that level makes a true / false prediction on the initial attribute feature map to obtain an initial prediction label.
[0085] Exemplarily, the initial prediction label obtained by the classifier corresponding to each level predicting the initial attribute feature map of that level, and the initial classification loss function corresponding to each level is calculated based on the difference between the initial prediction label and the true / false label as shown below:
[0086]
[0087] Among them, BCE is the cross-entropy loss function, and y ∈ {fake, real} is the true / false label of the sample image x.
[0088] Next, in step S704, an attribute feature map is obtained based on multiple initial attribute feature maps corresponding to multiple levels respectively.
[0089] As Figure 6 shown, and can be connected to obtain the attribute feature map F a . In other embodiments, it may also be to weight and then splice multiple initial attribute feature maps to obtain the attribute feature map.
[0090] In step S705, for each level, the initial attribute feature map and the initial content feature map corresponding to this level are input into the decoder corresponding to this level for reconstruction to obtain the reconstructed feature or the reconstructed image.
[0091] Among them, the purpose of the decoder is to ensure that the decoupled initial attribute feature map and the initial content feature map are complete representations of the sample image or the feature representation. Different constraints can be imposed on the decoupling of different levels to extend the reconstruction loss from the image level to the feature level, so as to extend the decoupling of the forged representation from a single level to multiple levels, rather than only using the reconstruction constraint at the image level.
[0092] Specifically, each level of feature representation has a corresponding decoder The decoder is designed to ensure that the information in the initial attribute feature map and the initial content feature map is a complete representation of the feature representation at this level, rather than the sample image, unless i = 1. That is, in the case of i = 1, the initial attribute feature map and the initial content feature map are input into the encoder of the first level for reconstruction to obtain the reconstructed image. In the case of i > 1, the initial attribute feature map and the initial content feature map are input into the encoder of the corresponding level for reconstruction to obtain the reconstructed feature. That is to say, except that the reconstruction target of the first level is the sample image, the reconstruction targets of other levels are the feature representations at these levels. The reconstruction loss function calculated based on the difference between the reconstruction features and the feature representations of multiple levels, and the reconstruction loss function calculated based on the difference between the reconstructed image and the sample image can be expressed as:
[0093]
[0094] Among them, Mse represents the Mean Square Error loss function.
[0095] In addition, in step S706, the key forgetting module obtains the influence degree of each pixel in the attribute feature map on the prediction result, determines the key pixels in the attribute feature map based on the influence degree of each pixel, and weakens the influence degree of the key pixels.
[0096] For the explanation of step S706, reference can be made to the relevant description of step S103 in the previous text, which will not be elaborated here.
[0097] In step S707, the classifier in the key forgetting module performs authenticity prediction on the weakened attribute feature map to obtain a prediction label.
[0098] As Figure 6 shown in the key forgetting module, after obtaining the forgotten feature map F a ′ through weakening processing, the attribute feature map F a can be concatenated with the forgotten feature map F a ′ to obtain the final attribute feature map F a ″, and the classifier is used to predict F a ″ to obtain a prediction label.
[0099] For the explanation of step S707, reference can be made to the relevant description of step S104 in the previous text, which will not be elaborated here.
[0100] In step S708, the forgery image detection model is trained based on the difference between the prediction label and the authenticity label.
[0101] Among them, the classification loss function calculated based on the difference between the prediction label and the authenticity label can be as shown in formula (4). There are multiple implementation methods for this step. Different overall loss functions can be calculated based on loss functions from different aspects.
[0102] As one implementation method, it can be to train the forgery image detection model based on the difference between the reconstructed image and the sample image, the difference between multiple levels of reconstructed features and feature representations, and the difference between the prediction label and the authenticity label.
[0103] Among them, the reconstruction loss function calculated based on the difference between the reconstructed image and the sample image, and the difference between multiple levels of reconstructed features and feature representations can be as shown in formula (9). It can be to perform weighted summation on the classification loss function and the reconstruction loss function to obtain the overall loss function, and train the forgery image detection model based on the overall loss function.
[0104] As one implementation method, it can be to train the forgery image detection model based on the difference between the prediction label and the authenticity label, and the difference between the initial prediction label and the authenticity label.
[0105] Among them, the initial classification loss function corresponding to each level calculated based on the differences between the initial prediction labels and the true / false labels at multiple levels can be as shown in formula (8). It can be a weighted sum of the classification loss function and the initial classification loss function to obtain the overall loss function, and the forgery image detection model is trained based on the overall loss function.
[0106] As an implementation, it can be to train the forgery image detection model based on the differences between the prediction labels and the true / false labels, as well as the irrelevance between the content feature maps and the initial attribute feature maps at multiple levels.
[0107] Among them, the irrelevance loss function calculated based on the irrelevance between the content feature maps and the initial attribute feature maps at multiple levels can be as shown in formula (7). It can be a weighted sum of the classification loss function and the irrelevance loss function to obtain the overall loss function, and the forgery image detection model is trained based on the overall loss function.
[0108] As an implementation, contrastive learning can also be introduced to further improve the generalization performance. For multiple sample images with the same true / false labels, calculate the intra-class similarity between the weakened attribute feature maps corresponding to the multiple sample images respectively; for multiple sample images with different true / false labels, calculate the inter-class similarity between the weakened attribute feature maps corresponding to the multiple sample images respectively. Through contrastive learning, maximize the intra-class similarity s p , minimize the inter-class similarity s n , where the similarity can be measured by the distance between the weakened attribute feature maps F a ′, so that the features within the class are closer and the features between classes are farther. Exemplarily, given a batch in a training sample set with K pairs of same-class samples (i.e., sample image pairs with the same true / false labels) and L pairs of different-class samples (i.e., sample image pairs with different true / false labels), the contrastive learning loss function can be expressed as:
[0109]
[0110] Among them, γ is the scaling factor and m is the threshold for similarity.
[0111] Train the forgery image detection model based on the intra-class similarity, inter-class similarity, and the differences between the prediction labels and the true / false labels. It can be a weighted sum of the classification loss function and the contrastive learning loss function to obtain the overall loss function, and the forgery image detection model is trained based on the overall loss function.
[0112] As an implementation, the overall loss function for the training process is the weighted sum of the four loss functions listed above and the classification loss function:
[0113]
[0114] Among them, λ 1 , λ 2 , λ 3 and λ 4 are hyperparameters used to measure the impact of each loss function on the overall loss function.
[0115] In other implementation manners, the above loss functions can also be combined in other forms to obtain the overall loss function.
[0116] For the training method of the forged image detection model based on the key forgetting mechanism in the embodiments of this specification, first, image feature information focusing on different aspects in the image is used as the decoupling target instead of the image itself to simplify the difficulty of decoupling. Then, the decoupling of the forged representation is extended to the decoupling of multi-level feature representations, and the Pearson correlation coefficient is introduced to better separate the attribute feature map from the content feature map, ensuring that the attribute feature map is as relevant to forgery as possible. Subsequently, a key forgetting mechanism is used to force the model to pay more attention to the common subtle forgery traces that have been ignored before, thereby enhancing the generalization ability of the model to unseen forgery types.
[0117] Next, the application stage of the forged image detection model will be described in conjunction with Figure 8 FIG. Figure 8 FIG. is a flowchart of a forged image detection method based on the key forgetting mechanism proposed in the embodiments of this specification. This method can be executed by any device, platform or device cluster with computing and processing capabilities, including steps S801 - S803 shown below.
[0118] As shown in Figure 8 FIG., in step S801, a target image to be detected and a forged image detection model trained by the method according to any one of the above embodiments are obtained. The forged image detection model includes at least an attribute extractor and a key forgetting module.
[0119] Among them, the target image to be detected is a single image, which can specifically be an image that may be forged or a video frame in a video that may be forged.
[0120] In step S802, the attribute extractor extracts attributes based on the target image to obtain a target attribute feature map, and the target attribute feature map is used to represent information related to forgery traces in the target image.
[0121] Among them, the attribute extractor is used to separate information related to forgery traces in the target image.
[0122] As an implementation manner, as shown in Figure 4As shown, the forged image detection model further includes an encoder. To reduce the difficulty of decoupling, the encoder can be used to convert the target image into a target feature representation, so that decoupling is performed based on the target feature representation rather than the target image. In this step, the encoder can extract features from the target image to obtain the target feature representation, and the attribute extractor can extract attributes from the feature representation to obtain the target attribute feature map. Compared with the dense and complex information in the target image, the information contained in the feature representation is less, which is more conducive to separating the target attribute feature map related to forgery.
[0123] As an implementation method, a multi-level decoupling method can be adopted to improve the correlation between the target attribute feature map obtained by decoupling and the forgery information. When the encoder extracts features from the target image, multiple levels of target feature representations can be obtained. The feature representations at different levels are used to reflect different aspects of the image feature information. For each level of the target feature representation, the corresponding attribute extractor at that level extracts attributes from the target feature representation to obtain the initial target attribute feature map corresponding to that level, and based on the multiple target initial attribute feature maps corresponding to multiple levels respectively, the target attribute feature map is obtained.
[0124] In step S803, the classifier in the key forgetting module performs authenticity prediction on the target attribute feature map to obtain the detection result, and the detection result is used to indicate whether the target image is a forged image.
[0125] In practice, different from the training stage, in the key forgetting module, the target attribute feature map does not need to be weakened, but directly uses the classifier in the key forgetting module to perform authenticity prediction on the target attribute feature map to obtain the detection result. The value output by the classifier can be a value within 0 to 1, and the size of the value represents the probability that the target image belongs to a forged image or a real image. For example, the closer the value is to 0, the more likely the model thinks the target image is a forged image, and the closer the value is to 1, the more likely the model thinks the target image is a real image. A threshold can be set, such as 0.7. When the value is greater than 0.7, it is determined that the detection result is that the target image is a real image, and when the value is less than 0.7, it is determined that the detection result is that the target image is a forged image.
[0126] As an implementation, a corresponding classifier can be connected after each of the above-mentioned target initial attribute feature maps. For the initial target attribute feature maps at each level, the corresponding classifier at that level predicts the authenticity of the initial target attribute feature map to obtain an initial detection result. The initial detection result corresponding to each level can include the probability that a target image belongs to a forged image or a real image. It can be to perform weighted summation on the initial detection result and each probability in the detection result to obtain a final probability value, and based on this probability value, determine whether the target image is a forged image, so as to further improve the accuracy of forged image detection.
[0127] The forged image detection training method based on the key forgetting mechanism provided by the technical solution of the embodiments of this specification. The forged image detection model used at least includes an attribute extractor and a key forgetting module. Since the model learns how to identify the features of more common and general forgery traces during the training phase, and at the same time, the significant features extracted are still retained in the target attribute feature map, it can achieve better detection effects for different types of forged images.
[0128] Figure 9 It is a schematic structural diagram of a training device for a forged image detection model based on the key forgetting mechanism in the embodiments of this specification. The forged image detection model at least includes an attribute extractor and a key forgetting module. This device can be applied to any device, platform or device cluster with computing and processing capabilities. The device includes:
[0129] A sample acquisition unit 901, configured to acquire a training sample set. The training sample set contains multiple sample images, and the sample images have authenticity labels, and the authenticity labels are used to label whether the sample images are forged images;
[0130] A feature extraction unit 902, configured to perform attribute extraction on the sample images by the attribute extractor to obtain attribute feature maps, and the attribute feature maps are used to represent information related to forgery traces in the sample images;
[0131] A weakening processing unit 903, configured to obtain the influence degree of each pixel in the attribute feature map on the prediction result by the key forgetting module, determine the key pixels in the attribute feature map based on the influence degree of each pixel, and perform weakening processing on the influence degree of the key pixels;
[0132] A prediction classification unit 904, configured to perform authenticity prediction on the weakened attribute feature map by the classifier in the key forgetting module to obtain a prediction label;
[0133] An iterative training unit 905, configured to train the forged image detection model based on the difference between the prediction label and the authenticity label.
[0134] In some embodiments, the forged image detection model further includes an encoder; a feature extraction unit 902, specifically configured to: extract features from a sample image by the encoder to obtain a feature representation; and extract attributes from the feature representation by an attribute extractor to obtain an attribute feature map.
[0135] In some embodiments, the forged image detection model further includes a content extractor; the feature extraction unit 902 is further configured to: extract content from the feature representation by the content extractor to obtain a content feature map, where the content feature map is used to represent information unrelated to forged traces in the sample image;
[0136] An iterative training unit 905, specifically configured to: train the forged image detection model based on the difference between the predicted label and the true / false label and the irrelevance between the content feature map and the attribute feature map.
[0137] In some embodiments, the forged image detection model further includes a decoder, and the feature extraction unit 902 is further configured to: input the attribute feature map and the content feature map into the decoder for reconstruction to obtain a reconstructed image; the iterative training unit 905, specifically configured to: train the forged image detection model based on the difference between the reconstructed image and the sample image and the difference between the predicted label and the true / false label.
[0138] In some embodiments, the forged image detection module includes attribute extractors of multiple levels; the feature extraction unit 902 is specifically configured to: extract features from a sample image by the encoder to obtain feature representations of multiple levels, where the feature representations of different levels are used to reflect different aspects of image feature information; for the feature representation of each level, extract attributes from the feature representation by the attribute extractor corresponding to the level to obtain an initial attribute feature map corresponding to the level; and obtain an attribute feature map based on the multiple initial attribute feature maps respectively corresponding to multiple levels.
[0139] In some embodiments, the forged image detection model further includes content extractors of multiple levels and decoders of multiple levels, and the number of attribute extractors, content extractors, and decoders is the same; the feature extraction unit 902 is specifically configured to: for the feature representation of each level, extract content from the feature representation by the content extractor corresponding to the level to obtain an initial content feature map corresponding to the level; for each level, input the initial attribute feature map and the initial content feature map corresponding to the level into the decoder corresponding to the level for reconstruction to obtain a reconstructed feature or a reconstructed image; the iterative training unit 905, specifically configured to: train the forged image detection model based on the difference between the reconstructed image and the sample image, the differences between the reconstructed features of multiple levels and the feature representations, and the difference between the predicted label and the true / false label.
[0140] In some embodiments, the forged image detection model further includes classifiers corresponding to multiple levels; the prediction classification unit 904 is further configured to: for the initial attribute feature map of each level, perform authenticity prediction on the initial attribute feature map by the classifier corresponding to the level to obtain an initial prediction label; the iterative training unit 905 is specifically configured to: train the forged image detection model based on the difference between the prediction label and the authenticity label, and the difference between the initial prediction label and the authenticity label.
[0141] In some embodiments, the prediction classification unit 904 is specifically configured to: splice the attribute feature map and the weakened attribute feature map; perform authenticity prediction on the spliced attribute feature map to obtain a prediction label.
[0142] In some embodiments, the iterative training unit 905 is further configured to: for multiple sample images with the same authenticity label, calculate the intra-class similarity between the weakened attribute feature maps respectively corresponding to the multiple sample images; for multiple sample images with different authenticity labels, calculate the inter-class similarity between the weakened attribute feature maps respectively corresponding to the multiple sample images; train the forged image detection model based on the intra-class similarity, the inter-class similarity, and the difference between the prediction label and the authenticity label.
[0143] In some embodiments, the influence degree of each pixel on the prediction result is the gradient corresponding to each pixel in the backpropagation process. The weakening processing unit 903 is specifically configured to: determine the gradient corresponding to each pixel in the attribute feature map in the backpropagation process based on the difference between the prediction result and the authenticity label; for each pixel of the attribute feature map, if the gradient of the pixel is higher than a preset threshold, determine the pixel as a key pixel; set the value of the channel corresponding to the key pixel in the attribute feature map to a weakening value.
[0144] Figure 10 It is a schematic structural diagram of a forged image detection device based on a key forgetting mechanism in an embodiment of this specification. The forged image detection model includes at least an attribute extractor and a key forgetting module. This device can be applied to any device, platform or device cluster with computing and processing capabilities. This device includes:
[0145] An image acquisition unit 001, configured to: acquire a target image to be detected and a forged image detection model trained by the method according to any one of claims 1-10. The forged image detection model includes at least an attribute extractor and a key forgetting module;
[0146] An attribute extraction unit 002, configured to: extract attributes from the target image by the attribute extractor to obtain a target attribute feature map, where the target attribute feature map is used to represent information related to forged traces in the target image;
[0147] The detection and classification unit 003 is configured to: use the classifier in the key forgetting module to perform authenticity prediction on the target attribute feature map to obtain a detection result, where the detection result is used to indicate whether the target image is a forged image.
[0148] One or more embodiments of this specification further provide a computer-readable storage medium, which stores a computer program that can be used to execute the above Figure 1 or Figure 7 The training method of the forged image detection model based on the key forgetting mechanism provided in any one of the embodiments, or execute the above Figure 8 The forged image detection method based on the key forgetting mechanism provided in any one of the embodiments.
[0149] One or more embodiments of this specification further provide a computing device, including a memory and a processor. An executable code is stored in the memory. When the processor executes the executable code, it implements the above Figure 1 or Figure 7 The training method of the forged image detection model based on the key forgetting mechanism provided in any one of the embodiments, or execute the above Figure 8 The forged image detection method based on the key forgetting mechanism provided in any one of the embodiments.
[0150] In the 1990s, it was obvious to distinguish whether an improvement to a technology was a hardware improvement (e.g., improvement to circuit structures such as diodes, transistors, switches, etc.) or a software improvement (improvement to method flows). However, with the development of technology, many method flow improvements today can be regarded as direct improvements to hardware circuit structures. Almost all designers obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement to a method flow cannot be implemented with a hardware entity module. For example, a Programmable Logic Device (PLD) (e.g., a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by a user's programming of the device. Designers can program by themselves to "integrate" a digital system onto a single PLD, without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compiler used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a Hardware Description Language (HDL). There is not only one type of HDL, but many types, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply performing some logical programming on the method flow using the above-mentioned several hardware description languages and programming it into an integrated circuit, it is easy to obtain a hardware circuit that implements the logical method flow.
[0151] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an Application Specific Integrated Circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to logically program the method steps to enable the controller to be implemented in the form of logic gates, switches, ASICs, programmable logic controllers, embedded microcontrollers, etc. to achieve the same function. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or the structures within the hardware component.
[0152] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a server system. Of course, this application does not exclude that with the development of future computer technologies, the computers for implementing the functions of the above embodiments can be, for example, personal computers, laptop computers, in-vehicle human-machine interaction devices, cellular phones, camera phones, smart phones, personal digital assistants, media players, navigation devices, email devices, game consoles, tablet computers, wearable devices, or any combination of these devices.
[0153] Although one or more embodiments of this specification provide method operation steps as described in the embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-creative means. The order of steps listed in the embodiments is only one way among many execution orders of steps and does not represent the only execution order. When the actual device or terminal product is executed, it may be executed in the order of the method shown in the embodiments or the drawings or executed in parallel (for example, in an environment of parallel processors or multi-threaded processing, or even in a distributed data processing environment). The terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, product or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or also includes elements inherent to such process, method, product or device. Without further limitation, it does not exclude the existence of additional identical or equivalent elements in the process, method, product or device including the said elements. For example, if terms such as first and second are used to denote names, they do not denote any specific order.
[0154] For convenience of description, when describing the above device, it is divided into various modules according to functions and described separately. Of course, when implementing one or more of this specification, the functions of each module can be implemented in the same or multiple software and / or hardware, or the modules implementing the same function can be realized by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.
[0155] The present invention is described with reference to the flowcharts and / or block diagrams of methods according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0156] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more processes and / or blocks Figure 1 in one or more processes and / or blocks Figure 1 specified in one or more blocks or multiple blocks.
[0157] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to generate a computer-implemented process, thereby providing steps for implementing the functions specified in one or more processes and / or blocks Figure 1 in one or more processes and / or blocks Figure 1 specified in one or more blocks or multiple blocks.
[0158] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0159] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). Memory is an example of computer-readable media.
[0160] Computer-readable media includes both permanent and non-permanent, removable and non-removable media and can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage, graphene storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0161] Those skilled in the art should understand that one or more embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, one or more embodiments of this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, one or more embodiments of this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0162] One or more embodiments of this specification can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. One or more embodiments of this specification can also be practiced in a distributed computing environment, where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0163] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments. In the description of this specification, the description of reference terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this specification. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0164] The above description is only for the embodiments of one or more embodiments of this specification and is not intended to limit one or more embodiments of this specification. For those skilled in the art, one or more embodiments of this specification can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the scope of the claims.
Claims
1. A training method for a forged image detection model based on a key forgetting mechanism, wherein the forged image detection model comprises at least an attribute extractor and a key forgetting module, and the method comprises: Acquire a training sample set, wherein the training sample set includes a plurality of sample images, wherein the sample images have authenticity labels, and the authenticity labels are used to mark whether the sample images are forged images; The attribute extractor extracts attributes based on the sample image to obtain an attribute feature map, wherein the attribute feature map is used to represent information related to the forgery trace in the sample image; The key forgetting module obtains the influence degree of each pixel in the attribute feature map on the prediction result, determines the key pixel in the attribute feature map based on the influence degree of each pixel, and weakens the influence degree of the key pixel; The classifier in the key forgetting module predicts the authenticity of the attribute feature graph after the weakening process to obtain a predicted label; The forged image detection model is trained based on the difference between the predicted label and the true and false label.
2. The method according to claim 1, wherein: The forged image detection model further includes an encoder; the attribute extractor extracts attributes based on the sample image to obtain an attribute feature map, including: The encoder extracts features from the sample image to obtain feature representation; The attribute extractor extracts attributes from the feature representation to obtain the attribute feature graph.
3. The method according to claim 2, wherein: The forged image detection model also includes a content extractor; the method also includes: The content extractor extracts content from the feature representation to obtain the content feature map, where the content feature map is used to represent information unrelated to the forgery trace in the sample image; The step of training the forged image detection model based on the difference between the predicted label and the true or false label includes: The forged image detection model is trained based on the difference between the predicted label and the true and false label and the independence between the content feature map and the attribute feature map.
4. The method according to claim 3, wherein: The forged image detection model also includes a decoder, and the method further includes: Inputting the attribute feature map and the content feature map into the decoder for reconstruction to obtain a reconstructed image; Based on the difference between the predicted label and the true and false label, the forged image detection model is trained, including: The forged image detection model is trained based on the difference between the reconstructed image and the sample image and the difference between the predicted label and the true and false label.
5. The method according to claim 2, wherein: The forged image detection module includes multiple levels of the attribute extractors; The encoder extracts features from the sample image to obtain feature representation, including: The encoder extracts features from the sample image to obtain feature representations at multiple levels, wherein the feature representations at different levels are used to reflect image feature information in different aspects; The attribute extractor extracts attributes from the feature representation to obtain the attribute feature graph, including: For each level of feature representation, an attribute extractor corresponding to the level extracts attributes from the feature representation to obtain an initial attribute feature graph corresponding to the level; The attribute feature map is obtained based on the multiple initial attribute feature maps corresponding to the multiple levels respectively.
6. The method according to claim 5, wherein: The forged image detection model further includes a plurality of levels of content extractors and a plurality of levels of decoders, and the number of the attribute extractors, the content extractors and the decoders is the same; the method further includes: For each level of feature representation, a content extractor corresponding to the level extracts content from the feature representation to obtain an initial content feature graph corresponding to the level; For each level, inputting the initial attribute feature map and the initial content feature map corresponding to the level into a decoder corresponding to the level for reconstruction to obtain a reconstructed feature or a reconstructed image; Based on the difference between the predicted label and the true and false label, the forged image detection model is trained, including: The forged image detection model is trained based on the difference between the reconstructed image and the sample image, the difference between the reconstructed features and the feature representations at multiple levels, and the difference between the predicted label and the true and false label.
7. The method according to claim 5, wherein: The forged image detection model also includes classifiers corresponding to multiple levels; the method also includes: For the initial attribute feature graph of each level, the classifier corresponding to the level predicts the authenticity of the initial attribute feature graph to obtain an initial prediction label; Based on the difference between the predicted label and the true and false label, the forged image detection model is trained, including: The forged image detection model is trained based on the difference between the predicted label and the true or false label, and the difference between the initial predicted label and the true or false label.
8. The method according to claim 1, wherein: The step of performing a true or false prediction on the weakened attribute feature graph to obtain a predicted label includes: splicing the attribute characteristic graph and the attribute characteristic graph after weakening processing; The authenticity of the spliced attribute feature graph is predicted to obtain the predicted label.
9. The method according to claim 1, wherein: The method further comprises: For the plurality of sample images with the same authenticity label, calculating the intra-class similarity between the weakened attribute feature graphs respectively corresponding to the plurality of sample images; For the plurality of sample images with different authenticity labels, calculating the inter-class similarity between the weakened attribute feature graphs respectively corresponding to the plurality of sample images; Based on the difference between the predicted label and the true and false label, the forged image detection model is trained, including: The forged image detection model is trained based on the intra-class similarity, the inter-class similarity, and the difference between the predicted label and the true and false label.
10. The method according to claim 1, wherein: The influence degree of each pixel on the prediction result is the gradient corresponding to each pixel in the back propagation process, and the obtaining of the influence degree of each pixel in the attribute feature map on the prediction result, determining the key pixel in the attribute feature map based on the influence degree of each pixel, and weakening the influence degree of the key pixel includes: Based on the difference between the prediction result and the true and false label, determining the gradient corresponding to each pixel in the attribute feature map during the back propagation process; For each pixel of the attribute feature map, if the gradient of the pixel is higher than a preset threshold, the pixel is determined as a key pixel; The value of the channel corresponding to the key pixel in the attribute feature map is set to a weakened value.
11. A forged image detection method based on a key forgetting mechanism, the method comprising: Acquire a target image to be detected and a forged image detection model trained based on the method according to any one of claims 1 to 10, wherein the forged image detection model comprises at least an attribute extractor and a key forgetting module; The attribute extractor extracts attributes based on the target image to obtain a target attribute feature map, wherein the target attribute feature map is used to represent information related to the forgery trace in the target image; The classifier in the key forgetting module predicts the authenticity of the target attribute feature map to obtain a detection result, and the detection result is used to indicate whether the target image is a forged image.
12. A computing device, comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method according to any one of claims 1 to 10 or the method according to claim 11 is implemented.
Citation Information
Patent Citations
Forged image recognition model training method and forged image recognition method
CN112686331A
Forged video detection method, electronic equipment and storage medium
CN112733733A
Image authenticity identification model training method, application method and device
CN115496963A
Method and device for detecting fake face changing image based on identity recognition probability distribution
CN116188439A
Image detection method and device, equipment and storage medium
CN119380421A