Image Tampering Detection Method and Device Based on Semantic-Irrelevant Feature Learning

Through the plug-in auxiliary module guidance method, the pre-trained benchmark semantic feature extractor and the comparison learning framework constrain the feature similarity, combined with boundary supervision, the generalization problem of the image tamper detection model is solved, and the better tampering area detection effect is achieved.

CN117095228BActive Publication Date: 2025-07-25INST OF COMPUTING TECH CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311117995.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-31
Publication Date
2025-07-25
Estimated Expiration
2043-08-31

AI Technical Summary

Technical Problem

The existing image tamper detection model has insufficient generalization capabilities, and it is difficult to effectively identify tampering areas at any location in the image, and it is difficult to unify the auxiliary task and the main task.

Method used

Using the plug-in auxiliary module guidance method, the similarity between the tamper trace features and the reference semantic features are constrained by pre-training the benchmark semantic feature extractor and the comparison learning framework, and combined with the tampering area boundary supervision, the feature conversion structure is designed to achieve the unity of auxiliary tasks and main tasks.

Benefits of technology

It improves the semantic irrelevant feature learning ability of the image tamper detection model, improves the generalization ability of the model in unknown scenarios, and ensures the accuracy and robustness of tampering area detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117095228B_ABST
    Figure CN117095228B_ABST
Patent Text Reader

Abstract

The present invention proposes an image forgery detection method and device based on semantic-agnostic feature learning, including extracting semantic features of the image to be detected through a pre-trained benchmark semantic feature encoder, using a contrastive learning framework to constrain the similarity between local region forgery trace features and benchmark semantic features, thereby directly restricting the semantic correlation in the forgery trace features; using forgery region boundary supervision to guide the model to mine the inconsistency between the real region and the forgery region features near the forgery region boundary, and designing a feature transformation structure to achieve the unity of the auxiliary task and the main task, so as to ensure that the forgery region boundary prediction task as the auxiliary task can provide gain for the forgery region prediction task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image forgery detection in image detection and classification, and particularly relates to an image forgery detection method and device based on semantic-irrelevant feature learning. Background Art

[0002] The wide application of image editing and forgery technologies has provided more possibilities for the content of our visual world, but these technologies have also brought major challenges to ensuring the authenticity and security of graphic content in various media. Although most forgery and editing occur in regions with strong correlation with semantic attributes, for example, people in images are often key elements and are also regions that attract more attention from image viewers, so forgery often occurs in the people region; however, in real scenarios, forgery operations may occur at any position in the image, which requires an image forgery detection model that can identify the forgery region in an image to pay more attention to forgery trace features rather than the limited semantic correlation existing in the limited training data, so as to ensure that the model has good generalization ability.

[0003] In order to ensure that the model can learn semantic-irrelevant forgery trace features to obtain better generalization ability, existing methods mainly use specific feature extraction modules to restrict the input semantic information, or avoid the guidance of semantic mask labels by converting the form of task objectives and modifying the structure of the segmentation network. The specifically designed feature extraction module can convert the RGB space containing rich semantic content into a semantic-irrelevant noise space or high-frequency information space. However, noise features and high-frequency information are easily erased or changed by post-processing operations (such as compression and blurring), and many other available forgery trace features are also directly filtered at this stage. Redesigning the network structure and converting the task objective can avoid the supervision of semantic mask labels, but when there is a conflict between the auxiliary task and the original task, it is difficult for the model to directly unify these two tasks. Summary of the Invention

[0004] The purpose of the present invention is to improve the ability of learning semantic-irrelevant features in the image forgery detection task and thus ensure the generalization of the detection method, and a method for image forgery detection guided by a plug-in auxiliary module is proposed.

[0005] Specifically, the present invention provides an image forgery detection method based on semantic-irrelevant feature learning, which includes:

[0006] A step of extracting benchmark semantic features, obtaining an image with a labeled forgery region label as training data, and using a benchmark semantic feature extraction model to extract the benchmark semantic features of the training data;

[0007] A step of extracting a boundary mask label of the forgery region, and extracting the boundary mask label of the forgery region in the training data according to the forgery region label;

[0008] Model detection step: construct an image forgery detection model including an encoder, a decoder, a classification head, and a segmentation head. The encoder extracts forgery trace features of the training data according to the forgery region label. The classification head obtains the training classification result of whether the training data has been forged based on the forgery trace features. The decoder and the segmentation head obtain the training region result of the forged part in the training data based on the forgery trace features. Construct a classification loss according to the forgery region label and the training classification result, and construct a segmentation loss according to the forgery region label and the training region result.

[0009] Auxiliary task guidance step: use a contrastive learning framework to constrain the similarity between the forgery trace features and the reference semantic features to obtain a contrastive learning loss. Use the boundary mask label to guide the model to mine the inconsistency between the real region and the forgery region features near the forgery region boundary to obtain a forgery region boundary prediction loss.

[0010] Model training execution step: train the image forgery detection model according to the segmentation loss, the classification loss, the contrastive learning loss, and the forgery region boundary prediction loss, and use the trained image forgery detection model to perform an image forgery detection task to obtain whether the target image is a forged image. If it is a forged image, the forged region in the target image is obtained together.

[0011] The image forgery detection method based on semantic-agnostic feature learning, wherein the reference semantic feature extraction step includes:

[0012] Obtain the encoders in multiple semantically segmented models with different structures, input the training data into each encoder respectively to obtain the semantic features extracted by each encoder, and use cosine similarity to constrain the semantic features extracted by all encoders to narrow the distance of the semantic features in the feature space. After the constraint is completed, select one of the encoders as the reference semantic feature extractor.

[0013] The image forgery detection method based on semantic-agnostic feature learning, wherein the forgery region boundary mask label extraction step includes:

[0014] Through the following formula based on the sliding window mechanism, obtain the boundary mask label G of whether each pixel position belongs to the forgery region boundary i,j :

[0015]

[0016] where W ij is a window centered on the pixel position (i, j), and (m, n) are the pixel positions in the window.

[0017] The described image forgery detection method based on semantic-agnostic feature learning, where the auxiliary task guiding step includes:

[0018] The contrastive learning framework performs contrastive learning based on local region feature blocks on the forgery trace feature F tr for each feature block t k located in the forgery region in p , whose positive sample t n is a randomly selected feature block in other locations also located in the forgery region, and whose negative sample t k is the feature block located in the non-forgery region and semantically closest to t

[0019] t n = t ij ,

[0020] s.t. (i, j) = argmax (i,j) {s p · s ij | s ij ∈ F se \ T s}

[0021] where s p and s ij are the local features in the reference semantic feature corresponding to t p and t ij respectively, T s is the set of local reference semantic feature vectors in the forgery region, and F se is the reference semantic feature;

[0022] Use the contrastive learning loss L SSM to learn semantic-agnostic features:

[0023]

[0024] P(t k ) is the set of positive samples of t k , N(t k ) is the set of negative samples of t k , t p is any positive sample in the set of positive samples, t a is any sample in the set of positive samples or the set of negative samples, and τ is the hyperparameter used for contrastive learning;

[0025] Convert the guidance of the forgery region boundary prediction task on features to the guidance of the forgery region prediction task on features through constrained convolution. The constrained convolution is:

[0026]

[0027] Input the forgery trace feature F tr into the constrained convolutional layer BayarConv. The parameter ω of the k-th convolution of the constrained convolution k satisfies the above constraints, where the spatial index (0, 0) represents the center position of the convolution kernel, and (x, y) represents any other position outside the center position of the convolution kernel; the constrained convolutional layer outputs the boundary trace feature F bd , complete the transformation of the feature, and then use the decoder to predict the feature to obtain the forgery area boundary prediction S. The total loss L including the cross-entropy loss and the mean error loss is used to jointly train the image forgery detection model:

[0028] L BGM =∑ (i,j) (L bce (S i,j , bin(G i,j ))·w i,j +L l1 (S i,j , G i,j ))

[0029] L = L seg +λ1L cls +λ2L SSM +λ3L BGM

[0030] where L seg and L cls are the segmentation and classification losses of the baseline semantic segmentation model, both using the cross-entropy loss. λ1, λ2, and λ3 are the weight hyperparameters for balancing the three tasks respectively.

[0031] The present invention also proposes an image forgery detection device based on semantic-agnostic feature learning, which includes:

[0032] A benchmark semantic feature extraction module for obtaining an image with a labeled forgery area label as training data and extracting the benchmark semantic features of the training data using a benchmark semantic feature extraction model;

[0033] A forgery area boundary mask label extraction module for extracting the boundary mask label of the forgery area in the training data according to the forgery area label;

[0034] A model detection module, used to construct an image tampering detection model including an encoder, a decoder, a classification head, and a segmentation head. The encoder extracts the tampering trace features of the training data according to the tampering region label. The classification head obtains the training classification result of whether the training data has been tampered with according to the tampering trace features. The decoder and the segmentation head obtain the training region result of the tampered part in the training data according to the tampering trace features. A classification loss is constructed according to the tampering region label and the training classification result, and a segmentation loss is constructed according to the tampering region label and the training region result.

[0035] An auxiliary task guiding module, used to use a contrast learning framework to constrain the similarity between the tampering trace features and the benchmark semantic features to obtain a contrast learning loss, and use the boundary mask label to guide the model to mine the inconsistency between the real region and the tampering region features near the tampering region boundary to obtain a tampering region boundary prediction loss.

[0036] A model training execution module, used to train the image tampering detection model according to the segmentation loss, the classification loss, the contrast learning loss, and the tampering region boundary prediction loss, and use the trained image tampering detection model to perform an image tampering detection task to obtain whether the target image is a tampered image. If it is a tampered image, the tampering region in the target image is obtained together.

[0037] The image tampering detection method based on semantic irrelevant feature learning, wherein the benchmark semantic feature extraction module includes:

[0038] Obtain the encoders in multiple semantic segmentation models with different structures, input the training data into each encoder respectively, obtain the semantic features extracted by each encoder, and use cosine similarity to constrain the semantic features extracted by all encoders to narrow the distance of the semantic features in the feature space. After the constraint is completed, select one of the encoders as the benchmark semantic feature extractor.

[0039] The image tampering detection device based on semantic irrelevant feature learning, wherein the tampering region boundary mask label extraction module includes:

[0040] Through the following formula based on the sliding window mechanism, obtain the boundary mask label G of whether each pixel position belongs to the tampering region boundary i,j :

[0041]

[0042] where, W ij is the window centered on the pixel position (i, j), and (m, n) is the pixel position in the window.

[0043] The image tampering detection device based on semantic irrelevant feature learning, wherein the auxiliary task guiding module includes:

[0044] The contrastive learning framework performs contrastive learning based on local region feature blocks for the tampering trace feature F tr For each feature block t k in the tampered area in p , its positive sample t n is a feature block randomly selected from other locations also in the tampered area, and its negative sample t k is the feature block that is located in the untampered area and has the closest semantics to t

[0045] t n = t ij ,

[0046] s.t. (i, j) = argmax (i,j) {s p · s ij | s ij ∈ F se \T s}

[0047] where s p and s ij are the local features in the reference semantic feature corresponding to t p and t ij respectively, T s is the set of local reference semantic feature vectors in the tampered area, and F se is the reference semantic feature;

[0048] Use the contrastive learning loss L SSM to learn features that are semantically irrelevant:

[0049]

[0050] P(t k ) is the set of positive samples of t k , N(t k ) is the set of negative samples of t k , t p is any positive sample in the set of positive samples, t a is any sample in the set of positive samples or the set of negative samples, and τ is the hyperparameter used for contrastive learning;

[0051] Convert the guidance of the tampering area boundary prediction task on features to the guidance of the tampering area prediction task on features through constrained convolution. The constrained convolution is:

[0052]

[0053] Take the tampering trace feature F trInput constraint convolutional layer BayarConv, where the parameters ω of the k-th layer of constrained convolution k Satisfy the above constraints, where the spatial index (0, 0) represents the center position of the convolutional kernel, and (x, y) represents any other position outside the center position of the convolutional kernel; the constrained convolutional layer outputs the boundary trace feature F bd , complete the transformation of the feature, and then use this decoder to predict this feature to obtain the tampered area boundary prediction S, and jointly train the image tampering detection model through the total loss L including this cross-entropy loss and this mean error loss:

[0054] L BGM = ∑ (i,j) (L bce (S i,j , bin(G i,j )) · w i,j + L l1 (S i,j , G i,j ))

[0055] L = L seg + λ1L cls + λ2L SSM + λ3L BGM

[0056] where L seg and L cls are the segmentation and classification losses of the baseline semantic segmentation model, both using cross-entropy loss, and λ1, λ2, and λ3 are the weight hyperparameters for balancing the three tasks respectively.

[0057] The present invention also proposes a server, as Figure 2 shown, including the image tampering detection device based on semantic-agnostic feature learning.

[0058] The present invention also proposes a storage medium for storing a computer program for executing the image tampering detection method based on semantic-agnostic feature learning.

[0059] As can be seen from the above solutions, the advantages of the present invention are as follows:

[0060] The semantic suppression module based on contrast learning of the present invention can directly constrain the feature extractor to ensure that it learns semantic-agnostic features; the boundary guidance module based on feature transformation can, on the basis of ensuring the unity of the auxiliary task and the main task, guide the model to mine the inconsistency of the features in the area near the boundary of the tampered area; the image tampering detection method guided by the plug-in auxiliary module realizes the adaptation of the auxiliary task module to the existing semantic segmentation framework, so that the plug-and-play auxiliary task module can effectively guide various semantic segmentation models to learn semantic-agnostic features to obtain better generalization ability. Description of the Drawings

[0061] Figure 1 This is the flowchart of the image tampering detection method guided by the plug-in auxiliary module of the present invention;

[0062] Figure 2 This is the schematic diagram of the server structure of the present invention. Detailed implementation manners

[0063] To achieve the above technical effects, the present invention includes the following key technical points:

[0064] Key point 1, the semantic suppression module based on contrast learning. The present invention first pre-trains a benchmark semantic feature extraction module, and then uses the semantic features extracted by it as a reference, and uses a contrast learning framework to constrain the similarity between the local area tampering trace features and the benchmark semantic features, thereby directly restricting the semantic correlation in the tampering trace features.

[0065] Key point 2, the boundary guidance module based on feature transformation. According to the existing tampering area mask label, this method uses a sliding window mechanism to calculate the label of the boundary of the tampering area, and then uses this as a supervision to guide the model to mine the inconsistency between the features of the real area and the tampering area near the boundary of the tampering area, and realizes the unification of the auxiliary task and the main task through a specially designed feature transformation structure.

[0066] Key point 3, the image tampering detection method guided by the plug-in auxiliary module. Without changing the existing semantic segmentation framework, this method directly constrains the feature extractor using a plug-and-play auxiliary task module, so as to better guide the model to learn semantically irrelevant features, and then obtain the generalization ability of detecting image tampering areas in unknown scenarios. The auxiliary task module includes Figure 1 the boundary guidance module and the semantic suppression module in it; the sliding window mechanism calculation module serves the boundary guidance module, and the benchmark semantic feature extraction module serves the semantic suppression module.

[0067] To make the above features and effects of the present invention more clearly and understandably described, specific embodiments are given below, and detailed descriptions are made in conjunction with the accompanying drawings of the specification as follows. This specification discloses one or more embodiments including the features of the present invention. The disclosed embodiments are only for illustrative purposes. The protection scope of the present invention is not limited to the disclosed embodiments, and the present invention is defined by the appended claims.

[0068] References to "one embodiment", "an embodiment", "example embodiment", etc. in the specification mean that the described embodiment may include specific features, structures, or characteristics, but not every embodiment must include these specific features, structures, or characteristics. Moreover, such expressions do not refer to the same embodiment. Further, when describing specific features, structures, or characteristics in combination with an embodiment, it has been shown that combining such features, structures, or characteristics into other embodiments is within the knowledge of those skilled in the art, whether or not explicitly described.

[0069] To solve the generalization problem of image tampering detection, the present invention proposes a semantic-irrelevant feature learning network guided by a plug-in auxiliary module. This method uses an unmodified semantic segmentation model as the backbone network and restricts the feature encoder of the model to learn only semantic-irrelevant features through modular auxiliary tasks. In the semantic suppression module, the semantic features of the image to be detected are extracted by a pre-trained benchmark semantic feature encoder, and a contrast learning framework is used to constrain the similarity between the local region tampering trace features and the benchmark semantic features, thereby directly restricting the semantic relevance in the tampering trace features; in the boundary guidance module, the boundary supervision of the tampering region is used to guide the model to mine the inconsistency between the features of the real region and the tampering region near the boundary of the tampering region, and a feature transformation structure is designed to unify the auxiliary task and the main task to ensure that the tampering region boundary prediction task as an auxiliary task can provide a gain for the tampering region prediction task.

[0070] The overall process of the image tampering detection method guided by the plug-in auxiliary module of the present invention is as Figure 1 shown: First, a benchmark semantic feature extraction module that is not affected by the model structure is pre-trained to extract the benchmark semantic features of the image to be detected. Second, according to the existing tampering region mask labels, the labels of the tampering region boundaries are calculated using a sliding window mechanism. Finally, a plug-in auxiliary task module is used to guide the semantic segmentation model to focus on semantic-irrelevant tampering traces. The following is a detailed introduction to each module of this method.

[0071] Benchmark semantic feature extraction step:

[0072] To obtain reliable benchmark semantic features F se and minimize the influence of the pre-training dataset and the encoder structure, we use the encoder in the pre-trained semantic segmentation model and fine-tune it.

[0073] In the present invention, the encoders in two semantic segmentation models with different structures pre-trained on different semantic segmentation task datasets are selected as the initial encoders. At this time, the features they extract contain rich semantic features. In order to reduce the influence of the pre-trained dataset and the encoder structure on the extracted benchmark semantic features, the cosine similarity is used to constrain the semantic features extracted by the two different encoders, hoping that the features extracted by the two encoders are as close as possible in the feature space. After the fine-tuning of this stage, one of the encoders is selected as the benchmark semantic feature extractor, because after fine-tuning, the effects of the two encoders are the same and either one can be chosen randomly.

[0074] After obtaining the benchmark semantic feature extractor, we input the image X to be detected into the benchmark semantic feature extractor to obtain the extracted benchmark semantic feature F se , and this feature will provide a reference for the constraint of the subsequent semantic suppression module.

[0075] Steps for extracting the boundary mask label of the tampered area:

[0076] Since the original annotation provided in the Chinese image tampering dataset is about the tampered area and does not include the annotation of the boundary of the tampered area, in this process, we calculate the tampered area label g provided in the dataset to obtain the boundary mask label G of the tampered area.

[0077] We do not directly obtain the boundary label reflecting the absolute boundary position using the common boundary detection algorithm because the regional annotation of the original data has the problem of being not refined, and in order to better adapt to the boundary guidance module, we use the sliding window mechanism to calculate the distance of each pixel position from the absolute boundary:

[0078]

[0079] Among them, W ij is the window centered on the pixel position (i, j), the window size can be a preset value, and (m, n) is the pixel position in the window. The boundary annotation of the tampered area calculated using this window mechanism provides supervision information for the subsequent boundary guidance module.

[0080] Steps for auxiliary task guidance:

[0081] Finally, the auxiliary tasks designed in the auxiliary module and the main task (tampered area prediction) will form a multi-task training framework together to jointly guide the encoder to learn semantically irrelevant features.

[0082] The semantic suppression module uses the contrast learning framework to perform contrast learning on local region feature blocks. For each feature block t tr in the tampering trace feature F k that is located in the tampered area, its positive sample t pFor a feature block located in the tampered area at another randomly selected position, its negative sample is t n is located in the original area that has not been tampered with, but is the one that is k semantically closest to t:

[0083] t n = t ij ,

[0084] s.t. (i, j) = argmax (i,j) {s p · s ij | s ij ∈ F se \ T s}

[0085] where s p and s ij are the local features in the reference semantic features corresponding to t p and t ij respectively, used to calculate the distance between semantics, and T s is the set of local reference semantic feature vectors in the tampered area. After the selection of positive and negative samples, the contrastive learning loss L SSM is used to learn features that are semantically irrelevant:

[0086]

[0087] where t k is the selected anchor sample, P(t k ) is the set of positive samples corresponding to the anchor sample, N(t k ) is the set of negative samples corresponding to the anchor sample, t p is any positive sample in the set of positive samples, t a is any sample in the set of positive samples or the set of negative samples, and τ is the hyperparameter used in contrastive learning.

[0088] The boundary guidance module uses a constrained but trainable constrained convolution to convert the guidance of the tampered area boundary prediction task on features to the guidance of the tampered area prediction task on features, achieving the unification of the two tasks. The constraints designed for the constrained convolution are as follows:

[0089]

[0090] Input the tampering trace feature F tr into the constrained convolution layer BayarConv. The parameters ω k of the k-th convolution of this constrained convolution satisfy the above constraints, where the spatial index (0, 0) represents the center position of the convolution kernel, and (x, y) represents any other position outside the center position of the convolution kernel. The constrained convolution layer outputs the boundary trace feature Fbd Complete the transformation of the features, and then use the decoder to predict the features to obtain the prediction S of the tampered region boundary. The model is trained jointly by cross-entropy loss and mean error loss:

[0091] L BGM =∑ (i,j) (L bce (S i,j , bin(G i,j ))·w i,j +L l1 (S i,j , G i,j )

[0092] During the training process, the multi-task loss is used to update the parameters of the entire model including the encoder, decoder, segmentation head, and classification head:

[0093] L = L seg +λ1L cls +λ2L SSM +λ3L BGM

[0094] where L seg and L cls are the segmentation and classification losses of the baseline semantic segmentation model, both using cross-entropy loss. λ1, λ2, and λ3 are weight hyperparameters for balancing the three tasks, and are all set to 0.1 by default.

[0095] The following is a system embodiment corresponding to the above method embodiment. This embodiment can be implemented in cooperation with the above embodiment. The relevant technical details mentioned in the above embodiment are still valid in this embodiment. To avoid repetition, they will not be elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied to the above embodiment.

[0096] The present invention also proposes an image tampering detection device based on semantic-agnostic feature learning, which includes:

[0097] A benchmark semantic feature extraction module, configured to obtain an image with a labeled tampered region label as training data, and extract the benchmark semantic features of the training data using a benchmark semantic feature extraction model;

[0098] A tampered region boundary mask label extraction module, configured to extract the boundary mask label of the tampered region in the training data according to the tampered region label;

[0099] A model detection module is used to construct an image tampering detection model including an encoder, a decoder, a classification head and a segmentation head, wherein the encoder extracts the tampering trace feature of the training data according to the tampering area label, the classification head obtains a training classification result of whether the training data has been tampered with according to the tampering trace feature, the decoder and the segmentation head obtain a tampered training area result in the training data according to the tampering trace feature, a classification loss is constructed according to the tampering area label and the training classification result, and a segmentation loss is constructed according to the tampering area label and the training area result;

[0100] The auxiliary task guiding module is used to use the contrastive learning framework to constrain the similarity between the tampering trace feature and the benchmark semantic feature to obtain the contrastive learning loss, and use the boundary mask label to guide the model to mine the inconsistency between the real area and the tampered area features near the boundary of the tampered area to obtain the tampered area boundary prediction loss;

[0101] The model training execution module is used to train the image tampering detection model according to the segmentation loss, the classification loss, the contrastive learning loss and the tampering area boundary prediction loss, and use the trained image tampering detection model to perform image tampering detection tasks to obtain whether the target image is a tampered image, and if it is a tampered image, the tampered area in the target image is also obtained.

[0102] The image tampering detection method based on semantic-independent feature learning, wherein the reference semantic feature extraction module comprises:

[0103] Obtain encoders from multiple semantic segmentation models with different structures, input training data into each encoder respectively, obtain the semantic features extracted by each encoder, use cosine similarity to constrain the semantic features extracted by all encoders to shorten the distance between semantic features in feature space; after the constraints are completed, select one encoder as the benchmark semantic feature extractor.

[0104] The image tampering detection device based on semantically irrelevant feature learning, wherein the tampering area boundary mask label extraction module comprises:

[0105] The following formula based on the sliding window mechanism is used to obtain the boundary mask label G of each pixel position, indicating whether it belongs to the boundary of the tampered area. i,j :

[0106]

[0107] Among them, W ij is a window centered at pixel position (i, j), and (m, n) is the pixel position in the window.

[0108] The image tampering detection device based on semantically irrelevant feature learning, wherein the auxiliary task guiding module comprises:

[0109] The contrastive learning framework performs contrastive learning based on local region feature blocks for the tampering trace feature F tr For each feature block t k in the tampering area in p , its positive sample t n is a feature block randomly selected from other locations that is also in the tampering area, and its negative sample t k is the feature block that is located in the non-tampered area and has the closest semantics to t

[0110] t n = t ij ,

[0111] s.t. (i, j) = argmax (i,j) {s p · s ij | s ij ∈ F se \T s}

[0112] where s p and s ij are the local features in the reference semantic features corresponding to t p and t ij respectively, T s is the set of local reference semantic feature vectors in the tampering area, and F se is the reference semantic feature;

[0113] Use the contrastive learning loss L SSM to learn features that are semantically irrelevant:

[0114]

[0115] P(t k ) is the set of positive samples of t k , N(t k ) is the set of negative samples of t k , t p is any positive sample in the set of positive samples, t a is any sample in the set of positive samples or the set of negative samples, and τ is the hyperparameter used in contrastive learning;

[0116] Convert the guidance of the tampering area boundary prediction task on features to the guidance of the tampering area prediction task on features through constrained convolution. The constrained convolution is:

[0117]

[0118] Convert the tampering trace feature F trInput the constrained convolutional layer BayarConv, and the parameter ω of the k-th constrained convolution k Satisfy the above constraints, where the spatial index (0, 0) represents the center position of the convolution kernel, and (x, y) represents any other position outside the center position of the convolution kernel; the constrained convolutional layer outputs the boundary trace feature F bd , complete the transformation of the feature, and then use the decoder to predict the feature to obtain the tampered area boundary prediction S. The image tampering detection model is jointly trained by the total loss L including the cross-entropy loss and the mean error loss:

[0119] LB GM =∑ (i,j) (L bce (S i,j , bin(G i,j ))·w i,j +L l1 (S i,j , G i,j ))

[0120] L = L seg +λ1L cls +λ2L SSM +λ3L BGM

[0121] where L seg and L cls are the segmentation and classification losses of the baseline semantic segmentation model, both using the cross-entropy loss. λ1, λ2, and λ3 are the weight hyperparameters for balancing the three tasks respectively.

[0122] The present invention also proposes a server, including the image tampering detection device based on semantic-agnostic feature learning.

[0123] The present invention also proposes a storage medium for storing a computer program for executing the image tampering detection method based on semantic-agnostic feature learning.

[0124] Although the embodiments of the present invention have been disclosed as above, they are not limited to only the applications listed in the specification and the embodiments. It can be fully applied to various fields suitable for the present invention. For those familiar with the field, additional modifications can be easily made. Therefore, without departing from the general concept defined by the claims and the equivalent scope, the present invention is not limited to the specific details and the illustrated and described examples here.

Claims

1. An image forgery detection method based on semantic-irrelevant feature learning, characterized in that Including: A benchmark semantic feature extraction step, obtaining an image with a labeled tampered area label as training data, and using a benchmark semantic feature extraction model to extract the benchmark semantic features of the training data; A tampered area boundary mask label extraction step, extracting the boundary mask label of the tampered area in the training data according to the tampered area label; A model detection step, constructing an image tampering detection model including an encoder, a decoder, a classification head, and a segmentation head. The encoder extracts the tampering trace features of the training data according to the tampered area label. The classification head obtains the training classification result of whether the training data has been tampered according to the tampering trace features. The decoder and the segmentation head obtain the training area result of the tampered area in the training data according to the tampering trace features. A classification loss is constructed according to the tampered area label and the training classification result, and a segmentation loss is constructed according to the tampered area label and the training area result; An auxiliary task guidance step, using a contrastive learning framework to constrain the similarity between the tampering trace features and the benchmark semantic features to obtain a contrastive learning loss, and using the boundary mask label to guide the model to mine the inconsistency between the real area near the tampered area boundary and the tampered area features to obtain a tampered area boundary prediction loss; A model training execution step, training the image tampering detection model according to the segmentation loss, the classification loss, the contrastive learning loss, and the tampered area boundary prediction loss, and using the trained image tampering detection model to execute the image tampering detection task to obtain whether the target image belongs to a tampered image. If it belongs to a tampered image, the tampered area in the target image is obtained together.

2. The image forgery detection method based on semantic-irrelevant feature learning according to claim 1, characterized in that The benchmark semantic feature extraction step includes: Obtaining the encoders in multiple semantic segmentation models with different structures, inputting the training data into each encoder respectively to obtain the semantic features extracted by each encoder, and using cosine similarity to constrain the semantic features extracted by all encoders to narrow the distance of the semantic features in the feature space; after the constraint is completed, selecting one of the encoders as the benchmark semantic feature extractor.

3. The method for detecting image forgery based on semantic-irrelevant feature learning according to claim 1, wherein The tampered area boundary mask label extraction step includes: The boundary mask label G indicating whether each pixel position belongs to the boundary of the tampered region is obtained through the following formula based on the sliding window mechanism i,j :[[]]END]] Among them, W ij is a window centered on the pixel position (i, j), (m, n) is the pixel position in the window, and g is the label of the tampered area.

4. The method for detecting image tampering based on semantic-irrelevant feature learning according to claim 1, wherein The auxiliary task guidance step includes: The contrastive learning framework performs contrastive learning based on local region feature blocks on the forgery trace feature F tr For each feature block t k located in the forgery area in p , its positive sample t n is a feature block randomly selected from other locations also in the forgery area, and its negative sample t k is the feature block located in the non-forged area and semantically closest to t: t n = t ij , such that (i,j) = argmax (i,j) {s p ·s ij |s ij ∈F se \T s} where s p and s ij are respectively the local features in the reference semantic feature corresponding to t p and t ij , T s is the set of local reference semantic feature vectors of the tampered area, and F se is the reference semantic feature; Use the contrastive learning loss L SSM Learn semantically irrelevant features: P(t k ) is the set of positive samples for t k , and N(t k ) is the set of negative samples for t k . Let t p be any positive sample in the set of positive samples, and t a be any sample in the set of positive samples or the set of negative samples. τ is a hyperparameter used for contrastive learning; Converting the guidance of the tampered area boundary prediction task on features to the guidance of the tampered area prediction task on features through a constrained convolution, and the constrained convolution is: Input the tampering trace feature F tr into the constrained convolutional layer BayarConv. The parameter ω of the k-th convolution of the constrained convolution k satisfies the above constraints, where the spatial index (0, 0) represents the center position of the convolution kernel, and (x, y) represents any other position outside the center position of the convolution kernel; the constrained convolutional layer outputs the boundary trace feature F bd , complete the transformation of the feature, and then use the decoder to predict the feature to obtain the tampering area boundary prediction S. The image tampering detection model is jointly trained by the total loss L including the cross-entropy loss and the mean error loss: L BGM = ∑ (i,j) (L bce (S i,j , bin(G i,j )) · w i,j + L l1 (S i,j , G i,j )) L = L seg + λ1L cls + λ2L SSM + λ3L BGM where L seg and L cls are the segmentation and classification losses of the baseline semantic segmentation model, both using cross-entropy loss, and λ1, λ2, and λ3 are the weight hyperparameters for balancing the three tasks, respectively.

5. An image forgery detection device based on semantic-irrelevant feature learning, characterized in that, Including: A benchmark semantic feature extraction module, used to obtain an image with a labeled tampered area label as training data, and use a benchmark semantic feature extraction model to extract the benchmark semantic features of the training data; A tampered area boundary mask label extraction module, extracting the boundary mask label of the tampered area in the training data according to the tampered area label; The model detection module is used to construct an image tampering detection model including an encoder, a decoder, a classification head, and a segmentation head. The encoder extracts the tampering trace features of the training data according to the tampering region label. The classification head obtains the training classification result of whether the training data has been tampered with based on the tampering trace features. The decoder and the segmentation head obtain the training region result of the tampered part in the training data according to the tampering trace features. A classification loss is constructed based on the tampering region label and the training classification result, and a segmentation loss is constructed based on the tampering region label and the training region result. The auxiliary task guidance module is used to use the contrast learning framework to constrain the similarity between the tampering trace features and the reference semantic features to obtain a contrast learning loss, and use the boundary mask label to guide the model to mine the inconsistency between the real region and the tampering region features near the tampering region boundary to obtain a tampering region boundary prediction loss. The model training execution module is used to train the image tampering detection model according to the segmentation loss, the classification loss, the contrast learning loss, and the tampering region boundary prediction loss, and use the trained image tampering detection model to perform the image tampering detection task to obtain whether the target image is a tampered image. If it is a tampered image, the tampered region in the target image is obtained together.

6. The image forgery detection method based on semantic-irrelevant feature learning according to claim 5, wherein The reference semantic feature extraction module includes: Obtain the encoders in multiple semantic segmentation models with different structures, input the training data into each encoder respectively, obtain the semantic features extracted by each encoder, and use cosine similarity to constrain the semantic features extracted by all encoders to narrow the distance of the semantic features in the feature space. After the constraint is completed, select one of the encoders as the reference semantic feature extractor.

7. The image forgery detection device based on semantic-irrelevant feature learning according to claim 5, wherein The tampering region boundary mask label extraction module includes: Obtain the boundary mask label G indicating whether each pixel position belongs to the boundary of the tampered area through the following formula based on the sliding window mechanism i,j : Among them, W ij is a window centered on the pixel position (i, j), and (m, n) is the pixel position in the window.

8. The image forgery detection device based on semantic-irrelevant feature learning according to claim 5, wherein The auxiliary task guidance module includes: This contrastive learning framework conducts contrastive learning based on local region feature blocks, for the forgery trace feature F tr For each feature block t k located in the forgery area in p , its positive sample t n is a feature block randomly selected from other locations that is also located in the forgery area, and its negative sample t k is the feature block located in the non-forged area and having the closest semantics to t t n = t ij , such that (i,j) = argmax (i,j) {s p ·s ij |s ij ∈F se \T s} where s p and s ij are respectively the local features in the reference semantic feature corresponding to t p and t ij , T s is the set of local reference semantic feature vectors of the tampered area, and F se is the reference semantic feature; Use the contrastive learning loss L SSN Perform the learning of semantically irrelevant features: P(t k ) is the set of positive samples for t k , N(t k ) is the set of negative samples for t k , t p is any positive sample in the set of positive samples, t a is any sample in the set of positive samples or the set of negative samples, and τ is a hyperparameter used for contrastive learning; Convert the guidance of the tampering region boundary prediction task on features to the guidance of the tampering region prediction task on features through constraint convolution. The constraint convolution is: Input the tampering trace feature F tr into the constrained convolutional layer BayarConv. The parameter ω of the k-th convolution of the constrained convolution k satisfies the above constraints, where the spatial index (0, 0) represents the center position of the convolution kernel, and (x, y) represents any other position outside the center position of the convolution kernel; the constrained convolutional layer outputs the boundary trace feature F bd , complete the transformation of the feature, then use the decoder to predict the feature to obtain the tampering area boundary prediction S, and jointly train the image tampering detection model through the total loss L including the cross-entropy loss and the mean error loss: L BGM = ∑ (i,j) (L bce (S i,j , bin(G i,j )) · w i,j + L l1 (S i,j , G i,j )) L = L seg + λ1L cls + λ2L SSM + λ3L BGM where L seg and L cls are the segmentation and classification losses of the baseline semantic segmentation model, both using cross-entropy loss. λ1, λ2, and λ3 are the weight hyperparameters for balancing the three tasks respectively, and W ij is a window centered at the pixel position (i, j).

9. A server, characterized in that, An image tampering detection device based on semantic irrelevant feature learning according to any one of claims 5-8.

10. A storage medium for storing a computer program for executing an image tampering detection method based on semantic irrelevant feature learning according to any one of claims 1-4.

Citation Information

Patent Citations

  • Robust tampered image positioning method based on multiple tasks and comparative learning

    CN116049772A

  • Image tampering detection method based on edge guidance and multi-level search

    CN116342601A