High generalization image tampering detection and positioning method based on multi-granularity representation decoupling and collaboration
Patent Information
- Application Number
- CN202611249318.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-18
- Publication Date
- 2026-09-22
AI Technical Summary
[0004]但是,现有深度学习方法仍存在不足:一方面,模型容易学习与训练数据分布或场景语义相关的偏置信息,在跨数据集、跨篡改类型条件下易出现误检和漏检;另一方面,在复杂边界、小尺寸篡改区域以及压缩、模糊、噪声等后处理干扰场景下,局部伪迹和边界细节容易在特征提取与融合过程中被削弱,导致定位结果边界模糊、区域不连续
[0079]1、本发明通过构建细粒度异常表征和粗粒度结构表征,在统一框架下同时建模篡改区域的局部边界伪迹和整体区域一致性信息,能够兼顾边界细节刻画与区域结构建模;
Smart Images

Figure CN122799243A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of image processing, multimedia forensics and image content security technology, and specifically relates to a highly generalized image tampering detection and localization method based on multi-granularity representation decoupling and collaboration. Background Technology
[0002] With the rapid development of mobile communication technology, mobile internet, and social media platforms, digital images have become one of the most widely disseminated forms of multimedia information in cyberspace, widely used in news media, e-commerce, public safety, and judicial evidence collection. However, the widespread use of image editing tools has significantly lowered the technical threshold for image forgery and tampering, making manipulation operations such as splicing, copying and pasting, and deletion repair easier to implement. Once tampered images are disseminated on communication networks or social media platforms, they will seriously threaten the authenticity of communication content, the credibility of digital media, and the reliability of judicial evidence collection.
[0003] Existing image tampering detection and localization methods mainly include traditional manual feature-based methods and deep learning-based methods. Traditional methods typically rely on prior manual features such as abnormal noise distribution, resampling traces, compression artifacts, and inconsistent lighting for detection. However, these methods depend on specific scenario assumptions and are easily affected by the diversity of tampering methods, differences in image sources, and post-processing operations in complex real-world environments, resulting in insufficient generalization ability and localization stability. Deep learning-based methods can automatically learn multi-level tampering representations and achieve pixel-level localization in an end-to-end manner, making them the current mainstream technical approach.
[0004] However, existing deep learning methods still have shortcomings: on the one hand, models tend to learn bias information related to the distribution of training data or scene semantics, making them prone to false positives and false negatives under cross-dataset and cross-tampering types conditions; on the other hand, in scenarios with complex boundaries, small-sized tampered regions, and post-processing interference such as compression, blurring, and noise, local artifacts and boundary details are easily weakened during feature extraction and fusion, leading to blurred boundaries and discontinuous regions in the localization results. Existing models such as TruFor, SparseViT, CoDE, DDG, Mesorch, and NCNeT mostly adopt single-granularity or simple multi-scale fusion, without explicit division of labor between local artifacts and global structure. The fusion mechanism is fixed, making it difficult to balance fine boundary characterization and region consistency modeling, thus limiting cross-scene generalization ability.
[0005] Therefore, there is an urgent need for a detection and localization method that can simultaneously characterize local fine-grained artifacts and global structure, reduce feature redundancy, adaptively fuse multi-granularity information, and maintain stability in complex scenes and cross-dataset conditions. Summary of the Invention
[0006] Technical Problem: To address the above problems, this invention proposes a highly generalizable image tampering detection and localization method based on multi-granularity representation decoupling and collaboration. By constructing fine-grained anomaly representations and coarse-grained structural representations, and explicitly decoupling and gating adaptively fusing different granularity representations, the method improves the generalization ability, boundary characterization ability, and regional localization stability of the image tampering detection and localization model in complex scenarios.
[0007] Technical solution: A highly generalized image tampering detection and localization method based on multi-granularity representation decoupling and collaboration, comprising the following steps:
[0008] The image to be detected is obtained and input into the image tampering detection and localization model to extract multi-level image features;
[0009] Based on the multi-level image features, a multi-granularity representation is constructed to obtain a fine-grained anomaly representation for representing local boundaries, texture anomalies, and fine-grained artifact information, and a coarse-grained structural representation for representing the internal consistency, contextual relationships, and overall structural information of the tampered region.
[0010] The fine-grained anomaly representation and the coarse-grained structural representation are explicitly decoupled to reduce information redundancy and functional coupling between different granular representations, resulting in decoupled fine-grained representation and decoupled coarse-grained representation.
[0011] Based on prior information about fake traces, gated adaptive fusion is performed on the decoupled fine-grained representation and the decoupled coarse-grained representation to obtain the fused tampering feature;
[0012] Based on the fused tampering features, a tampering region prediction result is generated, resulting in an image tampering detection result and a tampering region localization result.
[0013] By utilizing the image tampering detection results and tampering area location results, the detection and location of forged, spliced, copied and pasted, or deleted and repaired tampered areas in the image to be detected can be achieved.
[0014] Furthermore, the construction of multi-granularity representations includes the following steps:
[0015] The image to be detected is input into a sparse visual Transformer encoder, and shallow features are extracted through a pre-encoding stage to obtain shallow features used to construct artifact priors. ;
[0016] Based on the shallow features The noise response branch and compression response branch are constructed separately to obtain the noise response. and compression response ;
[0017] The noise response and compression response By merging, we obtain the pseudo-trace prior. The formula is:
[0018] ;
[0019] ;
[0020] ;
[0021] In the formula, , and These are mapping functions, This indicates a pooling operation. Indicates an upsampling operation. Indicates a channel splicing operation;
[0022] Based on the aforementioned artifact prior Guided sparse attention modeling to generate fine-grained anomaly representations and coarse-grained structure characterization .
[0023] Furthermore, the generation of fine-grained anomaly characterization and coarse-grained structural characterization includes the following steps:
[0024] The first The input features of each sparse coding unit are denoted as... Based on input features and artifact prior Estimate the selection probability corresponding to the candidate sparsity level The formula is:
[0025] ;
[0026] In the formula, This represents an adaptive selection function;
[0027] For the probability of selection Quantization is performed to obtain the sparsity level. The formula is:
[0028] ;
[0029] In the formula, Indicates quantization operation;
[0030] Based on the sparsity level For input features Perform sparse attention transformation to obtain the output features, as shown in the formula:
[0031] ;
[0032] In the formula, This represents the attention transformation controlled by the sparsity level;
[0033] Fine-grained anomaly representations of attention boundaries, texture breaks, and local statistical anomalies are generated using a dynamic sparse attention mechanism. By using a fixed sparse attention mechanism, coarse-grained structural representations with consistency of the region of interest, context dependency, and global structural constraints are generated. .
[0034] Furthermore, the fine-grained anomaly characterization and the coarse-grained structural characterization are explicitly decoupled, the steps of which include:
[0035] Characterization of fine-grained anomalies and coarse-grained structure characterization Global average pooling is performed, and the vectors are mapped to a unified embedding space through independent projection heads to obtain fine-grained embedding vectors. and coarse-grained embedding vectors The formula is:
[0036] ;
[0037] ;
[0038] In the formula, This indicates a global average pooling operation. and These represent the projection mappings for fine-grained branches and coarse-grained branches, respectively. This indicates a normalization operation;
[0039] Calculate the fine-grained embedding vector and coarse-grained embedding vectors explicit decoupling loss between The formula is:
[0040] ;
[0041] In the formula, Indicates batch size. and They represent the first The fine-grained embedding vector and coarse-grained embedding vector corresponding to each sample;
[0042] By minimizing the explicit decoupling loss, the similarity between fine-grained anomaly representations and coarse-grained structural representations in the embedding space is reduced, thereby enhancing the functional division of labor between representations of different granularities.
[0043] Further, gating adaptive fusion is performed, including the following steps:
[0044] By mapping multi-path features at different levels to a unified number of channels and a unified spatial scale, an aligned feature set is obtained. ;
[0045] Based on the aligned feature set and artifact prior Generate dynamic weights The formula is:
[0046] ;
[0047] In the formula, This represents a dynamic weight prediction function;
[0048] According to the dynamic weight Construct local branch representations respectively Context branch representation and basic fusion representation The formula is:
[0049] ;
[0050] ;
[0051] ;
[0052] In the formula, This represents the set of low-level detail feature indexes. Represents a set of high-level context feature indexes;
[0053] Based on the aforementioned basic fusion representation and artifact prior Constructing a gated guidance diagram The formula is:
[0054] ;
[0055] In the formula, Represents a mapping function. This represents the Sigmoid activation function;
[0056] According to the gate control guidance diagram A spatially relevant adaptive fusion of the local branch representation and the context branch representation is performed to obtain the fused tampering feature.
[0057] Furthermore, the fusion and tampering characteristics are obtained, and the steps include:
[0058] Representing local branches respectively and context branch representation Perform mapping to obtain local mapping features. and context mapping features The formula is:
[0059] ;
[0060] ;
[0061] In the formula, and These represent the local branch mapping function and the context branch mapping function, respectively.
[0062] A gate weight map is generated based on the local mapping features, context mapping features, and gate guidance graph. The formula is:
[0063] ;
[0064] In the formula, This represents the gating weight prediction function;
[0065] Based on the gate weight graph The local mapping features and context mapping features are weighted and fused to obtain the fused tampering features. The formula is:
[0066] ;
[0067] In the formula, Indicates the output mapping function, This represents element-wise multiplication;
[0068] Increase the fusion weight of local mapping features in locations near tampering boundaries, with complex local structures, or where artifact responses are significant; increase the fusion weight of context mapping features in locations that depend on overall structural constraints and regional stability.
[0069] Furthermore, the training steps for the image tampering detection and localization model include:
[0070] Calculate the primary localization loss based on the predicted and actual tampered areas. ;
[0071] The explicit decoupling loss is calculated based on the similarity between fine-grained anomaly representations and coarse-grained structural representations. ;
[0072] Calculate the gate regularization term based on the gate weight graph. This constrains the gating weights to prevent them from biasing towards a single branch too early in the training process.
[0073] The overall loss is calculated based on the principal localization loss, explicit decoupling loss, and gating regularization term. The formula is:
[0074] ;
[0075] In the formula, and These represent the balance coefficients of the explicit decoupling loss and the gated regularization term, respectively.
[0076] By minimizing the overall loss, the image tampering detection and localization model is trained to obtain an image tampering detection and localization model with cross-dataset generalization ability and fine localization ability.
[0077] Furthermore, the predicted tampered region includes a tampered probability map and a binarized tampered mask; wherein, the tampered probability map is used to represent the probability that each pixel in the image to be detected belongs to the tampered region, and the binarized tampered mask is obtained by thresholding the tampered probability map; based on the binarized tampered mask, the position, shape and boundary of the tampered region are determined, thereby achieving pixel-level localization of the tampered image.
[0078] Compared with the prior art, the present invention, employing the above technical solution, has the following technical effects:
[0079] 1. This invention constructs fine-grained anomaly representations and coarse-grained structural representations, and simultaneously models local boundary artifacts and overall regional consistency information of the tampered region within a unified framework, thus taking into account both boundary detail depiction and regional structure modeling.
[0080] 2. This invention reduces information redundancy and functional coupling between different granularities of representation by using explicit decoupling modules, enabling fine-grained branches to focus more on local texture, boundary transitions and statistical anomalies, and coarse-grained branches to focus more on regional consistency, contextual dependence and global structural constraints, thereby improving the model's ability to represent tamper-related clues.
[0081] 3. This invention uses a gated adaptive fusion module to dynamically adjust the fusion weights of features of different granularities according to the different needs for local detail information and context information at different spatial locations, thereby improving the positioning stability in complex boundaries, small-sized tampered areas and complex background interference scenarios. Attached Figure Description
[0082] Figure 1 This is a framework diagram of the high-generalization image tampering detection and localization method based on multi-granularity representation decoupling and collaboration of the present invention;
[0083] Figure 2 This is a visualization of the present invention and other models on different datasets;
[0084] Figure 3 This is a visualization of ablation experiments of the present invention under different module combinations;
[0085] Figure 4 This is a comparison chart of the robustness of the present invention with other models under different perturbation conditions. Detailed Implementation
[0086] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be further described in detail below with reference to the accompanying drawings. The described embodiments are only a part of the embodiments involved in this invention. All non-innovative embodiments based on these embodiments by other researchers in the art are within the protection scope of this invention. Furthermore, the step numbers in the embodiments of this invention are only for ease of explanation and do not limit the order of the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.
[0087] The highly generalized image tampering detection and localization method based on multi-granularity representation decoupling and collaboration described in this embodiment includes three stages: multi-granularity representation construction, explicit decoupling, and gated adaptive fusion.
[0088] In the multi-granularity representation construction stage, the image to be detected is input into the image tampering detection and localization model. Multi-level features are extracted through a sparse visual Transformer encoder, and fine-grained anomaly representations for representing boundary neighborhoods, texture breaks and local statistical anomalies are constructed, as well as coarse-grained structural representations for representing the consistency, contextual relationships and overall structural constraints within the tampered region.
[0089] In the explicit decoupling stage, fine-grained anomaly representations and coarse-grained structural representations are mapped to a unified embedding space. Explicit constraints reduce information redundancy and functional coupling between the two, enabling a clear division of labor between the two types of representations.
[0090] In the gated adaptive fusion stage, a gated fusion mechanism guided by artifact prior is introduced. Based on the different needs of local detail information and global context information at different spatial locations, the fusion weight of multi-granular features is dynamically adjusted, and finally, a tampering probability map and a binary tampering mask are generated to realize the detection and localization of tampered areas in the image.
[0091] In this embodiment, the specific implementation process of the high-generalization image tampering detection and localization method based on multi-granularity representation decoupling and collaboration includes the following steps:
[0092] Step 1: Obtain the image to be detected and input it into the image tampering detection and localization model to extract multi-level image features.
[0093] Furthermore, the image tampering detection and localization model includes a sparse ViT encoder, an explicit decoupling module, and a gated adaptive fusion module.
[0094] Given an input image, the backbone network is first used to extract features from the image to be detected, obtaining image features at different levels. Shallow features mainly include low-level information such as texture, edges, noise, and compression artifacts, while deep features mainly include region structure, contextual dependencies, and global semantic information. Joint modeling of features at different levels provides a foundation for subsequent multi-granularity representation construction.
[0095] Step 2: Construct multi-granularity representations based on multi-level image features to obtain fine-grained anomaly representations and coarse-grained structural representations.
[0096] Furthermore, the steps for constructing multi-granularity representations include:
[0097] The image to be detected is input into a sparse visual Transformer encoder, and shallow features are extracted through a pre-encoding stage to obtain shallow features used to construct artifact priors. ;
[0098] Considering that noise perturbation and compression distortion are typical and complementary tampering traces in the process of image manipulation, based on shallow features... The noise response branch and compression response branch are constructed separately to obtain the noise response. and compression response ;
[0099] noise response and compression response By merging them, a unified pseudo-trace prior is obtained. The formula is:
[0100] ;
[0101] ;
[0102] ;
[0103] In the formula, , and These are mapping functions, This indicates a pooling operation. Indicates an upsampling operation. This indicates a channel splicing operation.
[0104] In particular, the aforementioned artifact prior Instead of directly outputting the location of the tampering, it provides artifact cues related to the input content for subsequent sparse attention modeling, enabling the model to adaptively adjust the modeling granularity based on image content, artifact strength, and structural complexity.
[0105] Furthermore, the first The input features of each sparse coding unit are denoted as... Based on input features and artifact prior Estimate the selection probability corresponding to the candidate sparsity level The formula is:
[0106] ;
[0107] In the formula, This represents an adaptive selection function;
[0108] For the probability of selection Quantization is performed to obtain the sparsity level. The formula is:
[0109] ;
[0110] In the formula, Indicates quantization operation;
[0111] Based on sparsity level For input features Perform sparse attention transformation to obtain the output features, as shown in the formula:
[0112] ;
[0113] In the formula, This represents attention transformations controlled by sparse levels.
[0114] In this embodiment, the encoder primarily extracts basic visual features in the first two stages. The third stage employs a dynamic sparse attention mechanism to enhance the model's ability to perceive boundary neighborhoods, local anomalies, and fine-grained artifact cues, thereby obtaining fine-grained anomaly representations. The fourth stage employs a fixed sparse attention mechanism to aggregate region-level contextual information, enhancing the model's ability to model the consistency within tampered regions and the overall structural relationships, thus obtaining a coarse-grained structural representation. .
[0115] Step 3: Explicitly decouple the fine-grained anomaly characterization from the coarse-grained structural characterization.
[0116] Furthermore, the steps for explicitly decoupling fine-grained anomaly characterization and coarse-grained structural characterization include:
[0117] Fine-grained anomaly characterization and coarse-grained structural characterization are denoted as follows: and , respectively and Global average pooling is performed, and the vectors are mapped to a unified embedding space through independent projection heads to obtain fine-grained embedding vectors. and coarse-grained embedding vectors The formula is:
[0118] ;
[0119] ;
[0120] In the formula, This indicates a global average pooling operation. and These represent the projection mappings for fine-grained branches and coarse-grained branches, respectively. This indicates a normalization operation.
[0121] To further reduce redundant coupling between the two types of representations, the explicit decoupling loss is defined as minimizing their similarity in the normalized embedding space:
[0122] ;
[0123] In the formula, Indicates batch size. and They represent the first The fine-grained embedding vector and coarse-grained embedding vector corresponding to each sample.
[0124] By minimizing the explicit decoupling loss, the similarity between fine-grained anomaly representations and coarse-grained structural representations in the embedding space is reduced. This encourages fine-grained branches to focus more on boundary, texture, local anomaly, and artifact information, and encourages coarse-grained branches to focus more on region consistency, contextual relationships, and overall structural constraints. This reduces invalid repetitive encoding and enhances the complementarity between different granularity representations.
[0125] Step 4: Based on the prior information of the pseudo-trace, perform gated adaptive fusion on the decoupled fine-grained representation and coarse-grained representation.
[0126] Furthermore, the steps for gating adaptive fusion include:
[0127] By mapping multi-path features at different levels to a unified number of channels and a unified spatial scale, an aligned feature set is obtained. ;
[0128] Based on the aligned feature set and artifact prior Generate dynamic weights The formula is:
[0129] ;
[0130] In the formula, This represents a dynamic weight prediction function;
[0131] Based on dynamic weights Construct local branch representations respectively Context branch representation and basic fusion representation The formula is:
[0132] ;
[0133] ;
[0134] ;
[0135] In the formula, This represents the set of low-level detail feature indexes. This represents the set of high-level contextual feature indices. Local branches represent... The focus is primarily on local texture, boundary transitions, and fine-grained artifact information, with contextual branching representation. The main focus is on region consistency, global semantics, and structural constraints, with a fundamental fusion representation. The base response is used to integrate features at different levels.
[0136] Subsequently, based on the fundamental fusion representation and artifact prior Constructing a gated guidance diagram The formula is:
[0137] :
[0138] In the formula, Represents a mapping function. This represents the Sigmoid activation function.
[0139] Furthermore, the local branches are represented separately. and context branch representation Perform mapping to obtain local mapping features. and context mapping features The formula is:
[0140] ;
[0141] ;
[0142] In the formula, and These represent the local branch mapping function and the context branch mapping function, respectively.
[0143] Gating weight graph is generated based on local mapping features, context mapping features, and gating guidance graph. The formula is:
[0144] ;
[0145] In the formula, This represents the gating weight prediction function;
[0146] Based on gated weight graph The local mapping features and context mapping features are weighted and fused to obtain the fused tampering features. The formula is:
[0147] ;
[0148] In the formula, Indicates the output mapping function, This indicates element-wise multiplication.
[0149] In particular, gating weight graphs are used in locations near tampering boundaries, with complex local structures, or where artifact responses are significant. Increase the fusion weights of local mapping features; in locations where the overall structural constraints and regional stability are more dependent, gate the weight map. Increase the fusion weights of context mapping features. This allows the model to adaptively fuse features of different granularities based on spatial location differences.
[0150] Step 5: Generate the predicted tampered region based on the fused tampering features.
[0151] Furthermore, it will integrate tampering features. A prediction head is input to generate a tampering probability map, which represents the probability that each pixel in the image to be detected belongs to a tampered region. The tampering probability map is then thresholded to obtain a binarized tampering mask. Based on the binarized tampering mask, the location, shape, and boundary of the tampered region are determined, achieving pixel-level detection and localization of tampered regions such as those resulting from splicing, copying and pasting, and deletion repair in the image to be detected.
[0152] Step 6: Train the image tampering detection and localization model.
[0153] Furthermore, the training steps for the image tampering detection and localization model include:
[0154] Calculate the primary localization loss based on the predicted and actual tampered areas. ;
[0155] The explicit decoupling loss is calculated based on the similarity between fine-grained anomaly representations and coarse-grained structural representations. ;
[0156] Calculate the gate regularization term based on the gate weight graph. This constrains the gating weights to prevent them from biasing towards a single branch too early in the training process.
[0157] The overall loss is calculated based on the primary localization loss, explicit decoupling loss, and gating regularization term, using the following formula:
[0158] ;
[0159] In the formula, and These represent the balance coefficients of the explicit decoupling loss and the gating regularization term, respectively.
[0160] By minimizing the overall loss, the image tampering detection and localization model is trained so that the model can simultaneously learn local artifact clues and overall structural relationships of the tampered region, and output stable tampering detection results and tampered region localization results.
[0161] This embodiment employs a three-stage process—multi-granularity representation construction, explicit decoupling, and gated adaptive fusion—to generate image tampering detection and localization results with strong generalization and fine-grained localization capabilities. The multi-granularity representation construction stage simultaneously captures boundary details and regional structural information. The explicit decoupling stage reduces information redundancy between representations of different granularities, enhancing functional division of labor. The gated adaptive fusion stage dynamically adjusts the fusion ratio of local detail features and contextual structural features based on spatial location differences, thereby achieving accurate detection and localization of tampered areas.
[0162] To further illustrate the effectiveness of the highly generalized image tampering detection and localization method based on multi-granularity representation decoupling and collaboration described in this invention, the following examples are provided.
[0163] This embodiment employs a multi-source data joint training and cross-dataset testing experimental setup. The training set consists of TampCOCO, CASIA_V2, FantasticReality_v1, and IMD2020, while the test set includes Coverage, Columbia, CASIA_V1, and NIST16. The training data covers various tampering methods such as splicing, copying and pasting, deletion and repair, boundary blurring, and color inconsistency, providing diverse supervision samples for the model. The test data differs in distribution from the training data, used to verify the model's cross-dataset generalization ability.
[0164] For image tampering detection and localization tasks, this embodiment uses pixel-level F1 score, Intersection over Union (IoU), and Area Under the ROC Curve (AUC) as evaluation metrics. The F1 score measures the model's overall performance in terms of precision and recall; IoU measures the overlap between the predicted tampered region and the actual tampered region; and AUC measures the model's overall ability to distinguish between tampered pixels and real pixels under different threshold conditions.
[0165] Table 1. Comparison of F1 and IoU for different models with a fixed threshold of 0.5 on the test set.
[0166]
[0167] As shown in Table 1, the present invention exhibits good detection and localization performance on multiple test sets, particularly achieving high F1 and IoU values on the NIST16 and Coverage datasets, indicating that the present invention maintains good localization capabilities under complex tampering scenarios and copy-paste tampering scenarios. Furthermore, the present invention also maintains stable performance on the CASIA_V1 and Columbia datasets, demonstrating that the multi-granularity representation decoupling and collaboration mechanism has good cross-dataset adaptability.
[0168] The AUC comparison results of different models on the test set are shown in Table 2.
[0169] Table 2. Comparison of AUC of different models on the test set
[0170]
[0171] As shown in Table 2, the present invention also exhibits good performance in terms of the AUC metric. The AUC metric can reflect the overall separability and stability of the model's output probability map from multiple threshold perspectives. The results shown in Table 2 demonstrate that the tampering probability map generated by the present invention can effectively distinguish between tampered regions and real regions, providing a reliable foundation for the subsequent generation of a binarized tampering mask.
[0172] like Figure 2 As shown, under different datasets and complex scenarios, the tamper location results output by this invention are quite close to the true mask. Compared with other models, when there are complex backgrounds, small target scales, or irregular shapes in the image, this invention can better suppress background noise interference, reduce scattered false detections, and maintain the integrity and continuity of the predicted region. Figure 2 To further explain, the present invention can not only detect tampered areas, but also accurately delineate the boundaries of tampered areas.
[0173] Table 3 Ablation Experiment Results
[0174]
[0175] Table 3 shows that the F1 score, IoU, and AUC are generally improved after gradually adding multi-granularity representation construction and explicit decoupling modules to the basic model. Adding multi-granularity representation construction allows the model to better capture local anomaly clues in suspected tampered areas; adding explicit decoupling modules clarifies the functional division between fine-grained anomaly representations and coarse-grained structural representations. Further comparison of different fusion methods reveals that additive fusion and prior-gated fusion can improve localization performance to some extent, while stitching fusion is less effective. Prior-gated fusion achieves the highest F1 score, IoU, and AUC, indicating that artifact priors can effectively guide the model to dynamically adjust the fusion weights of local and contextual information based on different spatial locations, thereby improving localization quality.
[0176] like Figure 3 As shown, with the introduction of multi-granularity representation construction and explicit decoupling modules, the model prediction results are improved in terms of regional integrity, boundary clarity, and small-region tampering response capability. Compared with additive fusion, splicing fusion, and prior-gated fusion, the complete model using prior-gated fusion can reduce false detections in the background region and fragmentation of the predicted region, and its localization results are closer to the true mask, indicating that multi-granularity representation construction, explicit decoupling, and prior-gated fusion can synergistically improve the model's tampering detection and localization capabilities.
[0177] Furthermore, to verify the stability of the invention under common post-processing conditions, this embodiment performs robustness comparisons under three types of perturbations: JPEG compression, Gaussian blur, and Gaussian noise. The robustness comparison results are as follows: Figure 4 As shown, where Figure 4 (a) shows the robustness comparison results under JPEG compression perturbation. Figure 4 (b) shows the robustness comparison results under Gaussian blur perturbation. Figure 4 (c) in the figure represents the robustness comparison results under Gaussian noise perturbation.
[0178] Depend on Figure 4 It can be seen that under different perturbation conditions, the present invention maintains relatively stable detection and localization performance, indicating that the multi-granularity characterization decoupling and synergistic mechanism described in the present invention can enhance the robustness of the model under complex propagation scenarios such as compression distortion, boundary smoothing and noise pollution.
[0179] In this embodiment, through Figure 1 The overall framework shown completes multi-granularity representation construction, explicit decoupling, and gated adaptive fusion; through Figure 2 This document presents a visual comparison of the tampering localization capabilities of different models; it verifies the cross-dataset detection and localization capabilities of the models from the perspectives of F1, IoU, and AUC using Tables 1 and 2; and it further verifies the tampering localization capabilities of the models using Tables 3 and 4. Figure 3 Verify the effectiveness of each key module; through Figure 4The robustness of the model under common post-processing perturbations was verified. This demonstrates that the method of this invention can achieve high-generalization detection and fine-grained localization of image tampering regions in complex scenarios.
[0180] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A highly generalized image tampering detection and localization method based on multi-granularity representation decoupling and collaboration, characterized in that, The steps include the following: Acquire the image to be detected, input it into the image tampering detection and localization model, and extract multi-level image features; Based on the multi-level image features, a multi-granularity representation is constructed to obtain a fine-grained anomaly representation for representing local boundaries, texture anomalies, and fine-grained artifact information, and a coarse-grained structural representation for representing the internal consistency, contextual relationships, and overall structural information of the tampered region. The fine-grained anomaly representation and the coarse-grained structural representation are explicitly decoupled to reduce information redundancy and functional coupling between different granular representations, resulting in decoupled fine-grained representation and decoupled coarse-grained representation, forming a differentiated functional division of labor. Based on prior information about fake traces, gated adaptive fusion is performed on the decoupled fine-grained representation and the decoupled coarse-grained representation to obtain the fused tampering feature; Based on the fused tampering features, a tampering region prediction result is generated, resulting in an image tampering detection result and a tampering region localization result. By utilizing the image tampering detection results and tampering area location results, the detection and location of forged, spliced, copied and pasted, or deleted and repaired tampered areas in the image to be detected can be achieved.
2. The high-generalization image tampering detection and localization method according to claim 1, characterized in that, The steps for constructing the multi-granularity representation include: The image to be detected is input into a sparse visual Transformer encoder, and shallow features are extracted through a pre-encoding stage to obtain shallow features used to construct artifact priors. ; Based on the shallow features The noise response branch and compression response branch are constructed respectively to obtain the noise response. and compression response ; The noise response and compression response By merging, we obtain the pseudo-trace prior. The formula is: ; ; ; In the formula, , and These are mapping functions, This indicates a pooling operation. [·] indicates an upsampling operation, and [·] indicates a channel concatenation operation; Based on the aforementioned artifact prior Guided sparse attention modeling to generate fine-grained anomaly representations and coarse-grained structure characterization .
3. The high-generalization image tampering detection and localization method according to claim 2, characterized in that, The steps for generating the fine-grained anomaly characterization and coarse-grained structural characterization include: The first The input features of each sparse coding unit are denoted as... Based on input features and artifact prior Estimate the selection probability corresponding to the candidate sparsity level The formula is: ; In the formula, This represents an adaptive selection function; For the probability of selection Quantization is performed to obtain the sparsity level. The formula is: ; In the formula, Indicates quantization operation; Based on the sparsity level For input features Perform sparse attention transformation to obtain output features. The formula is: ; In the formula, This represents the attention transformation controlled by the sparsity level; Fine-grained anomaly representations of attention boundaries, texture breaks, and local statistical anomalies are generated using a dynamic sparse attention mechanism. By using a fixed sparse attention mechanism, coarse-grained structural representations with consistency of the region of interest, context dependency, and global structural constraints are generated. .
4. The high-generalization image tampering detection and localization method according to claim 3, characterized in that, The steps for explicitly decoupling the fine-grained anomaly characterization and the coarse-grained structural characterization include: Characterization of fine-grained anomalies and coarse-grained structure characterization Global average pooling is performed, and the vectors are mapped to a unified embedding space through independent projection heads to obtain fine-grained embedding vectors. and coarse-grained embedding vectors The formula is: ; ; In the formula, This indicates a global average pooling operation. and These represent the projection mappings for fine-grained branches and coarse-grained branches, respectively. express Normalization operation; Calculate the fine-grained embedding vector and coarse-grained embedding vectors explicit decoupling loss between The formula is: ; In the formula, Indicates batch size. Indicates transpose. and They represent the first The fine-grained embedding vector and coarse-grained embedding vector corresponding to each sample; By minimizing the explicit decoupling loss, the similarity between fine-grained anomaly representations and coarse-grained structural representations in the embedding space is reduced, thereby enhancing the functional division of labor between representations of different granularities.
5. The high-generalization image tampering detection and localization method according to claim 4, characterized in that, The steps for performing the gated adaptive fusion include: By mapping multi-path features at different levels to a unified number of channels and a unified spatial scale, an aligned feature set is obtained. ; Based on the aligned feature set and artifact prior Generate dynamic weights The formula is: ; In the formula, This represents a dynamic weight prediction function; According to the dynamic weight Construct local branch representations respectively Context branch representation and basic fusion representation The formula is: ; ; ; In the formula, This represents the set of low-level detail feature indexes. Represents a set of high-level context feature indexes; Based on the aforementioned basic fusion representation and artifact prior Constructing a gated guidance diagram The formula is: ; In the formula, Represents a mapping function. This represents the Sigmoid activation function.
6. The high-generalization image tampering detection and localization method according to claim 5, characterized in that, According to the gate control guidance diagram A spatially relevant adaptive fusion of the local branch representation and the context branch representation is performed to obtain the fused tampering feature.
7. The high-generalization image tampering detection and localization method according to claim 6, characterized in that, To obtain the fusion tampering features, the steps include: Representing local branches respectively and context branch representation Perform mapping to obtain local mapping features. and context mapping features The formula is: ; ; In the formula, and These represent the local branch mapping function and the context branch mapping function, respectively. Based on the local mapping features Context mapping features and gate control guidance diagram Generate a gated weight graph The formula is: ; In the formula, This represents the gating weight prediction function; Based on the gate weight graph The local mapping features and context mapping features are weighted and fused to obtain the fused tampering features. The formula is: ; In the formula, Indicates the output mapping function, This indicates element-wise multiplication.
8. The high-generalization image tampering detection and localization method according to claim 7, characterized in that, The adaptive fusion increases the fusion weight of local mapping features near tampering boundaries, where the local structure is complex, or where there are artifact responses; and increases the fusion weight of context mapping features at locations that depend on overall structural constraints and regional stability.
9. The high-generalization image tampering detection and localization method according to any one of claims 1 to 8, characterized in that, The training steps for the image tampering detection and localization model include: Calculate the primary localization loss based on the predicted and actual tampered areas. ; The explicit decoupling loss is calculated based on the similarity between fine-grained anomaly representations and coarse-grained structural representations. ; Calculate the gate regularization term based on the gate weight graph. This constrains the gating weights and prevents bias towards a single branch in the early stages of training. The overall loss is calculated based on the principal localization loss, explicit decoupling loss, and gating regularization term. The formula is: ; In the formula, and These represent the balance coefficients of the explicit decoupling loss and the gated regularization term, respectively. By minimizing the overall loss, the image tampering detection and localization model is trained to obtain an image tampering detection and localization model with cross-dataset generalization ability and fine localization ability.
10. The high-generalization image tampering detection and localization method according to any one of claims 1 to 8, characterized in that, The predicted tampering region includes a tampering probability map and a binarized tampering mask: The tampering probability map is used to represent the probability that each pixel in the image to be detected belongs to the tampered region, and the binarized tampering mask is obtained by thresholding the tampering probability map; The location, shape, and boundary of the tampered area are determined based on the binarized tampering mask, thereby achieving pixel-level localization of the tampered image.