Object-level prototype modeling-based copy movement tampering detection method and system
Through the alternating update of object-level prototype modeling and suspicious area re-identification methods, the error problem of similar area detection in copy-move tamper detection is solved, and tampering area detection is achieved with high accuracy and integrity.
Patent Information
- Application Number
- CN202510828070.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-08-29
AI Technical Summary
The prior art cannot accurately obtain the complete similar areas of the object in copy-move tamper detection, and the error of the detection result cannot be effectively corrected. Especially after the post-processing operation, the detection error cannot be guaranteed, and the completeness and accuracy of the detection cannot be guaranteed.
The alternating update module is designed to iteratively optimize the tampering domain and source domain prototype, and combine the suspicious area extraction module and the re-identification module to generate the final detection result through multi-scale multi-dimensional inconsistency detection and shallow feature fusion.
Through object-level prototype modeling, precise detection of replicated mobile tampering areas can be detected to reduce false positives, enhance target area detection, and ensure the structural integrity and accuracy of the detection results.
Smart Images

Figure CN120564017A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image tampering detection, and in particular to a copy-move tampering detection method and system based on object-level prototype modeling. Background Art
[0002] The copy-and-move operation is a common form of image tampering. This operation copies a region of an image (the source domain) and pastes it into another region of the same image (the tampered domain), changing the meaning of the entire image. Because the tampered region and the source domain in the copy-and-move tampering detection image are from the same image, their brightness and contrast are highly similar, which undoubtedly poses a challenge to copy-and-move forgery detection. Furthermore, forgers often perform post-processing operations on the tampered image, such as blurring and scaling, which also poses a major challenge to copy-and-move detection.
[0003] Currently, cutting-edge tampering detection methods essentially only calculate the correlation between pixels to obtain similar region features and achieve detection and differentiation between the source domain and the tampered domain. However, this method has the following problems:
[0004] First, the autocorrelation calculation method for images cannot accurately obtain the complete similar area of the object. The tampered area is obtained by copying the original domain to other locations in the image, so there are differences in the surrounding pixels of the source domain and the tampered domain. To hide the traces of tampering, tamperers usually perform post-processing operations such as blurring on the tampered image. After blurring, the differences in the surrounding pixels of the source domain and the tampered domain will lead to differences in the edge area information of similar target objects. In this case, detecting pixel areas only through pixel-by-pixel similarity calculations is very likely to cause errors in the detection of target edge areas, and cannot guarantee the integrity of the detection of similar objects.
[0005] Second, the detection results are not effectively corrected to obtain accurate results. Specifically, existing methods fail to reversely detect misjudged similar regions from the perspective of inconsistency, including supplementing missing target regions and removing false positive background regions to refine the detection results. The existing method ARNet refines the coarse similar regions obtained through similarity calculation by adding a skip residual structure. This only further extracts features from the coarse features, and does not further optimize the removal and supplementation of coarse similar regions from the essence of finding related regions. Summary of the Invention
[0006] The present invention aims to provide a copy-move tampering detection method and system based on object-level prototype modeling to address the issues raised in the aforementioned background technology. Specifically, the present invention designs an alternating update module that iteratively optimizes the tampered domain prototype and the source domain prototype to refine the features of similar regions; a suspicious region extraction module that extracts potential tampered regions based on multi-scale and multi-dimensional inconsistency detection; a suspicious region re-identification module that performs secondary verification of suspicious regions by combining shallow features; and a prototype-guided differentiation module that fuses multi-level features to generate the final detection results.
[0007] To achieve the above object, the present invention provides the following technical solutions: The copy-move tampering detection method based on object-level prototype modeling includes the following steps: Steps for correlation detection: Input the copied and moved tampered image, extract its features, and obtain the coarse similarity region feature Fc through correlation detection; Steps for prototype and similar area update: Initialize two sets of feature prototypes Pt and Ps, which are used to represent the tampered domain and the source domain respectively. Alternately update the two prototypes with the similar region feature Fc to generate similar region features Fct and Fcs; Steps for extracting suspicious areas: By analyzing the differences between Fct and Fcs, we can find the inconsistent regions between the two. Specifically, the regions detected as targets in both Fct and Fcs are high-confidence regions, and the features of high-confidence regions are recorded as Fh. In addition to the high-confidence regions, the other regions detected in Fct and Fcs are inconsistent regions, and the features of inconsistent regions are recorded as Fs. Steps for re-identifying suspicious areas: The suspicious area in Fs is located by extracting the mid- and deep-layer features from the input image through convolution operations, masking the irrelevant background, and obtaining the suspicious area feature Fac through correlation extraction and feature fusion; The steps of prototype guidance differentiation: A detection mask is generated based on the fusion features generated by Fh and Fac, and the intersection-over-union matching of the source domain and the tampered domain prototype obtained by the latest update is used to automatically distinguish the domain to which the target object belongs as the source domain or the tampered domain, and output the final detection result.
[0008] Furthermore, the specific methods for updating the prototype and similar region features include: Prototype updates: Using Pt and Ps as query features, and Fc as key and value in the cross-correlation modeling CCM process, through this mechanism, Pt and Ps are updated to generate P (t+1) and P (s+1) , the mathematical formula of the prototype update process is expressed as follows: ; ; Among them, P (t+1) and P (s+1) They represent the updated tampered domain prototype and source domain prototype respectively, conv( ) represents the convolution operation, T represents the transposition operation, and “·” represents the dot product operation. represents the square root normalization of the Fc dimension; Updates in similar areas: After completing the prototype update, use P (t+1) and P (s+1) To refine the coarse similarity region feature Fc; Specifically, using P (t+1) and P (s+1) To guide the feature Fc, generate similar region features Fct and Fcs associated with the tampered region and the source region respectively; by adopting this alternating update strategy, the feature representation is dynamically adjusted, and the mathematical formula of the similar region update process is as follows: ; ; in and are the dimension normalization factors of the tampered domain prototype and the source domain prototype, respectively.
[0009] Furthermore, in the step of extracting suspicious regions, the method for obtaining high confidence regions is as follows: the features Fct and Fcs are processed using the ReLU activation function to filter out background noise and retain only significant features; after activation, element-wise multiplication is performed between Fct and Fcs to obtain the high confidence region feature Fh.
[0010] Furthermore, in the step of suspicious region extraction, the suspicious region is defined as a location where there is a detection difference between Fct and Fcs, that is, one is detected as background and the other is detected as a target; the suspicious region is extracted by calculating the inconsistency between the corresponding locations in two similar regions with Fct and Fcs features. Specifically: Step (1), original scale consistency detection: Calculate the spatial inner product M1 and channel inner product M of Fct and Fcs at the corresponding position through the sliding window k , and perform convolution operation after splicing the two, and output the original scale consistent area M in Fct and Fcs at the original scale; Step (2), multi-scale sliding window detection: By sliding the window between M1 and M kRepeat the consistency detection operation of step (1) above to output the multi-scale consistent region ; Step (3), multi-scale feature merging and suspicious area positioning: Multi-scale consistent regions The consistent regions M at the original scale are merged to obtain the final consistent regions in Fct and Fcs at multiple scales. Subsequently, the inconsistent regions are separated from the consistent regions by reverse gating to obtain the suspicious regions.
[0011] Furthermore, the steps for re-identifying suspicious areas are as follows: The features F4 and F5 obtained by the fourth and fifth convolution operations of the input image are used to re-identify the suspicious areas. Specifically, the suspicious areas in Fs are mapped to the features F4 and F5 to locate the suspicious areas in F4 and F5. The elements in F4 and F5 corresponding to the suspicious locations are retained, while other elements including the background and high-confidence areas are masked. The masked F4 and F5 are defined as F4m and F5m respectively. Then, correlation extraction is performed on the mask features of F4m and F5m to obtain features F4c and F5c respectively. Finally, F4c and F5c are adaptively fused to generate Fac.
[0012] Furthermore, the steps of prototype-guided differentiation include: Feature fusion and classification: First, Fh and Fac are adaptively fused to obtain the fused feature Fan. The fused feature Fan is subjected to a softmax operation to implement element classification and obtain the detection mask Ma, which includes the two target objects O1 and O2. Prototype classification: Softmax operations are performed on the source domain prototype and the tampered domain prototype to achieve element classification, and the source domain and tampered domain detection masks Ms and Mt are obtained respectively; Target matching: Calculate the intersection and union ratio of the target parts Os and Ot in Ms and Mt with O1 respectively: If the intersection-over-union ratio of Os and O1 is greater than that of Ot and O1, then the target of O1 in Ma is the source domain and the target of O2 is the tampered domain; If the intersection-over-union ratio of Ot and O1 is greater than that of Os and O1, then the target of O1 in Ma is the tampered domain and the target of O2 is the source domain.
[0013] Furthermore, the total loss function is: ; Among them, Ldet represents the detection loss calculated using cross entropy, Lssim represents the use of SSIM loss to ensure the integrity of the spatial structure of the detected tampered area and the source area, and α is a learnable parameter; ; Among them, pi∈{0,1} indicates whether the i-th pixel belongs to the area to be detected. Indicates the probability that the pixel is predicted to belong to the target area, w1 and w2 are hyperparameters; ; Where x and y are the true value and the detection result respectively, u x and u y represents the mean of x and y, and represents the standard deviation of x and y, represents the covariance of x and y, and is a hyperparameter.
[0014] The present invention also provides a copy-move tampering detection system based on object-level prototype modeling, which is used to implement the copy-move tampering detection method based on object-level prototype modeling as described above. The detection system includes a coarse similarity region feature extraction module, an alternating update module, a suspicious region extraction module, a suspicious region re-identification module, and a prototype-guided differentiation module; The coarse similar region feature extraction module is used to extract the coarse similar region features Fc of the input image; The alternating update module is used to update the prototype and coarse similar region features Fc, and output similar region features Fct and Fcs associated with the tampered region and the source region; The suspicious region extraction module is used to extract high confidence region features Fh and inconsistent region features Fs from Fct and Fcs; The suspicious area re-identification module is used to re-identify the suspicious area extracted by the suspicious area extraction module to obtain the suspicious area feature Fac; The prototype-guided differentiation module generates a detection mask through the fusion features generated by the high-confidence region features Fh and the suspicious region features Fac, and automatically distinguishes the domain to which the target object belongs as the source domain or the tampered domain by using the intersection-over-union matching of the updated source domain and tampered domain prototypes, and outputs the final detection result.
[0015] Compared with the prior art, the present invention has the following beneficial effects: (1) We propose setting object-level prototypes of the source domain and the tampered domain to optimize similar region detection, ensuring the structural integrity of the detection object. By modeling the source domain prototype and mining the local inconsistencies between the source domain prototype and the similar region, we can obtain the characteristics of the tampered region. Similarly, by modeling the tampered domain prototype and mining the local inconsistencies between the tampered domain prototype and the similar region, we can obtain the characteristics of the source domain. By introducing the object-level modeling method of the source domain and the tampered domain prototype, we can obtain the detection target with complete structure through implicit optimization methods.
[0016] (2) By alternately updating the introduced source domain and tampered domain prototypes and the coarse similar region features obtained through correlation detection, the features Fct and Fcs representing the similar regions are obtained. These features are input into the suspicious region for suspicious region extraction. The regions where there are differences between the two are obtained through inconsistency mining. These different regions are then mapped to shallow features containing more detailed information through suspicious region re-identification, and further similar region detection is performed. The detection results are adaptively fused with the coarse similar features obtained through the similar region extraction module to supplement the missing regions and remove the background areas of false positives.
[0017] Unlike directly fusing Fct and Fcs optimized from the source and tampered domains, the method of this invention first explores inconsistencies between the two, then maps only the inconsistent regions and roughly similar regions to the shallow layer, further extracting correlations from the suppressed shallow features. The mapping process is essentially a decoupling operation, eliminating the influence of non-target background and high-confidence target regions. This ensures that correlation extraction focuses the network on suspicious regions, thereby resolving the problem of misjudgment caused by suspicious regions being obscured by high-confidence regions.
[0018] In addition, the selection of F4 and F5 for mapping is based on the fact that shallow features retain effective detail information, which is often lost in the deep layers of the network. Therefore, re-identifying similar regions in F4 and F5 and using them as supplementary high-confidence regions not only reduces false positives but also enhances the detection of target regions. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 It is a schematic diagram of the overall structure of the present invention; Figure 2 This is a visual display of the comparative experiment in the embodiment. DETAILED DESCRIPTION
[0020] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0021] See also Figure 1,The present invention provides a copy and move tampering detection method based on object-level ,prototype modeling, which is implemented through the following core modules: ,Alternating update module: iteratively optimizes the tampering domain prototype and the source ,domain prototype to refine the features of similar regions; ,Design of a suspicious region extraction module: extracts potential tampering regions based on ,multi-scale and multi-dimensional inconsistency detection; ,Design of a suspicious region re-identification module: combines shallow features to perform secondary verification on the ,suspicious region; ,Design of a prototype-guided differentiation module: integrates multi-level features to generate the ,final detection results.
[0022] As an example, a copy-move tampering detection method based on object-level prototype modeling includes the following steps: The step of correlation detection obtains the coarse similarity region feature Fc.
[0023] The prototype and similar region updating steps generate similar region features Fct and Fcs.
[0024] The step of extracting suspicious regions is to obtain high confidence regions and inconsistent regions in Fct and Fcs. The features of high confidence regions are recorded as Fh, and the features of inconsistent regions are recorded as Fs.
[0025] The step of suspicious area re-identification obtains the suspicious area in Fs and obtains the suspicious area feature Fac.
[0026] The prototype-guided differentiation step automatically distinguishes the domain to which the target object belongs as the source domain or the tampered domain based on Fh, Fac and the latest updated prototypes of the source domain and the tampered domain, and outputs the final detection result.
[0027] The implementation of each step is introduced in detail below.
[0028] 1. Steps of correlation detection: Input the copied, moved, and tampered image with a size of h*w, perform feature extraction on it, and obtain the coarse similarity region feature Fc through correlation detection. The size of Fc is H*W*C, where H, W, and C are height, width, and channel respectively.
[0029] 2. Steps for updating prototypes and similar areas: Initialize two sets of feature prototypes Pt and Ps, both of size H*W*C, to represent the tampered domain and the source domain respectively. Alternately update the two prototypes with the similar region feature Fc to generate similar region features Fct and Fcs, both of size H*W*C.
[0030] As a preferred embodiment, in step 2, an alternating update module is used to iteratively update the prototype and similar region features. The coarse similar region features Fc, the tampered domain prototype Pt, and the source domain prototype Ps are input into the alternating update module. The specific method is: Step 201: Prototype update: Using Pt and Ps as query features, and Fc as key and value in the cross-correlation modeling (CCM) process, cross-correlation modeling is a mechanism for modeling feature dependencies by calculating the cross-correlation between features. Through this mechanism, Pt and Ps are updated to make them more focused on key areas related to tampering (such as the source area and tampering area in copy-move tampering), generating P (t+1) and P (s+1) , with a size of H×W×C. The purpose of this iterative update strategy is to gradually refine the feature representation and force Pt and Ps to be closer to their corresponding target regions: Pt corresponds to the tampered region and Ps corresponds to the source domain.
[0031] The mathematical formula of this process is as follows: ; ; Among them, P (t+1) and P (s+1) They represent the updated tampered domain prototype and source domain prototype respectively, conv() represents the convolution operation, T represents the transposition operation, and “·” represents the dot product operation. represents the square root normalization of the Fc dimension.
[0032] Step 202: Update of similar regions: After completing the prototype update, use P (t+1) and P (s+1) To refine the coarse similarity region feature Fc. Specifically, use P (t+1) and P (s+1) To guide the feature Fc, generate similar region features Fct and Fcs associated with the tampered region and the source region respectively; by adopting this alternating update strategy, the feature representation is dynamically adjusted to make the model more sensitive to similar regions. The mathematical formula of this optimization process is as follows: ; ; Where Fct represents the updated tampering domain prototype P (t+1) Guided similar region features, Fcs represents the updated source domain prototype P (s+1) The sizes of the guided similarity region features, Fct and Fcs, are H×W×C. and are the dimension normalization factors of the tampered domain prototype and the source domain prototype, respectively.
[0033] The effectiveness of dynamic optimization is reflected in the alternating updates of the two prototypes and the similarity region feature Fc. Specifically, Fc, obtained through correlation detection, can guide Pt and Ps to focus on the tampered region and the source region, respectively, while establishing a clearer association between the two. In addition, the dynamically optimized Fct and Fcs provide feature descriptions of the similarity region from different perspectives. These features can effectively capture the details of the tampered region and the source region, and are further optimized by complementing each other, thereby obtaining a more refined similarity target.
[0034] 3. Steps for extracting suspicious areas: By analyzing the differences between Fct and Fcs, we can find inconsistent regions between the two. Specifically, the regions detected as targets in both Fct and Fcs are high-confidence regions, and the high-confidence region features are recorded as Fh; in addition to the high-confidence regions, the other regions detected in Fct and Fcs are inconsistent regions, and the inconsistent region features are recorded as Fs. Although these regions may appear insignificant compared to the consistent and dominant high-confidence regions, they often show subtle patterns that indicate tampering. The role of the suspicious region extraction module is to isolate these regions for further inspection to ensure that potential key information is not lost during the detection process.
[0035] As a preferred embodiment, in this step, the method for obtaining the high confidence region is: using the ReLU activation function to process the features Fct and Fcs to filter out background noise and retain only significant features; after activation, performing element multiplication between Fct and Fcs to obtain the high confidence region feature Fh.
[0036] As a preferred embodiment, in this step, the suspicious region is defined as a location where there is a detection difference between Fct and Fcs, that is, one is detected as background and the other is detected as a target; the suspicious region is extracted by calculating the inconsistency between corresponding locations in two similar regions with Fct and Fcs features. Specifically:
[0037] Step (1), original scale consistency detection: Calculate the spatial inner product M1 and channel inner product M of Fct and Fcs at the corresponding position through the sliding window k , and concatenate the two and perform a convolution operation to output the original scale consistent region M in Fct and Fcs at the original scale.
[0038] Step (2), multi-scale sliding window detection: By sliding the window between M1 and M k Repeat the consistency detection operation of step (1) above to output the multi-scale consistent region .
[0039] Step (3), multi-scale feature merging and suspicious area positioning: Multi-scale consistent regions The consistent regions M at the original scale are merged to obtain the final consistent regions in Fct and Fcs at multiple scales. Subsequently, the inconsistent regions are separated from the consistent regions by reverse gating to obtain the suspicious regions.
[0040] As an example, in this embodiment, multi-scale (1×1 and 2×2 sliding windows) and multi-dimensional (position and channel) extraction of suspicious regions is performed. Specifically, using a 1×1 sliding window, the inner product of the corresponding position (i, j) is calculated to obtain M1, which represents the spatially consistent region in Fct and Fcs detected at the original scale: ; Among them, i represents the i-th row, j represents the j-th column, represents the similar region features guided by the tampering domain prototype at the i-th row and j-th column; Represents the similar region features guided by the source domain prototype at the i-th row and j-th column; the size of M1 is H*W*1.
[0041] In addition, by calculating the inner product of the corresponding channel at the corresponding position, we can get M k , used to represent the consistent area in the channel dimension detected by Fct and Fcs in the channel dimension; ; Where k represents the k-th dimension channel, represents the similar region feature guided by the tampering domain prototype at the i-th row and j-th column on the k-th dimension channel, M represents the similar region feature guided by the source domain prototype at the i-th row and j-th column on the k-th dimension channel; k The size is H*W*C.
[0042] Splicing M1 and M k And perform the convolution operation to obtain the consistent area M in Fct and Fcs at the original scale, whose size is H*W*(C+1); ; Where conv represents the convolution operation and || represents the concatenation operation.
[0043] The inner product of multi-scale features (macro + channel level) is calculated through a sliding window to locate consistent areas, and then low-consistency areas are extracted through reverse gating to achieve accurate detection of tampered areas.
[0044] As an example, by using a sliding window of size 2×2 with a stride of 1, the consistent regions in Fct and Fcs are obtained. , the formula is as follows: ; ; ; Where v represents the horizontal coordinate of the 2*2 area after the sliding window is divided, and w represents the vertical coordinate of the position; represents the inner product of the phase position at the macroscopic scale, which is (H−1)×(W−1)×1. represents the inner product of the corresponding channel at the corresponding position, which is (H−1)×(W−1)×C. The size is (H − 1) × (W − 1) × (C + 1).
[0045] Will Reshape it to H×W×(C+1) and merge it with M to obtain the final consistent regions in Fct and Fcs at multiple scales; then, use reverse gating to obtain the suspicious regions, denoted as Fs, as shown below: ; Among them, Reshape represents the resampling operation.
[0046] 4. Steps for re-identifying suspicious areas: The suspicious area in Fs is located by extracting the middle and deep layer features (such as the fourth and fifth layer features) of the input image through convolution operations, masking the irrelevant background, and obtaining the suspicious area feature Fac through correlation extraction and feature fusion.
[0047] In this step, the suspicious region re-identification module is used to obtain suspicious regions, eliminate false positive regions, and extract the target region features Fa that are not included in the high-confidence region Fh. Fa complements the high-confidence regions and can generate fine-grained similar region pairs.
[0048] As a preferred embodiment, features F4 and F5, obtained by convolution of the input image through the fourth and fifth layers, are used to re-identify suspicious areas. Specifically, suspicious areas in Fs are mapped to features F4 and F5 to locate the suspicious areas in F4 and F5. Elements in F4 and F5 corresponding to the suspicious locations are retained, while other elements, including background and high-confidence areas, are masked. The masked F4 and F5 are defined as F4m and F5m, respectively.
[0049] Then, correlation extraction is performed on the mask features of F4m and F5m to obtain features F4c and F5c respectively. Finally, F4c and F5c are adaptively fused to generate Fac.
[0050] Step 5: Prototype guide differentiation steps: A detection mask is generated based on the fusion features generated by Fh and Fac, and the intersection-over-union matching of the source domain and the tampered domain prototype obtained by the latest update is used to automatically distinguish the domain to which the target object belongs as the source domain or the tampered domain, and output the final detection result.
[0051] In this step, within the prototype-guided differentiation module, the processing process is as follows: Feature fusion and classification: First, Fh and Fac are adaptively fused to obtain the fused feature Fan. The fused feature Fan is subjected to a softmax operation to implement element classification and obtain the detection mask Ma, which includes the two target objects O1 and O2. Prototype classification: Softmax operations are performed on the source domain prototype and the tampered domain prototype to achieve element classification, and the source domain and tampered domain detection masks Ms and Mt are obtained respectively; Target matching: Calculate the intersection and union ratio of the target parts Os and Ot in Ms and Mt with O1 respectively: If the intersection-over-union ratio of Os and O1 is greater than that of Ot and O1, then the target of O1 in Ma is the source domain and the target of O2 is the tampered domain; If the intersection-over-union ratio of Ot and O1 is greater than that of Os and O1, then the target of O1 in Ma is the tampered domain and the target of O2 is the source domain.
[0052] Finally, the source domain and the tampered domain target are distinguished through intersection-over-union matching.
[0053] To optimize the proposed network model, the total loss function is: ; Among them, Ldet represents the detection loss calculated using cross entropy, Lssim represents the use of SSIM loss to ensure the integrity of the spatial structure of the detected tampered area and the source area, and α is a learnable parameter; ; Among them, pi∈{0,1} indicates whether the i-th pixel belongs to the area to be detected. Indicates the probability that the pixel is predicted to belong to the target area, w1 and w2 are hyperparameters; ; Where x and y are the true value and the detection result respectively, u x and u y represents the mean of x and y, and represents the standard deviation of x and y, represents the covariance of x and y, and is a hyperparameter.
[0054] This network was implemented using the PyTorch deep learning framework. Experiments were conducted using the USC-ISI CMFD dataset, with the dataset split into a 6:2:2 ratio for training, validation, and testing. Training was performed with a batch size of 16 and the SGD optimizer with an initial learning rate of 1e-2. All experiments were conducted on an NVIDIA Tesla P100 GPU server.
[0055] As another embodiment of the present invention, a copy-move tampering detection system based on object-level prototype modeling is provided. Figure 1 As shown, the detection system includes a coarse similar region feature extraction module, an alternating update module, a suspicious region extraction module, a suspicious region re-identification module and a prototype-guided differentiation module.
[0056] The coarse similar region feature extraction module is used to extract the coarse similar region features Fc of the input image; The alternating update module is used to update the prototype and coarse similar region features Fc, and output similar region features Fct and Fcs associated with the tampered region and the source region; The suspicious region extraction module is used to extract high confidence region features Fh and inconsistent region features Fs from Fct and Fcs; The suspicious area re-identification module is used to re-identify the suspicious area extracted by the suspicious area extraction module to obtain the suspicious area feature Fac; The prototype-guided differentiation module generates a detection mask through the fusion features generated by the high-confidence region features Fh and the suspicious region features Fac, and automatically distinguishes the domain to which the target object belongs as the source domain or the tampered domain by using the intersection-over-union matching of the updated source domain and tampered domain prototypes, and outputs the final detection result.
[0057] This system can implement the copy-move tampering detection method based on object-level prototype modeling as described above. The specific processing methods and processes of each module can be found in the previous records and will not be repeated here.
[0058] experiment The technical effects of the present invention are verified by experimental data below.
[0059] 1. Dataset The USC-ISI dataset is used as the training dataset. The USC-ISI dataset contains 100,000 samples. In this embodiment, it is randomly divided into training set, validation set and test set with a ratio of 8:1:1. In addition, three additional datasets are used to evaluate the generalization ability of this method, including CASIA and CoMoFoD. CASIA contains 7491 real samples and 5123 tampered samples, covering copy-move forgery methods and other techniques such as cropping and deletion. In this embodiment, 1313 copy-move samples are selected for analysis. CoMoFoD contains 5000 tampered images, divided into 25 categories, which are post-processed such as rotation, scaling and blurring for robustness testing.
[0060] 2. Evaluation indicators To evaluate the effectiveness of the proposed method, this example uses precision, recall, and F-1 score as indicators to evaluate its performance. The formulas for precision and recall are as follows: Accuracy = TP / (TP+FP) Recall = TP / (TP+FN) Among them, TP, FP and FN represent the true positive, false positive and false negative detection regions, respectively.
[0061] This embodiment uses F1 to comprehensively evaluate the network model of the present invention (abbreviated as PD2Net), which represents the harmonic mean of precision and recall, as shown below: F1=2×(precision·recall) / (precision+recall).
[0062] 3. Comparison Method a) BusterNet: It integrates the features of similar regions and tampered regions to detect and distinguish between the original and tampered regions.
[0063] b) CMSDNet: It uses multi-scale feature extraction techniques to identify similar regions and crop these regions in the image for classification.
[0064] c) Doa-GAN and CNN-T GAN: A GAN structure is adopted to distinguish the source and tampered regions, where the generator consists of a dual-attention based feature extractor.
[0065] d) ARNet: extracts relevant features through multi-scale modeling and then optimizes through encoder-decoder to detect the original and tampered regions without distinction.
[0066] e) DRNet: Detects similar regions by calculating the autocorrelation between the decoupled shallow features and the last layer features.
[0067] f) CAMUNet: uses coordinate attention for multi-scale feature extraction and then fuses them through upsampling.
[0068] 4. Comparative experimental results Table 1 Comparative experimental results
[0069] As shown in Table 1, the proposed method outperforms other methods and provides more accurate results in distinguishing the source area from the tampered area.
[0070] On the USC-ISI dataset, compared with the suboptimal CNN-T GAN, the PD2Net of the present invention improved the F1 value of detecting the source area by 2.06%. In addition, to verify the generalization ability of the PD2Net proposed in the present invention, this embodiment was experimented on the CASIA and CoMoFod datasets, and the results were still the best overall. On the CASIA dataset, compared with the CNN-T GAN, the method proposed in the present invention improved the F1 score of detecting tampered areas by 5.55%. On the CoMoFoD dataset, compared with the CMSDNet, the method proposed in the present invention improved the F1 score of detecting tampered areas by 5.50%.
[0071] The comparative experiment in this embodiment is visualized as follows Figure 2 As shown in the figure, the first to the last columns are the input image, true value, BusterNet, CMSDNet, Doa-GAN, CNN-T GAN and the method proposed in the present invention, respectively. It can be observed that the method proposed in the present invention is able to detect and locate the source object with complete structure and the tampered area object. Specifically, in the second row of the figure, the edges of the masked flowers detected by the comparison network are relatively smooth and lack details. In contrast, the method of the present invention generates more refined detection results, which are closer to the actual situation, and the outlines of the petals can be clearly seen.
[0072] 5. Ablation Experiment Results To verify the effectiveness of each module, ablation experiments are performed.
[0073] a) Baseline: Prototypes are introduced and connected with similar region features.
[0074] b) Fc→Prototype: The prototype is introduced and updated through CCM operation, and the prototype is supervised for loss.
[0075] c) Update prototype → Fc: Perform CCM operations on the updated prototypes respectively and adaptively fuse the obtained features for detection and discrimination.
[0076] d) Mapping-free fusion: Add a suspicious region extraction module and directly fuse the output of the suspicious region extraction module and the high confidence region.
[0077] e) Mapping and Fusion: A suspicious region re-identification module is introduced, and the output of the suspicious region re-identification module is adaptively fused with the high-confidence region features. In the ablation experiment, the distinction between background, source, and tampered regions is achieved through prototype guidance.
[0078] Table 2 Ablation experiment results
[0079] In the ablation experiments, the distinction between background, source, and tampered regions was achieved through prototype guidance. Table 2 shows the ablation experiment results that verify the effectiveness of each module proposed in this paper. Cross-correlation modeling between prototypes and roughly similar regions improves the discrimination performance. Through iterative optimization, the prototypes and roughly similar regions gradually approach the target. Adding the suspicious region extraction module does not achieve significant improvement because it separates highly consistent and inconsistent detection results in two representations guided by different prototypes and then merges them without further processing the suspicious regions, resulting in minimal impact.
[0080] However, by adding the Suspicious Region Re-Identification Module, detection performance is further improved by re-identifying suspicious regions and merging them with high-confidence regions. This is because the present invention identifies suspicious regions and maps them to shallow features, thereby filtering out other regions and re-detecting only the suspicious regions. This avoids the suppression of suspicious regions by high-confidence regions and solves the problem of misidentification of suspicious regions. By using the Suspicious Region Extraction Module, supplementary regions are obtained for refining detection targets.
[0081] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A copy-move tampering detection method based on object-level prototype modeling, characterized in that: The following steps are involved: Steps for correlation detection: Input the copied and moved tampered image, extract its features, and obtain the coarse similarity region feature Fc through correlation detection; Steps for prototype and similar area update: Initialize two sets of feature prototypes Pt and Ps, which are used to represent the tampered domain and the source domain respectively. Alternately update the two prototypes with the similar region feature Fc to generate similar region features Fct and Fcs; Steps for extracting suspicious areas: By analyzing the differences between Fct and Fcs, we can find the inconsistent regions between the two. Specifically, the regions detected as targets in both Fct and Fcs are high-confidence regions, and the features of high-confidence regions are recorded as Fh. In addition to the high-confidence regions, the other regions detected in Fct and Fcs are inconsistent regions, and the features of inconsistent regions are recorded as Fs. Steps for re-identifying suspicious areas: The suspicious area in Fs is located by extracting the mid- and deep-layer features from the input image through convolution operations, masking the irrelevant background, and obtaining the suspicious area feature Fac through correlation extraction and feature fusion; The steps of prototype guidance differentiation: A detection mask is generated based on the fusion features generated by Fh and Fac, and the intersection-over-union matching of the source domain and the tampered domain prototype obtained by the latest update is used to automatically distinguish the domain to which the target object belongs as the source domain or the tampered domain, and output the final detection result.
2. The copy-move tampering detection method based on object-level prototype modeling according to claim 1 is characterized in that: The specific methods for updating prototype and similar area features include: Prototype updates: Using Pt and Ps as query features, and Fc as key and value in the cross-correlation modeling CCM process, through this mechanism, Pt and Ps are updated to generate P (t+1) and P (s+1) , the mathematical formula of the prototype update process is expressed as follows: ; ; Among them, P (t+1) and P (s+1) Represent the updated tampered domain prototype and source domain prototype respectively, conv( ) represents the convolution operation, T represents the transposition operation, "·" represents the dot product operation, represents the square root normalization of the Fc dimension; Updates in similar areas: After completing the prototype update, use P (t+1) and P (s+1) To refine the coarse similarity region feature Fc; Specifically, using P (t+1) and P (s+1) To guide the feature Fc, generate similar region features Fct and Fcs associated with the tampered region and the source region respectively; by adopting this alternating update strategy, the feature representation is dynamically adjusted, and the mathematical formula of the similar region update process is as follows: ; ; in and are the dimension normalization factors of the tampered domain prototype and the source domain prototype, respectively.
3. The copy-move tampering detection method based on object-level prototype modeling according to claim 1 is characterized in that: In the step of extracting suspicious regions, the method for obtaining high confidence regions is to process the features Fct and Fcs using the ReLU activation function to filter out background noise and retain only significant features; After activation, element-wise multiplication is performed between Fct and Fcs to obtain the high confidence region features Fh.
4. The copy-move tampering detection method based on object-level prototype modeling according to claim 1 is characterized in that: In the step of suspicious region extraction, the suspicious region is defined as the location where there is a detection difference between Fct and Fcs, that is, one is detected as background and the other is detected as target; the suspicious region is extracted by calculating the inconsistency between the corresponding locations in two similar regions with Fct and Fcs features. Specifically: Step (1), original scale consistency detection: Calculate the spatial inner product M1 and channel inner product M of Fct and Fcs at the corresponding position through the sliding window k , and perform convolution operation after splicing the two, and output the original scale consistent area M in Fct and Fcs at the original scale; Step (2), multi-scale sliding window detection: By sliding the window between M1 and M k Repeat the consistency detection operation of step (1) above to output the multi-scale consistent region ; Step (3), multi-scale feature merging and suspicious area positioning: Multi-scale consistent regions Merge with the original scale consistent region M to obtain the final consistent region in Fct and Fcs at multiple scales; Subsequently, the inconsistent regions were separated from the consistent regions by reverse gating to obtain the suspicious regions.
5. The copy-move tampering detection method based on object-level prototype modeling according to claim 1 is characterized in that: The steps for re-identifying suspicious areas are as follows: The features F4 and F5 obtained by the fourth and fifth convolution operations of the input image are used to re-identify the suspicious areas. Specifically, the suspicious areas in Fs are mapped to the features F4 and F5 to locate the suspicious areas in F4 and F5. The elements in F4 and F5 corresponding to the suspicious locations are retained, while other elements including the background and high-confidence areas are masked. The masked F4 and F5 are defined as F4m and F5m respectively. Then, correlation extraction is performed on the mask features of F4m and F5m to obtain features F4c and F5c respectively. Finally, F4c and F5c are adaptively fused to generate Fac.
6. The copy-move tampering detection method based on object-level prototype modeling according to claim 1 is characterized in that: The steps of prototype-guided differentiation include: Feature fusion and classification: First, Fh and Fac are adaptively fused to obtain the fused feature Fan. The fused feature Fan is subjected to a softmax operation to implement element classification and obtain the detection mask Ma, which includes the two target objects O1 and O2. Prototype classification: Softmax operations are performed on the source domain prototype and the tampered domain prototype to achieve element classification, and the source domain and tampered domain detection masks Ms and Mt are obtained respectively; Target matching: Calculate the intersection and union ratio of the target parts Os and Ot in Ms and Mt with O1 respectively: If the intersection-over-union ratio of Os and O1 is greater than that of Ot and O1, then the target of O1 in Ma is the source domain and the target of O2 is the tampered domain; If the intersection-over-union ratio of Ot and O1 is greater than that of Os and O1, then the target of O1 in Ma is the tampered domain and the target of O2 is the source domain.
7. The copy-move tampering detection method based on object-level prototype modeling according to claim 1 is characterized in that: The total loss function is: ; Among them, Ldet represents the detection loss calculated using cross entropy, Lssim represents the use of SSIM loss to ensure the integrity of the spatial structure of the detected tampered area and the source area, and α is a learnable parameter; ; Among them, pi∈{0,1} indicates whether the i-th pixel belongs to the area to be detected. Indicates the probability that the pixel is predicted to belong to the target area, w1 and w2 are hyperparameters, W and H are the width and height of the image; ; Where x and y are the true value and the detection result respectively, u x and u y represents the mean of x and y, and represents the standard deviation of x and y, represents the covariance of x and y, and is a hyperparameter.
8. A copy-move tampering detection system based on object-level prototype modeling, characterized in that: Used to implement the copy-move tampering detection method based on object-level prototype modeling as described in any one of claims 1 to 7, the detection system includes a coarse similar area feature extraction module, an alternating update module, a suspicious area extraction module, a suspicious area re-identification module and a prototype-guided differentiation module; The coarse similar region feature extraction module is used to extract the coarse similar region features Fc of the input image; The alternating update module is used to update the prototype and coarse similar region features Fc, and output similar region features Fct and Fcs associated with the tampered region and the source region; The suspicious region extraction module is used to extract high confidence region features Fh and inconsistent region features Fs from Fct and Fcs; The suspicious area re-identification module is used to re-identify the suspicious area extracted by the suspicious area extraction module to obtain the suspicious area feature Fac; The prototype-guided differentiation module generates a detection mask through the fusion features generated by the high-confidence region features Fh and the suspicious region features Fac, and automatically distinguishes the domain to which the target object belongs as the source domain or the tampered domain by using the intersection-over-union matching of the updated source domain and tampered domain prototypes, and outputs the final detection result.
Citation Information
Cited By
Tamper detection method, system, computing device, and readable storage medium based on sufficiency and necessity
CN122676309A