Image forgery positioning detection method based on collaborative difference optimization and multi-modal perception

Through the collaborative difference optimization and multimodal perception of image forgery positioning detection method, using cross-modal information encoder and feature enhancement module, the problems of false alarm and missed detection in image tampering detection are solved, and high-precision tampering area positioning is achieved.

CN120708040AActive Publication Date: 2025-09-26TIANJIN UNIVERSITY OF TECHNOLOGY +4

Patent Information

Application Number
CN202511186682.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-25
Publication Date
2025-09-26
Estimated Expiration
2045-08-25

AI Technical Summary

Technical Problem

Existing image tampering detection methods find it difficult to effectively separate the feature differences between forged and real areas, resulting in false positives and missed detections. Traditional models are unable to fully represent subtle tampering traces, resulting in insufficient detection accuracy and robustness.

Method used

An image forgery localization detection method based on collaborative difference optimization and multimodal perception is adopted. Multi-scale features are extracted through a cross-modal information encoder. The anomaly encoder and cross-scale feature enhancement module are combined to perform image refinement localization detection, and the loss function is used for iterative optimization.

Benefits of technology

It significantly improves the detection sensitivity of tampered areas of different sizes, solves the false detection problem of traditional single-modal methods, realizes double verification of tampering traces, and improves detection accuracy and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708040A_ABST
    Figure CN120708040A_ABST
Patent Text Reader

Abstract

The invention relates to an image forgery positioning detection method based on collaborative difference optimization and multi-modal perception, and belongs to the technical field of computer vision and multimedia evidence obtaining. The method comprises the following steps: obtaining an RGB image and carrying out random data enhancement; extracting a noise map by using a noise extractor; the RGB image and the noise graph generate multi-scale features through a pre-trained cross-modal information encoder; the multi-scale features are input into an exception encoder, the exception encoder comprises a multi-layer perceptron, a linear fusion block and a linear prediction block, multi-layer mapping features are obtained through processing of the multi-layer perceptron, and an initial prediction map is output; inputting the original RGB image into a ViT model to extract global features, and inputting the global features and the multilayer mapping features into a cross-scale feature enhancement module for interactive optimization to obtain a refined positioning map; fusing the initial prediction map and the refined positioning map to obtain a final prediction map; and iterative optimization is carried out through the loss function. According to the invention, the detection sensitivity of tampered areas with different sizes can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision and multimedia forensics, and specifically relates to an image forgery positioning detection method based on collaborative difference optimization and multimodal perception. Background Art

[0002] With the rapid development of digital technology, image tampering has become a common and complex challenge. The tampering process involves non-original modifications to digital images, including cutting and splicing, adding or removing elements, color and brightness adjustment, and image synthesis. These operations can not only be used for legitimate purposes such as artistic creation, but can also be abused maliciously, leading to the spread of false information. In recent years, deep learning technology has achieved remarkable results in image forgery localization, using convolutional neural networks as feature encoders to improve detection capabilities. However, existing methods rely on the design of high-level visual tasks and have difficulty effectively separating the feature differences between forged and authentic areas, resulting in false positives and missed detections. The fundamental reason is that tampering detection must prioritize local semantically agnostic cues (such as noise and texture) rather than semantic content. Traditional models cannot fully represent these subtle traces, limiting detection accuracy and robustness.

[0003] The core challenge of tampering detection lies in learning discriminative features between regions. It is necessary to extract long-range semantic information from a global perspective to identify inconsistencies. For example, analyzing the relationship between the object and the background can enhance sensitivity to tampering, but existing methods mostly rely on single-scale features or single-modal representations and cannot capture the synergistic effects of multiple modalities (such as noise statistics and edge texture). Tampered areas in images often exhibit noise characteristics at the boundaries, and single-modal strategies have difficulty matching these subtle differences, resulting in insufficient detection sensitivity. Therefore, an innovative strategy is urgently needed to combine multimodal perception with collaborative difference optimization to break through traditional limitations and achieve more accurate tampering localization. Summary of the Invention

[0004] In order to solve the above problems, the present invention provides an image forgery positioning detection method based on collaborative difference optimization and multimodal perception.

[0005] To achieve the above-mentioned purpose, the present invention is implemented through the following technical solutions: The present invention provides an image forgery location detection method based on collaborative difference optimization and multimodal perception, comprising the following steps: Obtain the RGB image to be detected and perform random data enhancement on the RGB image; The RGB image is passed through a noise extractor to extract noise information and obtain a noise map. The RGB image and noise map are input into a pre-trained cross-modal information encoder CMX to obtain multi-scale features. The multi-scale features are processed by an abnormal encoder to obtain an initial prediction map; the abnormal encoder includes a multi-layer perceptron MLP, a linear fusion block, and a linear prediction block; the multi-scale features are processed by the multi-layer perceptron MLP to obtain a multi-layer mapping feature; Input the RGB image X into the pre-trained ViT model to extract the global features of the image; input the global features and multi-layer mapping features into the cross-scale feature enhancement module for interactive optimization to obtain the refined localization map; The initial prediction map and the refined positioning map are fused to obtain the final positioning map; The above process is iteratively optimized through the loss function.

[0006] Furthermore, the random data enhancement includes random tampering operations based on semantic perception, basic image transformation, and imaging process simulation.

[0007] Furthermore, the RGB image X is processed by the pre-trained Noiseprint++ noise extractor, Srm noise extractor, and Bayar noise extractor to obtain the first noise auxiliary representation , Second Noise-Assisted Representation , the third noise auxiliary representation ; The first noise auxiliary representation , Second Noise-Assisted Representation and a third noise-assisted representation The mixed features are obtained through the fusion module FM ; Combine RGB image X and mixed features Input into the pre-trained cross-modal information encoder CMX to obtain the first feature , the second feature , the third characteristic And the fourth characteristic The cross-modal information encoder CMX uses a cross-attention mechanism to build an RGB image X and mixed features Channel attention association.

[0008] Furthermore, the linear fusion block includes a convolution block, batch normalization and a ReLU activation function; the linear prediction block includes a convolution block, a Dropout layer and an upsampling layer; The first feature , the second feature , the third characteristic And the fourth characteristic After weighted processing by the multi-layer perceptron MLP, multi-layer mapping features are obtained , ; Multi-layer mapping features After input into the linear fusion block for fusion, the fusion feature is obtained; the fusion feature is passed through the linear prediction block to obtain the initial prediction map .

[0009] Furthermore, the global features and multi-layer mapping features The refined localization map Y2 is input to the cross-scale feature enhancement module. The specific process is as follows: The first mapping feature , the second mapping feature , the third mapping feature , the fourth mapping feature According to their receptive fields, they are divided into local groups ( ) and regional groups ( ); and then respectively with the global features Perform a differential operation and multiply the differential result by the corresponding mapping feature element by element to obtain the enhanced feature , ; The formula is as follows: , in, Represents element-wise multiplication operation; Represents a differential operation; it will enhance the features With global features Fusion is performed to obtain local group fusion features and regional group fusion features ; The formula is as follows: , , in, Indicates splicing along the channel dimension; 、 、 、 Represent the first enhancement feature, the second enhancement feature, the third enhancement feature, and the fourth enhancement feature respectively; the local group fusion feature and regional group fusion features Processing is performed to obtain the refined positioning map Y2. The refined positioning map The mask G is used for training optimization.

[0010] Furthermore, the initial prediction map Multiply the refined positioning map Y2 element by element to get the final prediction map .

[0011] Furthermore, for the final prediction graph For each pixel in the image, calculate its absolute difference with the pixels in the neighborhood to form a texture vector , ;in, Represents the final prediction graph The first pixel in the neighborhood of the i-th pixel Texture vector for pixels; Represents the final prediction graph The first pixel in the neighborhood of the i-th pixel The texture vector of pixels, , , Represents the prediction result of the pixel, represents the middle pixel, Represents the middle pixel The jth pixel in the area, represents the k×k pixels in the neighborhood of the i-th pixel; The first noise auxiliary representation , Second Noise-Assisted Representation and a third noise-assisted representation The texture vectors of the three noise modal images are calculated to obtain the noise image texture vectors , the formula is as follows: , , in, represents the mth noise mode in the neighborhood of the i-th pixel Texture vector for pixels; represents the exponential function; represents a hyperparameter; Represents the texture vector of the i-th pixel under the m-th noise mode; Represents the texture vector of the j-th pixel under the m-th noise mode; Indicates the number of noise modal image types; represents the square of Euclidean distance; The joint consistency loss function is expressed as follows: , in, represents a binary boundary mask, represents the joint consistency loss function.

[0012] Furthermore, the binary boundary mask The acquisition process is as follows: the final prediction graph Convert to binary labels , distinguish the foreground and background, and obtain the binary matrix L; extract the boundary mask from the binary matrix L through the corrosion-expansion technique to obtain the binary boundary mask .

[0013] Furthermore, the loss function in step S6 also includes a weighted cross entropy loss function and a Dice loss function.

[0014] Furthermore, the total loss function is , ,in, 、 、 Represent the weighted cross entropy loss function respectively Weight coefficient, Dice loss function Weight coefficient, joint consistency loss function The weight coefficient of .

[0015] The advantages of the present invention are: This invention utilizes an innovative cross-scale feature enhancement module and visual transformers to capture long-range dependencies, overcoming the limitations of traditional single-scale feature analysis. This significantly improves detection sensitivity for tampered regions of varying sizes, making it particularly suitable for complex splicing and copy-and-move operations. The cross-modal information fusion module creatively integrates edge and noise feature analysis, achieving "double verification" of tampering traces by establishing feature consistency between noise maps. This effectively addresses the false detection issues of traditional single-modal methods (such as those relying solely on noise or texture analysis) and, for the first time, achieves dynamic co-optimization of edge and noise features, creating a new paradigm for "cross-feature collaborative analysis." This provides a scalable technical framework for image forensics, with future compatibility for other modal features, such as spectral features and deep learning fingerprints. Experimental data demonstrates that this solution surpasses existing technologies in key metrics such as F1 scores. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention.

[0017] Figure 1 is a flow chart of the steps of the method of the present invention; Figure 2 This is a visual comparison of edge refinement between the method of the present invention and the baseline model; Figure 3 A visual comparison diagram of false detection and missed detection between the method of the present invention and the existing method; Figure 4 This is a visual comparison diagram of the method of the present invention and the existing method. DETAILED DESCRIPTION

[0018] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments derived by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0019] Example 1 In this embodiment, Figure 1 As shown, the present invention provides an image forgery positioning detection method based on collaborative difference optimization and multimodal perception, the specific steps of which include: Step 1: Obtain the RGB image to be detected and perform random data enhancement on the RGB image.

[0020] Specifically, the random data augmentation includes semantically aware random tampering operations, basic image transformations, and imaging process simulation. The semantically aware random tampering operations include: RandomCopyMove: performing a copy-move operation on random image regions to simulate common splicing tampering; RandomInpainting: using a deep generative model to perform content restoration on randomly selected regions. The basic image transformations include: HorizontalFlip / VerticalFlip: horizontal / vertical flip enhancement; RandomBrightnessContrast: random adjustment of brightness and contrast. The imaging process simulation includes: GaussianBlur: Gaussian blur (σ=0.2); RandomScale: random scaling (scale factor=0.5); RandomCrop: random cropping (crop ratio=1); ImageCompression: JPEG compression (quality factor=0.5).

[0021] Step 2: The RGB image is passed through a noise extractor to extract noise information and obtain a noise map; the RGB image and noise map are input into the pre-trained cross-modal information encoder CMX to obtain multi-scale features.

[0022] Specifically, the RGB image X is processed by the pre-trained Noiseprint++ noise extractor, Srm noise extractor, and Bayar noise extractor to obtain the first noise auxiliary representation , Second Noise-Assisted Representation , the third noise auxiliary representation ; The first noise auxiliary representation , Second Noise-Assisted Representation and a third noise-assisted representation The mixed features are obtained through the fusion module FM ; The fusion module FM includes a first ConvBlock module, a second ConvBlock module, and a third ConvBlock module; the three ConvBlock modules each include a first convolutional layer, a second convolutional layer, and a third convolutional layer, and the first convolutional layer, the second convolutional layer, and the third convolutional layer each include a two-dimensional convolutional layer Conv2d, a two-dimensional batch normalization layer BatchNorm2d, and a ReLU activation function; the number of output channels of the first convolutional layer, the second convolutional layer, and the third convolutional layer are 24, 48, and 96, respectively; the first noise auxiliary representation , Second Noise-Assisted Representation , the third noise auxiliary representation After being processed by the first ConvBlock module, the second ConvBlock module, and the third ConvBlock module, three auxiliary representations are obtained. The three auxiliary representations are input into the 1×1 convolution unit to restore the channel to 3, and the three channel-restored auxiliary representations are obtained; the three channel-restored auxiliary representations are spliced ​​in the channel dimension to obtain the mixed feature. ; The RGB image X and the mixed features Input into the pre-trained cross-modal information encoder CMX to obtain the first feature , the second feature , the third characteristic And the fourth characteristic The cross-modal information encoder CMX uses a cross-attention mechanism to build an RGB image X and mixed features Channel attention association.

[0023] Step 3: Process the multi-scale features through an abnormal encoder to obtain an initial prediction map; the abnormal encoder includes a multi-layer perceptron MLP, a linear fusion block, and a linear prediction block; the multi-scale features are processed by the multi-layer perceptron MLP to obtain multi-layer mapping features.

[0024] Specifically, the linear fusion block includes a 1×1 convolution block, batch normalization, and a ReLU activation function; the linear prediction block includes a 1×1 convolution block, a Dropout layer, and an upsampling layer; The first feature , the second feature , the third characteristic And the fourth characteristic After weighted processing by the multi-layer perceptron MLP, multi-layer mapping features are obtained , , represents the first mapping feature; represents the second mapping feature; represents the third mapping feature; Represents the fourth mapping feature; the formula is as follows: , , , in, Indicates that the convolution kernel size is Convolution operation; Indicates that the convolution kernel size is Convolution operation; Represents batch normalization operation; Indicates the first processing of the i-th feature; represents the second processing of the i-th feature; Represents the Sigmoid activation function; Represents element-wise multiplication operation; represents the i-th feature, ; represents the first mapping feature, represents the second mapping feature, The third mapping feature, represents the fourth mapping feature; Multi-layer mapping features After input into the linear fusion block for fusion, the fusion feature is obtained; the fusion feature is passed through the linear prediction block to obtain the initial prediction map The specific process is: the input is four feature maps of different scales , , , , respectively representing multi-level features from shallow to deep layers, where shallow features contain rich spatial details, while deep features have stronger semantic information. Each feature first undergoes a linear transformation to uniformly map the channel dimension. Then, these processed features are concatenated in order from deep to shallow layers to form a tensor that integrates multi-scale information. Subsequently, the linear fusion block performs channel dimensionality reduction and feature integration to output the fused features. Finally, the linear prediction block predicts the category of each pixel to obtain the initial prediction map The entire process uses hierarchical feature fusion to retain high-resolution details in the shallow layer while integrating deep semantic information, thereby improving segmentation accuracy.

[0025] Step 4: Input the RGB image X into the pre-trained ViT model to extract the global features of the image; input the global features and multi-layer mapping features into the cross-scale feature enhancement module for interactive optimization to obtain the refined localization map.

[0026] Specifically, the global features and multi-layer mapping features The refined localization map Y2 is input to the cross-scale feature enhancement module. The specific process is as follows: The first mapping feature , the second mapping feature , the third mapping feature , the fourth mapping feature According to their receptive fields, they are divided into local groups ( ) and regional groups ( ); and then respectively with the global features Perform a differential operation and multiply the differential result by the corresponding mapping feature element by element to obtain the enhanced feature , ; The formula is as follows: , in, Represents a differential operation; it will enhance the features With global features Fusion is performed to obtain local group fusion features and regional group fusion features ; The formula is as follows: , , in, Indicates splicing along the channel dimension; 、 、 、 Represent the first enhancement feature, the second enhancement feature, the third enhancement feature, and the fourth enhancement feature respectively; the local group fusion feature and regional group fusion features After processing, the refined positioning map Y2 is obtained, and the formula is as follows: , in, Represents aggregation along the channel dimension; the refined positioning map The mask G is used for training optimization, and the joint loss function is as follows: , in 、 They are Dice loss and IOU loss, represents the joint loss function, and G represents the mask.

[0027] Step 5: Initial prediction graph Multiply the refined positioning map Y2 element by element to get the final prediction map .

[0028] Step 6: Iteratively optimize the above process through the loss function; Specifically, for the final prediction graph For each pixel in the image, calculate its absolute difference with the pixels in the neighborhood to form a texture vector , ;in, Represents the final prediction graph The first pixel in the neighborhood of the i-th pixel Texture vector for pixels; Represents the final prediction graph The first pixel in the neighborhood of the i-th pixel The texture vector of pixels, , , Represents the prediction result of the pixel, represents the middle pixel, Represents the middle pixel The jth pixel in the area, represents the k×k pixels in the neighborhood of the i-th pixel; The first noise auxiliary representation , Second Noise-Assisted Representation and a third noise-assisted representation The texture vectors of the three noise modal images are calculated to obtain the noise image texture vectors , the formula is as follows: , , in, represents the mth noise mode in the neighborhood of the i-th pixel Texture vector for pixels; represents the exponential function; represents a hyperparameter; Represents the texture vector of the i-th pixel under the m-th noise mode; Represents the texture vector of the j-th pixel under the m-th noise mode; Indicates the number of noise modal image types; represents the square of Euclidean distance; The joint consistency loss function is expressed as follows: , in, represents a binary boundary mask, represents the joint consistency loss function; the binary boundary mask The acquisition process is: The final prediction graph Convert to binary labels , distinguish the foreground and background, and obtain the binary matrix L, the formula is as follows: , in, Represents the pixel in the i-th row and j-th column; the boundary mask is extracted from the binary matrix L through the erosion-dilation technique to obtain the binary boundary mask .

[0029] Specifically, the loss function also includes the weighted cross entropy loss function and the Dice loss function: The formula of the weighted cross entropy loss function is as follows: , in, represents the weighted cross entropy loss; represents pixel i in mask G; represents pixel i in the RGB image X; 、 Represent the sample weight and positive class weight, which are set to 0.5 and 2.5 respectively to deal with the imbalance of sample data; Represents the logarithmic function, N represents the final prediction graph The total number of pixels in . This loss function optimizes the classification error pixel by pixel, but is sensitive to edge blur.

[0030] The formula of Dice loss function is as follows: , in, represents the Dice loss function; represents the smoothing coefficient.

[0031] Specifically, the total loss function is , ,in, 、 、 Represent the weighted cross entropy loss function respectively Weight coefficient, Dice loss function Weight coefficient, joint consistency loss function The weight coefficient of , 、 、 Set to 0.3, 0.5, and 0.2 respectively.

[0032] Example 2 This example evaluates the performance of the proposed method in the task of locating image tampering. The present invention reports the F1 score using both an optimal threshold and a fixed threshold of 0.5 for comparison. Table 1 shows the comparison data between the proposed method and other methods. Experiments were conducted on the Columbia, COVERAGE, CocoGlide, CasiaV1, and Nist16 datasets. Bold indicates the highest score, while underlined indicates the second highest score.

[0033] Table 1 Comparison of experimental data between the method of the present invention and the existing method Figure 2 The figure shows the difference in edge refinement between the model of the present invention and the TruFor model (baseline model). In the figure, the first row is the tampered image, and the following are the edge map refined by the baseline model, the edge map refined by this method, and the mask. It can be seen that this method can better highlight the details and completeness of the located area in terms of edge refinement than the baseline model. For example, in the first column, the edge refinement of this method is more complete than that of the baseline model. In the second column, the edges located by this method are better in detail than those of the baseline model. In the third column, the model of the present invention does not locate any unnecessary details.

[0034] Figure 3 A visual comparison diagram of the method of the present invention and the TruFor model (previous method) is shown. It can be seen from the figure that the method can effectively solve the problems of false detection and missed detection.

[0035] like Figure 4 The figure shows a visual comparison of the proposed method and existing methods. The first row shows the tampered image, the second row shows the mask, and the white area represents the tampered area. The third row shows the visualization of the proposed method, and the last four rows show the visualizations of other models. As can be seen from the figure, the proposed method locates the tampered area more accurately, clearly, and completely.

[0036] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art will be able to modify the technical solutions described in the aforementioned embodiments or substitute equivalents for some of the technical features. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

Claims

1. A method for detecting image forgery positioning based on collaborative difference optimization and multimodal perception, characterized in that: The following steps are involved: S1. Obtain the RGB image to be detected and perform random data augmentation on the RGB image; S2. Extract noise information from the RGB image through a noise extractor to obtain a noise map. Input the RGB image and noise map into a pre-trained cross-modal information encoder CMX to obtain multi-scale features. S3. The multi-scale features are processed by an abnormal encoder to obtain an initial prediction map; the abnormal encoder includes a multi-layer perceptron MLP, a linear fusion block, and a linear prediction block; the multi-scale features are processed by a multi-layer perceptron MLP to obtain a multi-layer mapping feature; S4. Input the RGB image X into the pre-trained ViT model to extract the global features of the image; The global features and multi-layer mapping features are input into the cross-scale feature enhancement module for interactive optimization to obtain the refined localization map; S5. Fusing the initial prediction map and the refined positioning map to obtain a final positioning map; S6. Iteratively optimize the above process through the loss function.

2. The image forgery location detection method based on collaborative difference optimization and multimodal perception according to claim 1 is characterized in that: The random data enhancement in step S1 includes random tampering operations based on semantic perception, basic image transformation, and imaging process simulation.

3. The image forgery location detection method based on collaborative difference optimization and multimodal perception according to claim 2 is characterized in that: Step S2 specifically includes: S21. The RGB image X is processed by the pre-trained Noiseprint++ noise extractor, Srm noise extractor, and Bayar noise extractor to obtain the first noise auxiliary representation , Second Noise-Assisted Representation , the third noise auxiliary representation ; The first noise auxiliary representation , Second Noise-Assisted Representation and a third noise-assisted representation The mixed features are obtained through the fusion module FM ; S22. Combine RGB image X and mixed features Input into the pre-trained cross-modal information encoder CMX to obtain the first feature , the second feature , the third characteristic And the fourth characteristic The cross-modal information encoder CMX uses a cross-attention mechanism to build an RGB image X and mixed features Channel attention association.

4. The image forgery location detection method based on collaborative difference optimization and multimodal perception according to claim 3 is characterized in that: Step S3 specifically includes: The linear fusion block includes a convolution block, batch normalization and a ReLU activation function; the linear prediction block includes a convolution block, a Dropout layer and an upsampling layer; The first feature , the second feature , the third characteristic And the fourth characteristic After weighted processing by the multi-layer perceptron MLP, multi-layer mapping features are obtained , ; Multi-layer mapping features After input into the linear fusion block for fusion, the fusion feature is obtained; the fusion feature is passed through the linear prediction block to obtain the initial prediction map .

5. The image forgery location detection method based on collaborative difference optimization and multimodal perception according to claim 4 is characterized in that: In step S4, the global features and multi-layer mapping features The refined localization map Y2 is input to the cross-scale feature enhancement module. The specific process is as follows: The first mapping feature , the second mapping feature , the third mapping feature , the fourth mapping feature According to their receptive fields, they are divided into local groups ( ) and regional groups ( ); and then respectively with the global features Perform a differential operation and multiply the differential result by the corresponding mapping feature element by element to obtain the enhanced feature , ; The formula is as follows: , in, Represents element-wise multiplication operation; Represents a differential operation; it will enhance the features With global features Fusion is performed to obtain local group fusion features and regional group fusion features ; The formula is as follows: , , in, Indicates splicing along the channel dimension; 、 、 、 Represent the first enhancement feature, the second enhancement feature, the third enhancement feature, and the fourth enhancement feature respectively; the local group fusion feature and regional group fusion features Processing is performed to obtain the refined positioning map Y2. The refined positioning map The mask G is used for training optimization.

6. The image forgery location detection method based on collaborative difference optimization and multimodal perception according to claim 5 is characterized in that: In step S5, the initial prediction map Multiply the refined positioning map Y2 element by element to get the final prediction map .

7. The image forgery location detection method based on collaborative difference optimization and multimodal perception according to claim 6 is characterized in that: Step S6 includes: For the final prediction graph For each pixel in the image, calculate its absolute difference with the pixels in the neighborhood to form a texture vector , ;in, Represents the final prediction graph The first pixel in the neighborhood of the i-th pixel Texture vector for pixels; Represents the final prediction graph The first pixel in the neighborhood of the i-th pixel The texture vector of pixels, , , Represents the prediction result of the pixel, represents the middle pixel, Represents the middle pixel The jth pixel in the area, represents the k×k pixels in the neighborhood of the i-th pixel; The first noise auxiliary representation , Second Noise-Assisted Representation and a third noise-assisted representation The texture vectors of the three noise modal images are calculated to obtain the noise image texture vectors , the formula is as follows: , , in, represents the mth noise mode in the neighborhood of the i-th pixel Texture vector for pixels; represents the exponential function; represents a hyperparameter; Represents the texture vector of the i-th pixel under the m-th noise mode; Represents the texture vector of the j-th pixel under the m-th noise mode; Indicates the number of noise modal image types; represents the square of Euclidean distance; The joint consistency loss function is expressed as follows: , in, represents a binary boundary mask, represents the joint consistency loss function.

8. The image forgery location detection method based on collaborative difference optimization and multimodal perception according to claim 7 is characterized in that: The loss function in step S6 also includes a weighted cross entropy loss function and a Dice loss function.

9. The image forgery location detection method based on collaborative difference optimization and multimodal perception according to claim 8, characterized in that: The total loss function is , ,in, 、 、 Represent the weighted cross entropy loss function respectively Weight coefficient, Dice loss function Weight coefficient, joint consistency loss function The weight coefficient of .

10. The image forgery location detection method based on collaborative difference optimization and multimodal perception according to claim 9 is characterized in that: The binary boundary mask The acquisition process is as follows: the final prediction graph Convert to binary labels , distinguish the foreground and background, and obtain the binary matrix L; extract the boundary mask from the binary matrix L through the corrosion-expansion technique to obtain the binary boundary mask .

Citation Information

Patent Citations

  • Multi-scale image tampering detection method based on mixed attention mechanism

    CN115578626A

  • Generalized deeply-forged image detection method and system based on noise perception

    CN118196865A

  • Weak supervision image tampering detection positioning method and system based on multiple modes and multiple scales

    CN119027789A

  • Multi-modal image tampering positioning method and system based on edge guidance

    CN120451483A

Cited By

  • Cotton disease and insect pest multi-modal image identification method and device based on multistage feature fusion

    CN121170508A