Multi-view feature supervision image tampering detection method and system based on coupling network

Through the multi-view feature supervision method of the coupled network, combined with the four-stage feature extraction and noise attention module, the local and global modeling problems of image tamper detection in the prior art are solved, and high-precision detection of tampering areas and boundaries are achieved.

CN120564013APending Publication Date: 2025-08-29SOUTH CHINA AGRICULTURAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510579532.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

The existing deep learning-based image tamper detection method is difficult to implement local and global spatial modeling at the same time. The boundary detection module cannot effectively capture gradient information in any direction, and it is insufficient generalization, making it difficult to accurately locate the tamper region and boundary.

Method used

A multi-view feature supervision method based on a coupled network is adopted, combined with a four-stage coupled feature extraction backbone network, boundary detection module and cross-stage noise attention module, and the accuracy of tampering area and boundary detection is improved through local and global features fusion and noise attention enhancement.

Benefits of technology

It enhances the detection accuracy of tampering areas and boundaries, can effectively detect tampering artifact traces, and improves the generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120564013A_ABST
    Figure CN120564013A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-view feature supervision image tampering detection method and system based on a coupling network, and the method comprises the steps: obtaining a real picture mask data set and a tampered picture mask data set, and constructing a tampered picture data set; combining a four-stage coupling feature extraction backbone network module, a boundary detection module and a cross-stage noise attention module to construct an image tampering detection network model based on multi-view feature supervision; and carrying out tampering region positioning detection processing and tampering boundary detection processing on the tampering picture data set based on the multi-view feature supervision image tampering detection network model to obtain a tampering region positioning result and a tampering boundary detection result. The method can fully and effectively detect tampering artifact traces left at the edge of the tampering area by tampering operation, and improves the precision of tampering area positioning detection and boundary detection. The multi-view feature supervision image tampering detection method and system based on the coupling network can be widely applied to the technical field of image recognition processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image recognition and processing technology, and in particular to a method and system for detecting image tampering based on multi-view feature supervision using a coupled network. Background Art

[0002] Existing image editing software technology can easily create realistic tampered images. At the same time, with the rapid development of deep learning, people can use advanced neural networks such as generative adversarial networks (GANs) and diffusion models to generate deep fake images that look "real" to people or generate images that do not exist in the real world. However, this has also led to the use of tampered images to forge certificates, create fake news and rumors, etc. Therefore, it is necessary to detect tampered images and further locate the tampered areas of the tampered images.

[0003] Most existing deep learning-based image tampering detection and localization methods are superior to those based on manual feature extraction. Most of them are based on a single convolutional neural network (CNN) and Transformer architecture, which makes it difficult for them to achieve local and global spatial modeling simultaneously. In addition, the Transformer architecture has quadratic complexity with respect to the input size, which increases memory requirements. Secondly, the boundary detection modules used in existing methods use CNNs architecture or Sobel operator integrated with CNNs to predict the probability of tampering boundaries in input images. Since their convolution kernels are optimized from random initialization or have fixed convolution kernel parameter values, they do not explicitly encode gradient information in any direction, making it difficult for them to focus on tampering edge-related features. Finally, methods that focus on specific forgery clues and specific forgery types are not optimal, and the generalization of the methods needs to be improved to cope with real-world scenarios. Summary of the Invention

[0004] In order to solve the above technical problems, the purpose of the present invention is to provide a multi-view feature-supervised image tampering detection method and system based on a coupled network, which can fully and effectively detect the tampering artifact traces left by the tampering operation at the edge of the tampered area, thereby improving the accuracy of tampering area positioning detection and tampering boundary detection.

[0005] The first technical solution adopted by the present invention is: a multi-view feature-supervised image tampering detection method based on a coupled network, comprising the following steps:

[0006] Obtain a dataset of real image masks and a dataset of tampered image masks, and construct a dataset of tampered images;

[0007] Combining the four-stage coupled feature extraction backbone network module, the boundary detection module and the cross-stage noise attention module, we constructed an image tampering detection network model based on multi-view feature supervision.

[0008] An image tampering detection network model based on multi-view feature supervision performs tampering area positioning detection and tampering boundary detection processing on the tampered image dataset to obtain tampering area positioning results and tampering boundary detection results.

[0009] Furthermore, the image tampering detection network model based on multi-view feature supervision specifically includes a four-stage coupled feature extraction backbone network module, a boundary detection module and a cross-stage noise attention module, wherein the first output end of the four-stage coupled feature extraction backbone network module is connected to the input end of the boundary detection module, and the second output end of the four-stage coupled feature extraction backbone network module is connected to the input end of the cross-stage noise attention module, wherein:

[0010] The four-stage coupling feature extraction backbone network module includes a first-stage coupling feature extraction module, a second-stage coupling feature extraction module, a third-stage coupling feature extraction module and a fourth-stage coupling feature extraction module;

[0011] The first-stage coupled feature extraction module, the second-stage coupled feature extraction module, the third-stage coupled feature extraction module, and the fourth-stage coupled feature extraction module all include a deep residual learning network module, a visual state space module, and a local and global feature fusion module;

[0012] The local and global feature fusion module includes a first local and global feature interactive attention module, a second local and global feature interactive attention module and a channel attention module;

[0013] The boundary detection module includes a central pixel differential convolution module, an axial pixel differential convolution module, a radial pixel differential convolution module and an edge decoder module;

[0014] The cross-stage noise attention module includes a noise extractor, a first matrix multiplication module and a second matrix multiplication module.

[0015] Furthermore, the loss function of the image tampering detection network model based on multi-view feature supervision includes a tampering area positioning loss function, a tampering boundary detection loss function, an image-level detection loss function, and a contrast loss function, and its expression is specifically as follows:

[0016] L=λ1L seg +λ2L edge +λ3L cls +λ4L con

[0017] L seg =L bce +L dice

[0018] Ledge =L bce +L dice

[0019] L cls =L bce ( cls ,lab

[0020]

[0021] In the above formula, L represents the loss function of the image tampering detection network model supervised by multi-view features, L seg represents the tampered area positioning loss function, L edge represents the tampering boundary detection loss function, L cls represents the image-level detection loss function, L con represents the contrast loss function, λ1, λ2, λ3, and λ4 represent weighting factors used to balance different modules, and L bce represents the binary cross entropy loss function, L dice represents the Dice loss function, cls Represents the predicted score, and lab represents the true label. k represents the feature embedding of a pixel in the feature map, P(t k ) indicates that k The set of pixel embeddings with the same label, N(t k ) shows t k The pixel embedding set with different labels, τ represents the temperature hyperparameter, t p Indicates that k Pixel embeddings with the same label, t a Indicates that k Pixel embeddings with different labels.

[0022] Furthermore, the image tampering detection network model based on multi-view feature supervision performs tampering region positioning detection processing and tampering boundary detection processing on the tampered image dataset to obtain tampering region positioning results and tampering boundary detection results. This step specifically includes:

[0023] Input the tampered image dataset into the image tampering detection network model based on multi-view feature supervision;

[0024] A four-stage coupled feature extraction backbone network module of an image tampering detection network model based on multi-view feature supervision performs feature extraction processing on a tampered image dataset to obtain a four-stage coupled feature map, which includes a first-stage feature map, a second-stage feature map, a third-stage feature map, and a fourth-stage feature map;

[0025] The boundary detection module of the image tampering detection network model based on multi-view feature supervision performs tampering boundary detection processing on the first-stage feature map and the second-stage feature map to obtain the tampering boundary detection result;

[0026] The cross-stage noise attention module of the image tampering detection network model based on multi-view feature supervision performs tampering area positioning detection processing on the four-stage coupled feature map and the tampered image dataset to obtain the tampering area positioning result.

[0027] Furthermore, the four-stage coupled feature extraction backbone network module of the image tampering detection network model based on multi-view feature supervision performs feature extraction processing on the tampered image dataset to obtain the four-stage coupled feature map. This step specifically includes:

[0028] The tampered image dataset is input into the four-stage coupled feature extraction backbone network module;

[0029] Based on the first-stage coupled feature extraction module of the four-stage coupled feature extraction backbone network module, global and local feature extraction is performed on the tampered image dataset to obtain the first-stage feature map;

[0030] The second-stage coupled feature extraction module based on the four-stage coupled feature extraction backbone network module performs global and local feature extraction on the first-stage feature map to obtain the second-stage feature map;

[0031] The third-stage coupled feature extraction module based on the four-stage coupled feature extraction backbone network module performs global and local feature extraction processing on the second-stage feature map to obtain the third-stage feature map;

[0032] The fourth-stage coupled feature extraction module based on the four-stage coupled feature extraction backbone network module performs global and local feature extraction processing on the third-stage feature map to obtain the fourth-stage feature map.

[0033] Furthermore, the first-stage coupled feature extraction module based on the four-stage coupled feature extraction backbone network module performs global and local feature extraction processing on the tampered image dataset to obtain the first-stage feature map, which specifically includes:

[0034] Input the tampered image dataset into the first stage coupled feature extraction module;

[0035] Based on the deep residual learning network module coupled with the feature extraction module in the first stage, local feature extraction processing is performed on the tampered image dataset to obtain a local feature map;

[0036] Based on the visual state space module coupled with the feature extraction module in the first stage, global feature extraction processing is performed on the tampered image dataset to obtain a global feature map;

[0037] Based on the local and global feature fusion module of the first-stage coupled feature extraction module, the local feature map and the global feature map are fused to obtain the first-stage feature map.

[0038] Furthermore, the local and global feature fusion module based on the first-stage coupled feature extraction module fuses the local feature map with the global feature map to obtain the first-stage feature map, which specifically includes:

[0039] Input the local feature map and the global feature map into the local and global feature fusion module;

[0040] Based on the first local and global feature interactive attention module and the second local and global feature interactive attention module of the local and global feature fusion module, if the input of the first local and global feature interactive attention module is the local feature map and the input of the second local and global feature interactive attention module is the global feature map, the global feature map is reweighted to the local feature map; if the input of the first local and global feature interactive attention module is the global feature map and the input of the second local and global feature interactive attention module is the local feature map, the local feature map is reweighted to the global feature map, and the local features fused with global information and the global features fused with local information are output;

[0041] Based on the channel attention module of the local and global feature fusion module, the local features fused with global information and the global features fused with local information are spliced ​​and convolved to obtain the first-stage feature map.

[0042] Furthermore, the boundary detection module of the image tampering detection network model based on multi-view feature supervision performs tampering boundary detection processing on the first-stage feature map and the second-stage feature map to obtain the tampering boundary detection result, which specifically includes:

[0043] Input the first stage feature map and the second stage feature map into the boundary detection module;

[0044] Based on the center pixel differential convolution module of the boundary detection module, the difference between the center pixel and the surrounding pixels of the first stage feature map and the second stage feature map is captured to obtain the center pixel differential convolution result;

[0045] The axial pixel differential convolution module based on the boundary detection module captures the difference between the central pixel and the surrounding axial pixels of the first-stage feature map and the second-stage feature map, and obtains the axial pixel differential convolution result;

[0046] The radiation pixel differential convolution module based on the boundary detection module captures the difference between the central pixel and the pixels in the outward diverging direction of the first-stage feature map and the second-stage feature map to obtain the radiation pixel differential convolution result;

[0047] The central pixel differential convolution result, the axial pixel differential convolution result and the radiation pixel differential convolution result are input into the edge decoder module for decoding and addition to obtain the tampering boundary detection result.

[0048] Furthermore, the cross-stage noise attention module of the image tampering detection network model based on multi-view feature supervision performs tampering area positioning detection processing on the four-stage coupled feature map and the tampered image dataset to obtain the tampering area positioning result. This step specifically includes:

[0049] The four-stage coupled feature map and the tampered image dataset are input into the cross-stage noise attention module;

[0050] The noise extractor based on the cross-stage noise attention module extracts the cross-correlation matrix of the noise-sensitive fingerprint from the tampered image dataset and obtains the attention map between the intermediate feature map and the noise-sensitive fingerprint;

[0051] Based on the first matrix multiplication module of the cross-stage noise attention module, the attention map and the first-stage feature map are subjected to noise enhancement processing between the real area and the tampered area to obtain a preliminary pixel embedding feature map;

[0052] Based on the second matrix multiplication module of the cross-stage noise attention module, the preliminary pixel embedding feature map and the four-stage coupling feature map are subjected to noise enhancement processing between the real area and the tampered area to obtain the pixel embedding feature map;

[0053] The pixel embedding feature map is added to the four-stage coupled feature map and a supervised contrast loss is calculated to obtain the tampered area positioning result.

[0054] The second technical solution adopted by the present invention is: a multi-view feature-supervised image tampering detection system based on a coupled network, comprising:

[0055] The first module is used to obtain a real image mask dataset and a tampered image mask dataset to construct a tampered image dataset;

[0056] The second module is used to combine the four-stage coupled feature extraction backbone network module, the boundary detection module and the cross-stage noise attention module to build an image tampering detection network model based on multi-view feature supervision;

[0057] The third module is used to implement an image tampering detection network model based on multi-view feature supervision, which performs tampering area positioning detection and tampering boundary detection processing on the tampered image dataset to obtain tampering area positioning results and tampering boundary detection results.

[0058] The beneficial effects of the method and system of the present invention are as follows: the present invention obtains a real picture mask dataset and a tampered picture mask dataset to construct a tampered picture dataset, further combines a four-stage coupled feature extraction backbone network module, a boundary detection module and a cross-stage noise attention module to construct an image tampering detection network model based on multi-view feature supervision, the four-stage coupled feature extraction backbone network module extracts the image tampering information from the micro level and the macro level respectively, and enhances the model's comprehensive understanding of the image by interactively fusing and complementing the information at the two levels, and the boundary detection module inputs the shallow features into the image tampering detection network model. This module can fully and effectively detect the tampering artifacts left by the tampering operation on the edge of the tampered area, and at the same time guide the model to discover subtle differences between the internal and external areas of the tampered area boundary, thereby improving the accuracy of edge detection. The cross-stage noise attention module avoids the loss of subtle tampering traces in the tampered image as the model network deepens, so that subtle tampering traces can still be perceived in the deep layer of the model. Finally, the image tampering detection network model based on multi-view feature supervision performs tampering area positioning detection and tampering boundary detection on the tampered image dataset, thereby improving the accuracy of tampering area positioning detection and tampering boundary detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 1 is a flowchart of the steps of the multi-view feature-supervised image tampering detection method based on a coupled network of the present invention;

[0060] Figure 2 This is a structural block diagram of the multi-view feature-supervised image tampering detection system based on a coupled network of the present invention;

[0061] Figure 3 Schematic diagram of a tampered image and a real image and their corresponding masks provided by a specific embodiment of the present invention;

[0062] Figure 4 2 is a schematic diagram of the structure of the multi-view feature-supervised image tampering detection network model provided by the specific implementation of the present invention;

[0063] Figure 5 It is a structural diagram of a coupling feature extraction module provided by a specific embodiment of the present invention;

[0064] Figure 6 It is a structural diagram of a local and global feature fusion module provided by a specific embodiment of the present invention;

[0065] Figure 7It is a structural diagram of the local and global feature interactive attention module provided by a specific embodiment of the present invention;

[0066] Figure 8 It is a structural diagram of a boundary detection module provided by a specific embodiment of the present invention;

[0067] Figure 9 Schematic diagram of the structure of the cross-stage noise attention module provided by the specific implementation of the present invention;

[0068] Figure 10 It is a visualization diagram of positioning results of different methods provided by the specific implementation of the present invention. DETAILED DESCRIPTION

[0069] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. The step numbers in the following embodiments are provided for ease of description only and do not limit the order of the steps. The order of execution of the steps in the embodiments can be adaptively adjusted based on the understanding of those skilled in the art.

[0070] First of all, it should be noted that the existing technology proposes a two-stream ResNet network to extract the RGB information and noise information of the image respectively, and adds the Sobel operator and edge residual block to enhance the boundary features of the tampered area, extract the boundary artifacts and noise inconsistency of the tampered area, or proposes a progressive mechanism to generate four tampering masks of different scales, where each mask is used as a priori to help predict the next scale mask. At the same time, a spatial-channel correlation module is designed to focus on the spatial and channel aspects of the extracted features, and a noise fingerprint extraction (Noiseprint++) is proposed to capture low-level forgery traces in the image and fuse the RGB semantic information with the low-level clues of the noise fingerprint. At the same time, the confidence map is introduced in the tampering localization to significantly reduce the false alarm rate.

[0071] However, most of them are based on a single convolutional neural network (CNNs) and Transformer architecture, making it difficult for them to achieve local and global spatial modeling at the same time. In addition, the boundary detection module used in the existing methods uses a CNNs architecture or an architecture module that integrates the Sobel operator and CNNs to predict the probability of tampering boundaries in the input image. Since their convolution kernels are optimized from random initialization or the convolution kernel parameter values ​​are fixed, there is no explicit encoding of gradient information in any direction, making it difficult for them to focus on tampering edge-related features.

[0072] Based on this, the method proposed in the embodiment of the present invention includes a feature extraction backbone network based on the coupling of ResNet and VMamba, a tampering boundary detection module based on a pixel differential convolution operator, and a cross-stage noise attention module. Specifically, ResNet and VMamba are used to extract local image detail information and global context information respectively, and then the two types of information are interactively supplemented, allowing the method to locate the tampered area under the correct understanding of the overall image information. At the same time, compared with the self-attention mechanism in Transformer, SS2D in Vmamba reduces the global information calculation complexity from quadratic to linear. In order to effectively detect tampering artifacts such as subtle pixel changes at the image tampering boundary, the pixel differential convolution operator is used in the tampering boundary detection module to guide the encoder to discover subtle differences between the internal and external areas of the tampered area boundary. In order to further improve the generalization of the method, the present invention uses a cross-stage noise attention module integrated with Noiseprint++ to highlight the inconsistency of information such as noise between the real area and the tampered area, and at the same time adds pixel-based contrast learning to suppress the semantic information of the image.

[0073] Reference Figure 1 The present invention provides a multi-view feature-supervised image tampering detection method based on a coupled network, which includes the following steps:

[0074] S100, obtaining a real image mask dataset and a tampered image mask dataset, and constructing a tampered image dataset;

[0075] Specifically, eight public tampered image datasets (DEFACTO-84k, CASIAv1, Columbia, Coverage, Dso-1, Wild, IMD, NIST16) are used. These datasets include real images and tampered images. The tampered images include three basic tampering types: splicing, copy-move and repair, which are widely used in many methods to generate Mask masks of real images, such as Figure 3 As shown in (b) in the figure, a mask of the tampered image is generated, such as Figure 3 As shown in (a) in .

[0076] S200, combining the four-stage coupled feature extraction backbone network module, the boundary detection module and the cross-stage noise attention module to build an image tampering detection network model based on multi-view feature supervision;

[0077] In this embodiment, if Figure 4As shown, the image tampering detection network model based on multi-view feature supervision specifically includes a four-stage coupled feature extraction backbone network module, a boundary detection module and a cross-stage noise attention module, the first output end of the four-stage coupled feature extraction backbone network module is connected to the input end of the boundary detection module, the second output end of the four-stage coupled feature extraction backbone network module is connected to the input end of the cross-stage noise attention module, wherein the four-stage coupled feature extraction backbone network module includes a first-stage coupled feature extraction module, a second-stage coupled feature extraction module, a third-stage coupled feature extraction module and a fourth-stage coupled feature extraction module; the first-stage coupled feature extraction The extraction module, the second-stage coupled feature extraction module, the third-stage coupled feature extraction module and the fourth-stage coupled feature extraction module all include a deep residual learning network module, a visual state space module and a local and global feature fusion module; the local and global feature fusion module includes a first local and global feature interactive attention module, a second local and global feature interactive attention module and a channel attention module; the boundary detection module includes a center pixel differential convolution module, an axial pixel differential convolution module, a radiation pixel differential convolution module and an edge decoder module; the cross-stage noise attention module includes a noise extractor, a first matrix multiplication module and a second matrix multiplication module.

[0078] Furthermore, it should be noted that the embodiment of the present invention uses binary cross entropy loss L bce , Dice loss L dice To train the tampering area positioning and tampering boundary detection tasks, where L seg =L bce +L dice , L edge =L bce +L dice , where the binary cross entropy loss function and Dice loss function are defined as:

[0079]

[0080]

[0081] Among them G i represents ground-truth, 0 represents original pixel, 1 represents forged pixel, M i is the predicted mask.

[0082] At the same time, binary cross entropy loss is used to train the detection task, L cls =L bce ( cls ,lab).

[0083] Where cls represents the predicted score, lab represents the true label, 0 represents the original image, and 1 represents the forged image. The total loss function of the proposed model includes the tampered area positioning loss L seg , tampering boundary detection loss L edge , image-level detection loss L cls And the contrast loss L con , whose expression is:

[0084] L=λ1L seg +λ2L edge +λ3L cls +λ4L con

[0085] In the above formula, L represents the loss function of the image tampering detection network model supervised by multi-view features, L seg represents the tampered area positioning loss function, L edge represents the tampering boundary detection loss function, L cls represents the image-level detection loss function, L con represents the contrast loss function, λ1, λ2, λ3, and λ4 represent weighting factors for balancing different modules. In the experiments of the specific embodiments of the present invention, the default values ​​of λ1, λ2, λ3, and λ4 are set to 2, 1, 1, and 1, respectively.

[0086] S300, an image tampering detection network model based on multi-view feature supervision performs tampering region positioning detection processing and tampering boundary detection processing on the tampered image dataset to obtain tampering region positioning results and tampering boundary detection results.

[0087] First of all, it should be noted that Figure 4 As shown in , the image tampering detection network model based on multi-view feature supervision is mainly composed of a feature extraction backbone network, a boundary detection module, and a cross-stage noise attention module. Figure 5As shown, each stage of the feature extraction backbone network consists of a ResNet (deep residual learning network module), a VMamba (visual state space module), and a local and global feature fusion module (LGFM). Within each stage of the backbone network, the LGFM fuses and supplements the local and global features extracted by the dual branches of the ResNet and VMamba modules before feeding them into the next stage. The features from the first and second stages (F1 and F2) are fed into the boundary detection module (EDM) to detect tampering boundaries. This is because tampering operations often leave detectable artifacts at the boundaries of tampered regions, and shallow features contain more detail and edge information than deep features. A cross-stage noise attention module (ASNAM) is then introduced to enhance the differences in noise and other information between authentic and tampered regions while suppressing semantic information in the image. This is primarily because most tampering operations occur in areas that exhibit strong correlations with semantic attributes. Therefore, suppressing semantic information allows the correct reflection of information related to tampering traces, improving the generalization of the method.

[0088] S310, inputting the tampered image dataset into the image tampering detection network model based on multi-view feature supervision;

[0089] S320, a four-stage coupled feature extraction backbone network module of an image tampering detection network model based on multi-view feature supervision performs feature extraction processing on the tampered image dataset to obtain a four-stage coupled feature map, wherein the four-stage coupled feature map includes a first-stage feature map, a second-stage feature map, a third-stage feature map, and a fourth-stage feature map;

[0090] Specifically, the tampered image data set is input into the four-stage coupled feature extraction backbone network module; the first-stage coupled feature extraction module based on the four-stage coupled feature extraction backbone network module performs global and local feature extraction processing on the tampered image data set to obtain a first-stage feature map; the second-stage coupled feature extraction module based on the four-stage coupled feature extraction backbone network module performs global and local feature extraction processing on the first-stage feature map to obtain a second-stage feature map; the third-stage coupled feature extraction module based on the four-stage coupled feature extraction backbone network module performs global and local feature extraction processing on the second-stage feature map to obtain a third-stage feature map; the fourth-stage coupled feature extraction module based on the four-stage coupled feature extraction backbone network module performs global and local feature extraction processing on the third-stage feature map to obtain a fourth-stage feature map.

[0091] Among them, for the four-stage coupled feature extraction backbone network module, since the first-stage coupled feature extraction module, the second-stage coupled feature extraction module, the third-stage coupled feature extraction module and the fourth-stage coupled feature extraction module all include a deep residual learning network module, a visual state space module and a local and global feature fusion module, the structure of the first-stage coupled feature extraction module is explained here, and the tampered image dataset is input into the first-stage coupled feature extraction module; based on the deep residual learning network module of the first-stage coupled feature extraction module, the tampered image dataset is subjected to local feature extraction processing to obtain a local feature map; based on the visual state space module of the first-stage coupled feature extraction module, the tampered image dataset is subjected to global feature extraction processing to obtain a global feature map; based on the local and global feature fusion module of the first-stage coupled feature extraction module, the local feature map and the global feature map are fused to obtain the first-stage feature map.

[0092] More specifically, the local feature map and the global feature map are input into the local and global feature fusion module; based on the first local and global feature interactive attention module and the second local and global feature interactive attention module of the local and global feature fusion module, if the input of the first local and global feature interactive attention module is the local feature map, and the input of the second local and global feature interactive attention module is the global feature map, the global feature map is reweighted to the local feature map; if the input of the first local and global feature interactive attention module is the global feature map, and the input of the second local and global feature interactive attention module is the local feature map, the local feature map is reweighted to the global feature map, and the local features fused with global information and the global features fused with local information are output; based on the channel attention module of the local and global feature fusion module, the local features fused with global information and the global features fused with local information are spliced ​​and convolved to obtain the first stage feature map.

[0093] In this embodiment, the local and global feature fusion module (LGFM) proposed in the embodiment of the present invention is specifically as follows: Figure 6 As shown in , it includes a local and global feature interactive attention module (Attention Module) and a channel attention module (SE), where the interactive attention module is specifically as follows Figure 7As shown in the figure, the input features of the local and global feature interactive attention module come from the local features of each ResNet stage and the global features of each VMamba stage. In the interactive attention module, when x comes from a ResNet block and y comes from a VMamba block, the global information of VMamba is used to reweight the local features of ResNet. When X comes from a VMamba block and Y comes from a ResNet block, the local features of ResNet are used to reweight the global information of VMamba. In the local and global feature fusion module, after two processes, global to local and local to global, the ResNet features that fuse global information and the VMamba features that fuse local information are obtained. The channel attention module then learns the weight relationship between channels to enhance information interaction between different channels and adjust the feature response of each channel so that the network can focus more on important feature channels. Finally, the two features are concatenated and convolved to obtain the final output of the module.

[0094] S330, a boundary detection module of the image tampering detection network model based on multi-view feature supervision performs tampering boundary detection processing on the first-stage feature map and the second-stage feature map to obtain a tampering boundary detection result;

[0095] Specifically, the first-stage feature map and the second-stage feature map are input into the boundary detection module; based on the center pixel differential convolution module of the boundary detection module, the difference between the center pixel and the surrounding pixels of the first-stage feature map and the second-stage feature map is captured to obtain the center pixel differential convolution result; based on the axial pixel differential convolution module of the boundary detection module, the difference between the center pixel and the surrounding axial pixels of the first-stage feature map and the second-stage feature map is captured to obtain the axial pixel differential convolution result; based on the radiation pixel differential convolution module of the boundary detection module, the difference between the center pixel and the pixels in the outward diverging direction of the first-stage feature map and the second-stage feature map is captured to obtain the radiation pixel differential convolution result; the center pixel differential convolution result, the axial pixel differential convolution result and the radiation pixel differential convolution result are input into the edge decoder module for decoding and addition to obtain the tampering boundary detection result.

[0096] In this embodiment, the tampering boundary detection module (BDM) proposed in the embodiment of the present invention is constructed based on pixel differential convolution (PDC). It focuses on the differences between adjacent pixel values ​​of the image, which makes it particularly suitable for areas where subtle pixel changes such as tampering artifacts at the image boundaries need to be detected. Traditional convolution operations may ignore subtle changes between adjacent pixels, while pixel differential convolution directly focuses on these differences and guides the encoder to discover subtle differences in the areas inside and outside the tampering area boundary, thereby improving the accuracy of edge detection. Figure 8 As shown, the embodiment of the present invention uses three types of pixel differential convolution, namely central PDC (CPDC), axial PDC (APDC) and radial PDC (RPDC), to capture the difference between the central pixel and the surrounding pixels, the difference between the axial pixels around the central pixel, and the difference between the central pixel and the pixels in the outward diverging direction, and then convolve them with the corresponding convolution weights to obtain F C 、F A 、F R , the three are added together to get the output of the PDC module. The central PDC (CPDC) operation can be expressed as:

[0097]

[0098] where x c represents the central pixel of the local area, x i represents the corresponding surrounding pixels, w i The values ​​represent the learnable convolution weights.

[0099] The axial PDC (APDC) operation can be expressed as:

[0100]

[0101] where x i and x n is a pair of adjacent pixels in the local area, w i The values ​​represent the learnable convolution weights.

[0102] The Radiative PDC (RPDC) operation can be expressed as:

[0103]

[0104] where x r is x i is a pair of adjacent pixels in the direction away from the center of the local area, w i The values ​​represent the learnable convolution weights.

[0105] S340, a cross-stage noise attention module of the image tampering detection network model based on multi-view feature supervision, performs tampering area positioning detection processing on the four-stage coupled feature map and the tampered image dataset to obtain the tampering area positioning result.

[0106] Specifically, the four-stage coupled feature map and the tampered image dataset are input into the cross-stage noise attention module; based on the noise extractor of the cross-stage noise attention module, the cross-correlation matrix of the noise-sensitive fingerprint is extracted from the tampered image dataset to obtain the attention map between the intermediate feature map and the noise-sensitive fingerprint; based on the first matrix multiplication module of the cross-stage noise attention module, the attention map and the first-stage feature map are subjected to noise enhancement processing between the real area and the tampered area to obtain a preliminary pixel embedding feature map; based on the second matrix multiplication module of the cross-stage noise attention module, the preliminary pixel embedding feature map and the four-stage coupled feature map are subjected to noise enhancement processing between the real area and the tampered area to obtain a pixel embedding feature map; the pixel embedding feature map and the four-stage coupled feature map are added and a supervised contrast loss calculation is performed to obtain the tampered area positioning result.

[0107] In this embodiment, the present invention proposes an ASNAM (Across-stage Noise Attention Module), such as Figure 9 As shown. It aims to highlight the inconsistency between the correct area and the tampered area of ​​the picture (such as noise, JPEG compression). The embodiment of the present invention uses Noiseprint++ as a noise extractor, which can not only capture the information of the camera model, but also capture the information of its editing history. This information is an important basis for the method to locate the tampered area, which will be lost as the network deepens, while in the shallow feature layer, this information is better preserved. At the same time, tampering operations generally occur in block areas and rarely occur in single pixels. Therefore, by calculating the cross-correlation matrix between the first-stage feature F1 block and the noise-sensitive fingerprint block extracted by Noiseprint++ of the image, an attention map between the intermediate feature map and the noise-sensitive fingerprint is obtained. This attention map is then used to compare with F all Perform matrix multiplication in blocks to enhance F all The difference in noise and other information between the real area and the tampered area is finally compared with the original input F all The final output of the module is obtained by adding them together. all The real and tampered pixel features are embedded in the dataset and their contrast loss is calculated so that the feature distributions between the two regions can be well separated, further improving the generalization performance.

[0108] Since different tampering operations will leave different forgery footprints in the tampered area of ​​the image, the embodiment of the present invention assumes that regardless of the forgery type, the deep features corresponding to the tampered area and the real area of ​​the image have certain differences. At the same time, knowing the mask corresponding to each image and the corresponding real label embedded in each pixel feature, the embodiment of the present invention adds supervised contrastive learning loss to enable the feature distribution between the two areas to be well separated, such as Figure 10For a single image sample, the final contrast loss L is obtained by calculating the average of all pixel embedding vector contrast loss values con The calculation formula is:

[0109]

[0110] where t p Indicates that k Other pixels with the same label are embedded in the set. Similarly, t a Indicates that k The set of other pixel embeddings with different labels. Where · represents the inner product equivalent of the cosine similarity between two vectors, and τ is a temperature hyperparameter.

[0111] Finally, it should be noted that the tampered region localization in the embodiment of the present invention is a pixel-level binary classification task, that is, determining whether the corresponding pixel is original or tampered. The F1 score (the harmonic mean of precision and recall), MCC score (a balanced binary classification indicator that combines true positives, true negatives, false positives, and false negatives), and IOU score (the degree of overlap between the positioning result and the true label) are used here. These three evaluation indicators are common evaluation criteria in the tampered region localization task, and most existing tampered region localization tasks also use these criteria.

[0112] Specifically, its expression is:

[0113]

[0114] In addition, the image-level detection task in the embodiments of the present invention is also a binary classification task, that is, determining whether an image is original or tampered with. The F1 score and AUC score (an evaluation metric for measuring the quality of a binary classification model) are used here. These two evaluation metrics are common evaluation criteria in classification tasks and are also widely used in existing tampering detection technologies.

[0115] In summary, the differences between the embodiments of the present invention and the prior art are:

[0116] 1) For the parallel extraction of local detail information and global context information of tampered images, compared with the existing technology, the embodiments of the present invention address the problem of low image tampering detection and positioning accuracy from both micro and macro perspectives to improve the model's image tampering detection and positioning capabilities. In reality, tampered images are mostly tampered with based on semantic information, so tampering traces exist at both the micro and macro levels. By extracting local detail information and global context information of tampered images, the two different levels of tampering information are integrated, thereby enhancing the model's image tampering detection and positioning capabilities.

[0117] 2) For the cross-stage noise attention module. Compared with the existing technology, the embodiment of the present invention addresses the problem of poor generalization of image tampering detection and positioning. Starting from highlighting the inconsistency of information such as noise and color between the real area and the tampered area, and suppressing the semantic information of the image, the model's detection and positioning capabilities for unknown tampering operations are improved. An attention map is calculated between the shallow features and the noise information extracted by the noise extractor, and this attention map is used to enhance the inconsistency of information such as noise between the real area and the tampered area, thereby enhancing the generalization of the model.

[0118] Therefore, compared with the prior art, the embodiments of the present invention have the following beneficial effects:

[0119] 1) The embodiment of the present invention designs ResNet and VMamba network structures, and uses the advantages of the characteristics of the two networks to extract image tampering information from the micro level and macro level respectively, and enhances the model's comprehensive understanding of the image by interactively fusing and complementing the information at the two levels. There are many types of image tampering operations, but each type of tampering operation will still leave detectable traces at the micro level. At the same time, most tampering operations are based on tampering of macro objects, which will also leave tampering information at the macro level. Unlike the existing technology, the present invention does not extract tampering microscopic traces at the micro level, but extracts tampering information that combines the micro and macro levels, thereby improving the detection and positioning performance of the model.

[0120] 2) The present invention designs a tampering edge detection module based on pixel differential convolution. Shallow features are input into the module to fully and effectively detect the tampering artifacts left at the edge of the tampered area by the tampering operation. The tampering artifacts at the boundary of the tampered image are almost imperceptible to the human eye. Traditional convolution operations ignore subtle changes between adjacent pixels, while pixel differential convolution directly focuses on these differences and guides the model to discover subtle differences between the internal and external areas of the tampered area boundary, thereby improving the accuracy of edge detection and further enhancing the detection and positioning performance of the model.

[0121] 3) This embodiment of the present invention also incorporates a cross-stage noise attention module to further highlight inconsistencies in noise, color, and other information between the real and manipulated regions of the tampered image. This module can, on the one hand, highlight semantically irrelevant tampering traces, improving the model's detection and positioning performance and generalization. Furthermore, this cross-stage design prevents subtle tampering traces in the manipulated image from being lost as the model network deepens, enabling the perception of subtle tampering traces even at deeper layers of the model.

[0122] We compared our method with other methods in two scenarios: first, pre-training on the DEFACTO-84k dataset and evaluating it on other datasets to verify the generalization of our method compared to other methods. Second, splitting the dataset into training and test sets at a ratio of 9:1 to verify the robustness and image-level detection performance of our method compared to other methods and to verify the effectiveness of our proposed module.

[0123] The experiment of the present invention compares the detection results of some of the most representative methods (Span, MVSS, MVSS++, PSCC, OSN, TruFor, MM) in the field of existing image tampering detection and positioning. Table 1 is the positioning result across data sets, where Table 1 is the pixel-level F1 performance of image forgery positioning calculated using a fixed threshold of 0.5. Table 2 is the image-level detection result. It can be seen that the detection performance of the present invention exceeds the existing technology in most cases. Table 3 is the robustness test result. For common image post-processing interference (Gaussian blur with different kernel sizes, JPEG compression with different quality factors, Gaussian noise, median blur) and social media post-processing operations, the embodiment of the present invention can still effectively locate the tampered image. Table 4 is the ablation result of different modules proposed by the present invention, which shows the contribution of different modules proposed by the present invention to the overall model. It can be seen that the modules proposed by the present invention are effective in detecting and positioning tampered images. Figure 10 This is a visualization of the positioning results of our method and other methods. It is clear from the figure that the mask predicted by our method is closer to the true mask of the tampered image than the other methods. Regardless of the size of the tampered area, our method can basically locate the tampered area.

[0124] Table 1. Positioning results across datasets

[0125] Model Casiav1 Columbia Coverage Dso-1 Wild IMD Nist16 average Span 0.139 0.220 0.170 0.081 0.150 0.112 0.127 0.143 MVSS 0.260 0.129 0.225 0.159 0.185 0.194 0.148 0.186 MVSS++ 0.224 0.252 0.203 0.205 0.228 0.191 0.227 0.219 PSCC 0.166 0.405 0.264 0.205 0.247 0.208 0.199 0.242 OSN 0.127 0.094 0.153 0.156 0.065 0.084 0.101 0.111 TruFor 0.279 0.303 0.244 0.123 0.234 0.217 0.275 0.239 MM 0.372 0.309 0.228 0.071 0.225 0.184 0.290 0.240 The present invention 0.245 0.355 0.284 0.283 0.256 0.219 0.235 0.268

[0126] Table 2 Image-level detection results data table

[0127]

[0128] In Table 2, since the Span and OSN methods do not have image-level detection heads, for fair comparison, the embodiment of the present invention adds a detection head with the same structure as the present invention to their methods, which are represented as Span* and OSN* respectively. The F1 and AUC are both calculated with a fixed threshold of 0.5.

[0129] Table 3 Robustness test results data table

[0130]

[0131]

[0132] Table 3 shows the performance comparison between the proposed method and other methods under various distortions and social media post-processing operations on the Nist16 dataset. The results are expressed as F1 scores with a fixed threshold of 0.5.

[0133] Table 4 Ablation result data of different modules proposed in the present invention

[0134] Variants Columbia Coverage Wild Dso-1 average Baseline 0.9823 0.4370 0.5293 0.6185 0.6418 w / o VMamba 0.9767 0.4196 0.4534 0.5828 0.6081 w / oResNet 0.9645 0.4169 0.3559 0.3945 0.5327 w / oASNAM 0.9739 0.4150 0.5049 0.5282 0.6055 w / o BDM 0.9817 0.3840 0.5166 0.5283 0.6027 w / ocloss 0.9815 0.3833 0.4999 0.5098 0.5936 w / oPDCwSobel 0.9811 0.4122 0.5239 0.5569 0.6185

[0135] In Table 4, the ablation study of removing the module is conducted on four datasets: Columbia, Coverage, Wild, and Dso-1. The F1 results are calculated with a fixed threshold of 0.5.

[0136] In summary, the embodiment of the present invention proposes a multi-view feature learning network based on coupled ResNet and VMamba to detect image tampering and locate tampered areas. The purpose is to solve the problem of low accuracy of existing image tampering detection and positioning methods by extracting local and global information of the image, detecting tampering boundaries, and enhancing the noise between the real area and the tampered area. By combining the local detail information and global context information of the image, a more comprehensive image representation is provided, and a tampering boundary detection module based on pixel differential convolution and a cross-stage noise attention module are designed to improve the detection and positioning capabilities of tampered images.

[0137] Reference Figure 2 , a multi-view feature-supervised image tampering detection system based on coupled networks, including:

[0138] The first module 201 is used to obtain a real image mask dataset and a tampered image mask dataset to construct a tampered image dataset;

[0139] The second module 202 is used to combine the four-stage coupled feature extraction backbone network module, the boundary detection module and the cross-stage noise attention module to build an image tampering detection network model based on multi-view feature supervision;

[0140] The third module 203 is used to perform tampering region positioning detection processing and tampering boundary detection processing on the tampered image dataset based on the image tampering detection network model supervised by multi-view features, and obtain tampering region positioning results and tampering boundary detection results.

[0141] The contents of the above method embodiments are all applicable to the present system embodiments. The functions specifically implemented by the present system embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0142] The above is a specific description of the preferred implementation of the present invention, but the invention is not limited to the embodiments. Those skilled in the art can make various equivalent modifications or substitutions without violating the spirit of the present invention. These equivalent modifications or substitutions are all included in the scope defined by the claims of this application.

Claims

1. A multi-view feature-supervised image tampering detection method based on a coupled network, characterized by: The following steps are involved: Obtain a dataset of real image masks and a dataset of tampered image masks, and construct a dataset of tampered images; Combining the four-stage coupled feature extraction backbone network module, the boundary detection module and the cross-stage noise attention module, we constructed an image tampering detection network model based on multi-view feature supervision. An image tampering detection network model based on multi-view feature supervision performs tampering area positioning detection and tampering boundary detection processing on the tampered image dataset to obtain tampering area positioning results and tampering boundary detection results.

2. The multi-view feature-supervised image tampering detection method based on a coupled network according to claim 1 is characterized in that: The image tampering detection network model based on multi-view feature supervision specifically includes a four-stage coupled feature extraction backbone network module, a boundary detection module and a cross-stage noise attention module, wherein the first output end of the four-stage coupled feature extraction backbone network module is connected to the input end of the boundary detection module, and the second output end of the four-stage coupled feature extraction backbone network module is connected to the input end of the cross-stage noise attention module, wherein: The four-stage coupling feature extraction backbone network module includes a first-stage coupling feature extraction module, a second-stage coupling feature extraction module, a third-stage coupling feature extraction module and a fourth-stage coupling feature extraction module; The first-stage coupled feature extraction module, the second-stage coupled feature extraction module, the third-stage coupled feature extraction module, and the fourth-stage coupled feature extraction module all include a deep residual learning network module, a visual state space module, and a local and global feature fusion module; The local and global feature fusion module includes a first local and global feature interactive attention module, a second local and global feature interactive attention module and a channel attention module; The boundary detection module includes a central pixel differential convolution module, an axial pixel differential convolution module, a radial pixel differential convolution module and an edge decoder module; The cross-stage noise attention module includes a noise extractor, a first matrix multiplication module and a second matrix multiplication module.

3. The multi-view feature-supervised image tampering detection method based on a coupled network according to claim 2 is characterized in that: The loss functions of the image tampering detection network model based on multi-view feature supervision include tampering area positioning loss function, tampering boundary detection loss function, image-level detection loss function and contrast loss function, and their expressions are as follows: L=λ1L seg +λ2L edge +λ3L cls +λ4L con L seg =L bce +L dice L edge =L bce +L dice L cls =L bce (cls,lab) In the above formula, L represents the loss function of the image tampering detection network model supervised by multi-view features, L seg represents the tampered area positioning loss function, L edge represents the tampering boundary detection loss function, L cls represents the image-level detection loss function, L con represents the contrast loss function, λ1, λ2, λ3, and λ4 represent weighting factors used to balance different modules, and L bce represents the binary cross entropy loss function, L dice represents the Dice loss function, cls represents the predicted score, lab represents the true label, t k represents the feature embedding of a pixel in the feature map, P(t k ) indicates that k The set of pixel embeddings with the same label, N(t k ) shows t k The pixel embedding set with different labels, τ represents the temperature hyperparameter, t p Indicates that k Pixel embeddings with the same label, t a Indicates that k Pixel embeddings with different labels.

4. The multi-view feature-supervised image tampering detection method based on a coupled network according to claim 3 is characterized in that: The image tampering detection network model based on multi-view feature supervision performs tampering region positioning detection processing and tampering boundary detection processing on the tampered image dataset to obtain tampering region positioning results and tampering boundary detection results. This step specifically includes: Input the tampered image dataset into the image tampering detection network model based on multi-view feature supervision; A four-stage coupled feature extraction backbone network module of an image tampering detection network model based on multi-view feature supervision performs feature extraction processing on a tampered image dataset to obtain a four-stage coupled feature map, which includes a first-stage feature map, a second-stage feature map, a third-stage feature map, and a fourth-stage feature map; The boundary detection module of the image tampering detection network model based on multi-view feature supervision performs tampering boundary detection processing on the first-stage feature map and the second-stage feature map to obtain the tampering boundary detection result; The cross-stage noise attention module of the image tampering detection network model based on multi-view feature supervision performs tampering area positioning detection processing on the four-stage coupled feature map and the tampered image dataset to obtain the tampering area positioning result.

5. The method for detecting image tampering based on multi-view feature supervision of coupled networks according to claim 4 is characterized in that: The four-stage coupled feature extraction backbone network module of the multi-view feature-supervised image tampering detection network model performs feature extraction processing on the tampered image dataset to obtain the four-stage coupled feature map. This step specifically includes: The tampered image dataset is input into the four-stage coupled feature extraction backbone network module; Based on the first-stage coupled feature extraction module of the four-stage coupled feature extraction backbone network module, global and local feature extraction is performed on the tampered image dataset to obtain the first-stage feature map; The second-stage coupled feature extraction module based on the four-stage coupled feature extraction backbone network module performs global and local feature extraction on the first-stage feature map to obtain the second-stage feature map; The third-stage coupled feature extraction module based on the four-stage coupled feature extraction backbone network module performs global and local feature extraction processing on the second-stage feature map to obtain the third-stage feature map; The fourth-stage coupled feature extraction module based on the four-stage coupled feature extraction backbone network module performs global and local feature extraction processing on the third-stage feature map to obtain the fourth-stage feature map.

6. The method for detecting image tampering based on multi-view feature supervision of coupled networks according to claim 5, characterized in that: The first-stage coupled feature extraction module based on the four-stage coupled feature extraction backbone network module performs global and local feature extraction processing on the tampered image dataset to obtain the first-stage feature map. This step specifically includes: Input the tampered image dataset into the first stage coupled feature extraction module; Based on the deep residual learning network module coupled with the feature extraction module in the first stage, local feature extraction processing is performed on the tampered image dataset to obtain a local feature map; Based on the visual state space module coupled with the feature extraction module in the first stage, global feature extraction processing is performed on the tampered image dataset to obtain a global feature map; Based on the local and global feature fusion module of the first-stage coupled feature extraction module, the local feature map and the global feature map are fused to obtain the first-stage feature map.

7. The method for detecting image tampering based on multi-view feature supervision of coupled networks according to claim 6, characterized in that: The step of fusing the local feature map and the global feature map based on the local and global feature fusion module of the first-stage coupled feature extraction module to obtain the first-stage feature map specifically includes: Input the local feature map and the global feature map into the local and global feature fusion module; Based on the first local and global feature interactive attention module and the second local and global feature interactive attention module of the local and global feature fusion module, if the input of the first local and global feature interactive attention module is the local feature map and the input of the second local and global feature interactive attention module is the global feature map, the global feature map is reweighted to the local feature map; if the input of the first local and global feature interactive attention module is the global feature map and the input of the second local and global feature interactive attention module is the local feature map, the local feature map is reweighted to the global feature map, and the local features fused with global information and the global features fused with local information are output; Based on the channel attention module of the local and global feature fusion module, the local features fused with global information and the global features fused with local information are spliced ​​and convolved to obtain the first-stage feature map.

8. The method for detecting image tampering based on multi-view feature supervision of coupled networks according to claim 7, characterized in that: The boundary detection module of the image tampering detection network model based on multi-view feature supervision performs tampering boundary detection processing on the first-stage feature map and the second-stage feature map to obtain the tampering boundary detection result. This step specifically includes: Input the first stage feature map and the second stage feature map into the boundary detection module; Based on the center pixel differential convolution module of the boundary detection module, the difference between the center pixel and the surrounding pixels of the first stage feature map and the second stage feature map is captured to obtain the center pixel differential convolution result; The axial pixel differential convolution module based on the boundary detection module captures the difference between the central pixel and the surrounding axial pixels of the first-stage feature map and the second-stage feature map, and obtains the axial pixel differential convolution result; The radiation pixel differential convolution module based on the boundary detection module captures the difference between the central pixel and the pixels in the outward diverging direction of the first-stage feature map and the second-stage feature map to obtain the radiation pixel differential convolution result; The central pixel differential convolution result, the axial pixel differential convolution result and the radiation pixel differential convolution result are input into the edge decoder module for decoding and addition to obtain the tampering boundary detection result.

9. The method for detecting image tampering based on multi-view feature supervision of coupled networks according to claim 8, characterized in that: The cross-stage noise attention module of the image tampering detection network model based on multi-view feature supervision performs tampering area positioning detection processing on the four-stage coupled feature map and the tampered image dataset to obtain the tampering area positioning result. This step specifically includes: The four-stage coupled feature map and the tampered image dataset are input into the cross-stage noise attention module; The noise extractor based on the cross-stage noise attention module extracts the cross-correlation matrix of the noise-sensitive fingerprint from the tampered image dataset and obtains the attention map between the intermediate feature map and the noise-sensitive fingerprint; Based on the first matrix multiplication module of the cross-stage noise attention module, the attention map and the first-stage feature map are subjected to noise enhancement processing between the real area and the tampered area to obtain a preliminary pixel embedding feature map; Based on the second matrix multiplication module of the cross-stage noise attention module, the preliminary pixel embedding feature map and the four-stage coupling feature map are subjected to noise enhancement processing between the real area and the tampered area to obtain the pixel embedding feature map; The pixel embedding feature map is added to the four-stage coupled feature map and a supervised contrast loss is calculated to obtain the tampered area positioning result.

10. A multi-view feature-supervised image tampering detection system based on coupled networks, characterized by: Includes the following modules: The first module is used to obtain a real image mask dataset and a tampered image mask dataset to construct a tampered image dataset; The second module is used to combine the four-stage coupled feature extraction backbone network module, the boundary detection module and the cross-stage noise attention module to build an image tampering detection network model based on multi-view feature supervision; The third module is used to implement an image tampering detection network model based on multi-view feature supervision, which performs tampering area positioning detection and tampering boundary detection processing on the tampered image dataset to obtain tampering area positioning results and tampering boundary detection results.