Image detection method and device, electronic equipment and storage medium
By extracting multi-scale frequency domain and spatial features, combined with heterogeneous high-resolution backbone networks and dual attention fusion mechanisms, an image tampering detection report is generated, which solves the problem of insufficient robustness in existing technologies and achieves higher detection accuracy and interpretability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING UNIV OF POSTS & TELECOMM
- Filing Date
- 2026-01-07
- Publication Date
- 2026-05-15
AI Technical Summary
Existing technologies lack robustness in image tampering detection, cannot provide evidence of tampering type or interpretability, and are insufficient to meet the needs of high-reliability scenarios such as judicial evidence collection.
By acquiring the image to be detected, multi-scale frequency domain and spatial features are extracted. Multi-scale frequency domain and noise feature maps are generated using a heterogeneous high-resolution backbone network. Combined with frequency-space gating mechanism and dual attention fusion mechanism, enhanced spatial feature map and noise feature map are generated. Finally, a tampering probability map is generated through a consistency-aware feature pyramid network to provide an image tampering detection report.
It improves the robustness and interpretability of image tampering detection, and the generated detection report can provide users with intuitive evidence support, making it suitable for fields with high requirements for the interpretability of results.
Smart Images

Figure CN122048815A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of data processing technology, and in particular to an image detection method, apparatus, electronic device and storage medium. Background Technology
[0002] With the increasing popularity and advancement of image editing, image tampering has become easier and more frequent. Deep learning algorithms can directly segment the tampered regions of the original image. However, directly identifying tampered regions using deep learning algorithms suffers from insufficient robustness and cannot provide evidence of the type of tampering or its interpretability.
[0003] Therefore, improving the robustness of image tampering detection has become an urgent technical problem to be solved. Summary of the Invention
[0004] In view of this, the purpose of this disclosure is to provide an image detection method, apparatus, electronic device and storage medium to solve or partially solve the above-mentioned technical problems.
[0005] To achieve the above objectives, a first aspect of this disclosure proposes an image detection method, the method comprising: The image to be detected is acquired, and the image to be detected is transformed to obtain a multi-scale frequency domain feature map. The image to be detected is then extracted to obtain an initial noise feature map. The image to be detected is processed to obtain a multi-scale spatial feature map, and the initial noise feature map is processed to obtain a multi-scale noise feature map. The multi-scale frequency domain feature map is added to the multi-scale spatial feature map to obtain an enhanced spatial feature map. The enhanced spatial feature map and the multi-scale noise feature map are fused to obtain an initial fused feature map. The initial fused feature map is then aggregated to obtain a target fused feature map. Based on the target fusion feature map, a tampering probability map is determined, and the tampering probability map is used as the image tampering detection result. An image detection report is generated based on the multi-scale frequency domain feature map, the initial noise feature map, the initial fusion feature map, and the tampering probability map.
[0006] Based on the same inventive concept, a second aspect of this disclosure proposes an image detection apparatus, comprising: The preprocessing module is configured to acquire the image to be detected, perform transformation processing on the image to be detected to obtain a multi-scale frequency domain feature map, and extract the image to be detected to obtain an initial noise feature map; The resolution processing module is configured to perform resolution processing on the image to be detected to obtain a multi-scale spatial feature map, and to perform resolution processing on the initial noise feature map to obtain a multi-scale noise feature map. The fusion processing module is configured to add the multi-scale frequency domain feature map to the multi-scale spatial feature map to obtain an enhanced spatial feature map, perform fusion processing on the enhanced spatial feature map and the multi-scale noise feature map to obtain an initial fusion feature map, and perform aggregation processing on the initial fusion feature map to obtain a target fusion feature map; The image detection module is configured to determine a tampering probability map based on the target fusion feature map, use the tampering probability map as the image tampering detection result, and generate an image detection report based on the multi-scale frequency domain feature map, the initial noise feature map, the initial fusion feature map, and the tampering probability map.
[0007] Based on the same inventive concept, a third aspect of this disclosure proposes an electronic device including a memory, a processor, and a computer program stored in the memory and executable by the processor, wherein the processor implements the method described above when executing the computer program.
[0008] Based on the same inventive concept, a fourth aspect of this disclosure provides a non-transitory computer-readable storage medium that stores computer instructions for causing a computer to perform the methods described above.
[0009] As can be seen from the above description, the image detection method, apparatus, electronic device, and storage medium provided in this disclosure include: acquiring an image to be detected; performing transformation processing on the image to be detected to obtain a multi-scale frequency domain feature map, which can capture the frequency domain features in the image to be detected; extracting and processing the image to be detected to obtain an initial noise feature map, which can capture the noise features in the image to be detected; performing resolution processing on the image to be detected to obtain a multi-scale spatial feature map, which can reflect the spatial variation features at different resolution levels; and performing resolution processing on the initial noise feature map to obtain a multi-scale noise feature map, which can reflect the noise variation features at different resolution levels. By performing resolution processing on the image, the problem of missed detection or false detection caused by the limitations of single-scale features can be effectively avoided. Adding the multi-scale frequency domain feature map to the multi-scale spatial feature map to obtain an enhanced spatial feature map, which integrates information from the frequency domain and spatial domain, and can enhance the expressive ability of the features to the image content. An initial fused feature map is obtained by fusing the enhanced spatial feature map and the multi-scale noise feature map. This initial fused feature map integrates multiple types of feature information, more accurately reflecting the true state of the image to be detected, reducing the influence of external factors on the image tampering detection results, and improving robustness. The initial fused feature map is then aggregated to obtain the target fused feature map. A tampering probability map is determined based on the target fused feature map and used as the image tampering detection result. An image detection report is generated based on the multi-scale frequency domain feature map, the initial noise feature map, the initial fused feature map, and the tampering probability map. This image detection report provides users with intuitive evidence, making the image tampering detection results more convincing and facilitating its application in fields with high requirements for interpretability. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in this disclosure or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 This is a flowchart of an image detection method according to an embodiment of the present disclosure; Figure 2 This is a schematic diagram of the structure of the image tampering detection network according to an embodiment of the present disclosure; Figure 3 This is a schematic diagram of the dual attention mechanism according to an embodiment of the present disclosure; Figure 4 This is a schematic diagram of the structure of a multi-agent interpretability generation system according to an embodiment of the present disclosure; Figure 5This is a schematic diagram of the structure of the image detection device according to an embodiment of the present disclosure; Figure 6 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present disclosure. Detailed Implementation
[0012] To make the objectives, technical solutions, and advantages of this disclosure clearer, the following detailed description is provided in conjunction with specific embodiments and the accompanying drawings.
[0013] It should be noted that, unless otherwise defined, the technical or scientific terms used in the embodiments of this disclosure should have the ordinary meaning understood by one of ordinary skill in the art to which this disclosure pertains. The terms "first," "second," and similar terms used in the embodiments of this disclosure do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0014] Based on the description of the background technology, current image tampering detection methods are mainly divided into two categories: one is based on handcrafted features, such as using Discrete Cosine Transform (DCT) coefficients to detect frequency domain discontinuities, or extracting noise residuals through constrained convolution to identify sensor noise anomalies; the other is based on deep learning, which directly uses convolutional neural networks in the RGB space to perform end-to-end tampering region segmentation.
[0015] However, the former is sensitive to complex post-processing (e.g., blurring, recompression) and has weak generalization ability; the latter, while performing well on specific datasets, relies excessively on high-level semantic features and ignores underlying physical traces, resulting in insufficient cross-domain robustness. Some methods attempt to fuse multimodal cues (e.g., spatial + noise), but most employ simple splicing or shallow fusion strategies, making it difficult to effectively align heterogeneous features and easily causing information loss or semantic conflicts. Furthermore, related technologies generally lack interpretability, only outputting tampering masks and failing to provide tampering type judgments or supporting evidence, making it difficult to meet the needs of high-credibility scenarios such as judicial evidence collection. With the popularization and development of generative artificial intelligence, image tampering (e.g., copy-paste forgery, region splicing, content erasure, etc.) is becoming increasingly covert and difficult to detect, posing a serious challenge to traditional manual review.
[0016] As mentioned above, how to improve the robustness of image tampering detection has become an important research question.
[0017] Based on the above description, such as Figure 1 As shown, the image detection method proposed in this embodiment includes: Step 101: Obtain the image to be detected, perform transformation processing on the image to be detected to obtain a multi-scale frequency domain feature map, and perform extraction processing on the image to be detected to obtain an initial noise feature map.
[0018] In practice, the image to be detected is acquired. Multimodal preprocessing is performed on the image to be detected to obtain multi-scale frequency domain feature maps. and initial noise feature map .
[0019] Step 102: Perform resolution processing on the image to be detected to obtain a multi-scale spatial feature map, and perform resolution processing on the initial noise feature map to obtain a multi-scale noise feature map.
[0020] In practice, the image to be detected Inputting into the first neural network extracts multi-scale spatial feature maps The initial noise feature map Input the second neural network to extract multi-scale noise feature maps The first neural network and the second neural network are two heterogeneous high-resolution backbone networks.
[0021] Step 103: Add the multi-scale frequency domain feature map to the multi-scale spatial feature map to obtain an enhanced spatial feature map; perform fusion processing on the enhanced spatial feature map and the multi-scale noise feature map to obtain an initial fused feature map; and perform aggregation processing on the initial fused feature map to obtain a target fused feature map.
[0022] In practice, a frequency-space gating mechanism is used to transform multi-scale frequency domain feature maps. Injected step-by-step into spatial feature maps of the corresponding scale In the process, an enhanced spatial feature map is generated. , here No additional processing is required.
[0023] By employing a dual-attention fusion mechanism, the spatial feature map is enhanced. With multi-scale noise feature map The initial fused feature map is obtained by performing cross-modal adaptive fusion processing. .
[0024] The initial fused feature map Input a Consistency-Aware Feature Pyramid Network (CAFPN) to generate a single high-resolution target fusion feature map. .
[0025] Step 104: Determine the tampering probability map based on the target fusion feature map, use the tampering probability map as the image tampering detection result, and generate an image detection report based on the multi-scale frequency domain feature map, the initial noise feature map, the initial fusion feature map, and the tampering probability map.
[0026] In practice, a mask prediction head is used, based on the target fusion feature map. Generate a tampering probability map with the same resolution as the image to be detected, and use the tampering probability map as the image tampering detection result.
[0027] Based on multiple intelligent agents, an image detection report is generated according to the multi-scale frequency domain feature map, the initial noise feature map, the initial fused feature map, and the tampering probability map.
[0028] Figure 2 This is a schematic diagram of the structure of an image tampering detection network according to an embodiment of this disclosure. Figure 2 As shown, in S101, a multi-scale frequency domain feature map (image after DCT) and an initial noise feature map are determined based on the image to be detected. In S102, the image to be detected is input into a first neural network (HRNet-W32 architecture) to generate a multi-scale spatial feature map, and the initial noise feature map is input into a second neural network (HRNet-W18 architecture) to generate a multi-scale noise feature map. In S103, the multi-scale frequency domain feature map is input into a lightweight convolution to generate a mapped frequency domain feature map. The mapped frequency domain feature map is used as a query parameter, and the multi-scale spatial feature map is used as a key parameter and a value parameter. A spatial attention map is determined based on the query parameter and the key parameter. In S103, the multi-scale frequency domain feature map is added to the multi-scale spatial feature map through a frequency domain enhancement submodule to obtain an enhanced spatial feature map. In S103, the enhanced spatial feature map and the multi-scale noise feature map are fused using a dual attention mechanism to obtain an initial fused feature map. In S103, the initial fused feature map is aggregated using a consistency-aware feature pyramid network to obtain a target fused feature map. In S104, the upsampling prediction head is used to determine the tampering probability map based on the target fusion feature map, and the prediction result is output as sample label.
[0029] Through the above embodiments, an image to be detected is acquired, and a multi-scale frequency domain feature map is obtained by transforming the image to be detected, which can capture the frequency domain features in the image to be detected. An initial noise feature map is obtained by extracting noise features from the image to be detected, which can capture the noise features in the image to be detected. A multi-scale spatial feature map is obtained by resolving the resolution of the image to be detected. The multi-scale spatial feature map can reflect the spatial variation features at different resolution levels. Similarly, a multi-scale noise feature map is obtained by resolving the initial noise feature map. This multi-scale noise feature map can reflect the noise variation features at different resolution levels. Resolving the image resolution can effectively avoid the problem of missed or false detections caused by the limitations of single-scale features. Adding the multi-scale frequency domain feature map to the multi-scale spatial feature map yields an enhanced spatial feature map. The enhanced spatial feature map integrates information from both the frequency and spatial domains, which can enhance the expressive power of the features on the image content. Finally, the enhanced spatial feature map and the multi-scale noise feature map are fused to obtain an initial fused feature map. This initial fused feature map integrates multiple types of feature information, which can more accurately reflect the true state of the image to be detected, reduce the influence of external factors on the image tampering detection results, and improve robustness. The initial fused feature map is aggregated to obtain the target fused feature map. A tampering probability map is determined based on the target fused feature map, and this map is used as the image tampering detection result. An image detection report is generated based on the multi-scale frequency domain feature map, the initial noise feature map, the initial fused feature map, and the tampering probability map. This image detection report provides users with intuitive evidence, making the image tampering detection results more convincing and facilitating its application in fields where interpretability is crucial.
[0030] In some embodiments, step 101 includes: Step 1011: Convert the image to be detected into a brightness image, and use a discrete cosine transform encoder to divide the brightness image into multiple pixel blocks.
[0031] In practice, assume the image to be detected is ,in, Indicates the height of the image to be detected. This indicates the width of the image to be detected.
[0032] Convert the image to be detected into a brightness image. Specifically, the red, green, and blue channels are extracted from the image to be detected, and a brightness image is obtained by weighting the red, green, and blue channels. ,in, For brightness image, It is a red channel. For green channel, The blue channel.
[0033] Brightness image The input is a multi-channel discrete cosine transform encoder (DCT encoder), which divides the brightness image into multiple pixel blocks. These pixel blocks can be non-overlapping. Pixel block.
[0034] Step 1012: Perform a two-dimensional discrete cosine transform on each pixel block to obtain frequency domain coefficients, and map the frequency domain coefficients to the feature space to obtain pixel block features.
[0035] In practical implementation, when multiple pixel blocks are non-overlapping... When dealing with pixel blocks, each pixel block can be represented as ,in, Rows representing pixel blocks , Rows representing pixel blocks .
[0036] The frequency domain coefficients are obtained by performing a two-dimensional discrete cosine transform on each pixel block. , in, These are frequency domain coefficients, which represent the brightness image. Central OK Non-overlapping columns Specific frequency components within a pixel block The frequency domain coefficients constitute part of the transformed feature map. Indicates in Spatial coordinates within a pixel block, used to identify the position of a pixel, with values ranging from 0 to 7. and For frequency index, . and Scaling factor , .
[0037] To reduce computational complexity and focus on the sensitive areas to be tampered with, only the top left corner is retained. The low-frequency coefficients (16 in total) are calculated, and the low-frequency coefficient in the upper left corner is flattened into a 16-dimensional vector. This vector is then mapped to a 64-dimensional feature space through a learnable linear projection layer. Specifically, mapping the frequency domain coefficients to the feature space yields pixel block features. ,in, , For trainable parameters, This indicates matrix vectorization operations.
[0038] Step 1013: Reassemble the pixel block features of all pixel blocks to obtain an initial frequency domain feature map, and perform pooling processing on the initial frequency domain feature map to obtain a multi-scale frequency domain feature map.
[0039] In practice, the pixel block features of all pixels are recombined to obtain the initial frequency domain feature map. .
[0040] The initial frequency domain feature map is pooled to obtain a multi-scale frequency domain feature map. Specifically, three average pooling operations (kernel size 2, stride 2) are applied sequentially to obtain frequency domain feature maps at three scales. The first-scale frequency domain feature map... The resolution is Second-scale frequency domain feature map The resolution is Third-scale frequency domain feature map The resolution is .
[0041] To align with the output scale of the subsequent backbone network, bilinear interpolation is used for upsampling. , , Finally, a set of multi-scale frequency domain feature maps is obtained. .
[0042] Step 1014: Input the image to be detected into a differentiable convolutional layer, and use the differentiable convolutional layer to perform calculations on the image to be detected to obtain an initial noise feature map.
[0043] In practice, the image to be detected is input into a differentiable Bayar convolutional layer that satisfies zero-sum constraints. Let the convolutional layer be... The weights must satisfy the requirements for each channel. have .
[0044] To achieve zero-sum constraints while preserving gradient differentiability, the following parameterization form is actually used. ,in, These are the original learnable parameters. This represents the convolution kernel after specific constraints and transformations. It is a matrix of all ones. This represents broadcast subtraction along the channel dimension.
[0045] Force the center point during initialization. The remaining positions are initialized with Gaussian random numbers with a mean of 0 and a standard deviation of 0.01. This setting makes the convolution operation approximate the current pixel minus the mean of its local neighborhood, thus highlighting local anomalies.
[0046] Then, using the adjusted convolution kernel Perform a convolution operation on the image to be detected, and output the initial noise feature map. .
[0047] The above scheme converts the image to be detected into a luminance image, and then uses a discrete cosine transform encoder to divide the luminance image into multiple pixel blocks. A two-dimensional discrete cosine transform is performed on each pixel block to obtain frequency domain coefficients, which are then mapped to the feature space to obtain pixel block features. The pixel block features of all pixel blocks are recombined to obtain an initial frequency domain feature map, and pooling is applied to this initial frequency domain feature map to obtain a multi-scale frequency domain feature map. This multi-scale frequency domain feature map can capture the frequency domain features in the image to be detected and reflects the frequency domain variation features at different resolution levels. The image to be detected is then input into a differentiable convolutional layer, which performs operations on the image to obtain an initial noise feature map. This initial noise feature map can capture the noise features in the image to be detected.
[0048] In some embodiments, step 102 includes: Step 1021: Input the image to be detected into the first neural network, generate a preset number of spatial feature maps using a preset number of high-resolution channels in the first neural network, and use the preset number of spatial feature maps as multi-scale spatial feature maps.
[0049] Step 1022: Input the initial noise feature map into the second neural network, generate a preset number of noise feature maps using a preset number of high-resolution channels in the second neural network, and use the preset number of noise feature maps as multi-scale noise feature maps.
[0050] In practical implementation, the first and second neural networks can be high-resolution backbone networks (HRNet architecture). The HRNet architecture employs a parallel multi-resolution representation architecture, simultaneously maintaining multiple feature branches of different resolutions throughout the entire forward propagation process, and achieving high-precision feature fusion through periodic cross-resolution information exchange. The above structure unfolds sequentially from left to right, containing several stages, each consisting of multiple parallel convolutional units, corresponding to different spatial resolution paths.
[0051] The input image first passes through an initial convolutional layer and downsampling operation to generate a basic feature map before entering the first stage. In the first stage, only the highest resolution path is retained (e.g., This consists of a series of standard convolutional units stacked together to extract fine spatial details. As the network depth increases, new low-resolution paths are gradually introduced in subsequent stages: for example, the second stage adds... The resolution path, in its third phase, is further expanded to... This forms a multi-scale parallel processing structure.
[0052] Each convolutional unit is typically a residual block or a bottleneck structure, and its basic form can be represented as follows: ,in, It uses a 1×1 convolution kernel for channel compression; It uses a 3×3 convolution kernel for spatial feature extraction; For the input feature map, To output the feature map, residual connections are used to ensure efficient gradient backpropagation.
[0053] Within each stage, paths of different resolutions interact with each other through upsampling and downsampling. Specifically, feature maps from high-resolution paths are passed to adjacent low-resolution paths via upsampling operations (e.g., bilinear interpolation or deconvolution) to enhance spatial detail perception; conversely, semantically rich features from low-resolution paths are downsampled by convolution with a stride of 2 and then injected into higher-resolution paths to improve contextual understanding.
[0054] This bidirectional information flow can be mathematically expressed as: upsampling fusion: Downsampling fusion: .in, Indicates an upsampling operation. This indicates a downsampling operation. Together, they construct a cross-resolution information bridging mechanism to ensure that each resolution path continuously shares structural and semantic information during the evolution process.
[0055] Finally, after all stages are completed, the network retains the output of the highest resolution path as the backbone feature because it contains both rich spatial details and global contextual information. In this embodiment of the disclosure, the image to be detected is... The input is an HRNet architecture, and the output is a set of feature maps at four key resolutions, denoted as . , This represents the base downsampling rate.
[0056] The feature maps in the feature map set are multi-scale spatial feature maps, used for content modeling and structural analysis in subsequent tasks. This structural design effectively avoids the spatial information loss problem caused by early downsampling in traditional CNNs, achieving a balance between high resolution preservation and multi-scale fusion, and providing powerful representation capabilities for complex visual tasks.
[0057] In practice, the image to be detected is input into the first high-resolution backbone network (using the HRNet-W32 architecture). Starting from the input resolution, the first high-resolution backbone network maintains four parallel high-resolution representation streams (with a resolution of [missing information - likely a specific value]). , , , The process involves exchanging multi-scale information through repeated cross-stage fusion modules. Ultimately, the outputs of the four resolution streams serve as multi-scale spatial feature maps. Among them, the first-scale spatial feature map Second-scale spatial feature map Third-scale spatial feature map Fourth-scale spatial feature map .
[0058] The initial noise feature map is input into a second high-resolution backbone network (using a lightweight HRNet-W18 architecture). The second high-resolution backbone network has a smaller number of channels, resulting in higher computational efficiency. It also outputs noise feature maps at four scales. Among them, the first-scale noise feature map Second-scale noise feature map Third-scale noise feature map Fourth-scale noise feature map .
[0059] The above scheme involves inputting the image to be detected into a first neural network, which generates a predetermined number of spatial feature maps using a predetermined number of high-resolution channels. These spatial feature maps are then used as multi-scale spatial feature maps. This multi-scale spatial feature map reflects the spatial variation characteristics at different resolution levels. Similarly, the initial noise feature map is input into a second neural network, which generates a predetermined number of noise feature maps using a predetermined number of high-resolution channels. These noise feature maps are then used as multi-scale noise feature maps. This multi-scale noise feature map reflects the noise variation characteristics at different resolution levels. By processing the image resolution, the model can adapt to tampering operations at different scales, effectively avoiding the problems of missed or false detections caused by the limitations of single-scale features.
[0060] In some embodiments, step 103 includes: Step 1031: Perform instance normalization processing on the multi-scale frequency domain feature map to obtain the normalized frequency domain feature map.
[0061] In practice, the multi-scale frequency domain feature map is subjected to instance normalization to obtain a normalized frequency domain feature map, thereby eliminating statistical differences between different samples. ,in, This is the normalized frequency domain feature map. This is a multi-scale frequency domain feature map.
[0062] Step 1032: Map the normalized frequency domain feature map to the feature space corresponding to the multi-scale spatial feature map to obtain the mapped frequency domain feature map.
[0063] In practice, depth-separable convolutional layers (with kernel size of ) are used. The output channel is the corresponding multi-scale spatial feature map. (Number of channels), mapping the normalized frequency domain feature map to the multi-scale spatial feature map. The same feature space yields a mapped frequency domain feature map. ,in, This is the mapped frequency domain feature map. This is the normalized frequency domain feature map.
[0064] Step 1033: Use the mapped frequency domain feature map as a query parameter, and the multi-scale spatial feature map as a key parameter and a value parameter, and determine the spatial attention map based on the query parameter and the key parameter.
[0065] In practice, the mapped frequency domain feature map The multi-scale spatial feature map will be used as a query parameter. As key and value parameters.
[0066] To reduce computational cost, multi-scale spatial feature maps conduct Convolutional projection yields the projected query parameters. and projected key parameters ,in, This is a learnable weight matrix.
[0067] Determine the spatial attention map based on the projected query parameters and projected key parameters. ,in, This is a spatial attention map. It is an all-1 vector used to sum along the channel dimension to obtain a single-channel attention weight map.
[0068] Step 1034: Generate an enhanced spatial feature map based on the spatial attention map and the multi-scale spatial feature map.
[0069] In practice, an enhanced spatial feature map is generated based on the spatial attention map and the multi-scale spatial feature map. Specifically, the spatial attention map... Compressed to using the Sigmoid function Intervals, and multi-scale spatial feature maps Element-wise multiplication yields the enhanced spatial feature map. ,in, To enhance the spatial feature map, For multi-scale spatial feature maps, This is a spatial attention map. This is the Sigmoid function.
[0070] The above process enables each item to dynamically enhance the corresponding spatial semantic response based on the frequency domain anomaly region, thereby improving the ability to locate tampering traces.
[0071] The above scheme involves instance normalization of the multi-scale frequency domain feature map to obtain a normalized frequency domain feature map. This normalized frequency domain feature map is then mapped to the feature space corresponding to the multi-scale spatial feature map, resulting in a mapped frequency domain feature map. The mapped frequency domain feature map is used as a query parameter, and the multi-scale spatial feature map is used as both a key and value parameter. A spatial attention map is then determined based on the query and key parameters. Finally, an enhanced spatial feature map is generated based on the spatial attention map and the multi-scale spatial feature map. This enhanced spatial feature map integrates information from both the frequency and spatial domains, enhancing the feature's ability to express image content.
[0072] In some embodiments, step 103 includes: Step 1035: The enhanced spatial feature map and the multi-scale noise feature map are spliced together to obtain a spliced feature map.
[0073] In practice, for each scale This will enhance the spatial feature map. and multi-scale noise feature maps The spliced feature map is obtained by splicing along the channel dimension. .
[0074] Step 1036: Determine the feature correlation between the first position and the second position in the spliced feature map, normalize the feature correlation to obtain the attention weight, and determine the position attention parameter based on the attention weight.
[0075] In practice, Figure 3 This is a schematic diagram of the dual attention mechanism according to an embodiment of this disclosure. Figure 3 As shown, dual modulation is performed using a positional attention module (PAM) and a channel attention module (CAM). Specifically, the positional attention parameters are determined by the positional attention module (PAM), and the channel attention parameters are determined by the channel attention module (CAM). The positional attention parameters and the channel attention parameters are then fused to obtain an initial fused feature map.
[0076] First position Second position It is a spliced feature map For any two spatial locations, determine the feature correlation between the first and second positions in the spliced feature map. ,in, First position With the second position The correlation of features between them , , It is a learnable projection matrix.
[0077] Attention weights are obtained by normalizing the feature correlations. ,in, For attention weights, First position With the second position The correlation of features between them.
[0078] The positional attention parameters are determined based on the attention weights. ,in, For positional attention parameters, For attention weights, .
[0079] Step 1037: Determine channel descriptors from the spliced feature map using global average pooling, determine channel weights based on the channel descriptors, and determine channel attention parameters based on the channel weights.
[0080] In practice, channel descriptors are determined from the spliced feature map using global average pooling. .
[0081] The channel weights at each location are obtained by passing the channel descriptor through a two-layer fully connected network. ,in, Channel weights for each position, , , , For ReLU function, This is the Sigmoid function.
[0082] Determine the channel attention parameters based on the channel weights. ,in, For channel attention parameters, Channel weights for each position, This is for splicing feature maps.
[0083] Step 1038: Determine the initial fused feature map based on the position attention parameters and the channel attention parameters.
[0084] In specific implementation, such as Figure 3As shown, the position attention parameters output by the Position Attention Module (PAM) are... Channel attention parameters output by the Channel Attention Module (CAM) Add them together, and through Convolutional integration yields an initial fused feature map. ,in, This is the initial fused feature map. For positional attention parameters, These are the channel attention parameters. Feature map concatenation. Compared with the original HRNet-W32 The number of channels is consistent across stages.
[0085] The above scheme involves concatenating the enhanced spatial feature map and the multi-scale noise feature map to obtain a concatenated feature map. The feature correlation between the first and second positions in the concatenated feature map is determined, and the feature correlation is normalized to obtain attention weights. Positional attention parameters are then determined based on these attention weights. Channel descriptors are determined from the concatenated feature map using global average pooling, channel weights are determined based on the channel descriptors, and channel attention parameters are determined based on the channel weights. An initial fused feature map is then determined based on the positional attention parameters and channel attention parameters. In this way, by fusing the enhanced spatial feature map and the multi-scale noise feature map, the initial fused feature map integrates multiple types of feature information, which can more accurately reflect the true state of the image to be detected, reduce the influence of external factors on the image tampering detection results, and improve robustness.
[0086] In some embodiments, step 103 includes: Step 103A: Map the number of channels of the initial fused feature map to a preset dimension to obtain lateral features.
[0087] In practice, the Consistency Aware Feature Pyramid Network (CAFPN) includes a horizontal projection module and a top-down smooth fusion module.
[0088] In the horizontal projection module, through four independent Convolutional layers initially fuse feature maps at each scale. The number of channels is uniformly mapped to 256 dimensions to obtain the lateral features. .
[0089] Step 103B: Determine the lowest resolution feature based on the horizontal feature, determine the highest resolution feature based on the lowest resolution feature, and use the highest resolution feature as the target fusion feature map.
[0090] In practice, within the top-down smooth fusion module, top-down feature fusion is performed starting from the level of the lowest resolution features: [The text then repeats the last two sentences, which are likely errors in the original Chinese.] ,in, Indicates by The smoothing unit consists of convolutional layers, LayerNorm normalization layers, and the GELU activation function.
[0091] for Calculate the resolution features for each resolution level sequentially: ,in, Bilinear interpolation is used, and the interpolation parameters are set to... This is to ensure the continuity of spatial alignment.
[0092] Finally, the highest resolution features output from the highest resolution layer are determined. The target fusion feature map is output by the Consistency Aware Feature Pyramid Network (CAFPN).
[0093] The above scheme maps the number of channels in the initial fused feature map to a preset dimension to obtain lateral features. Based on the lateral features, a minimum resolution feature is determined, and based on the minimum resolution feature, a maximum resolution feature is determined, which is then used as the target fused feature map. In this way, by aggregating the initial fused feature map into the target fused feature map, the initial fused feature map can be transformed into a single-scale feature map.
[0094] In some embodiments, step 104 includes: Step 1041: Perform a size transformation on the target fusion feature map to obtain a target fusion feature map of a preset size, wherein the target fusion feature map of the preset size is consistent with the size of the image to be detected.
[0095] In practice, the target feature map is fused. Perform a single bilinear upsampling operation to restore the spatial resolution of the target fused feature map to the size of the image to be detected. Among them, the target fusion feature map of the preset size is: .
[0096] Step 1042: Perform prediction processing on the target fusion feature map of the preset size to obtain the tampering probability of each pixel block, generate a tampering probability map based on the tampering probability of each pixel block, and use the tampering probability map as the image tampering detection result.
[0097] In practice, the target feature map of the preset size will be fused. Input mask prediction head. The mask prediction head contains, in sequence: a convolutional layer with padding of 1s and an output of 256 channels; a LayerNorm normalization layer, which normalizes along the channel dimension; a GELU activation function; and a convolutional layer that compresses the number of channels to 1.
[0098] Fusion feature maps of a target of a preset size using a mask prediction head Prediction processing is performed to obtain the tampering probability of each pixel block. The tampering probability of each pixel block is a single-channel tensor. A tampering probability map is generated based on the tampering probability of each pixel block, and this tampering probability map is used as the image tampering detection result.
[0099] Step 1043: The first intelligent agent performs semantic parsing on the multi-scale frequency domain feature map, the initial noise feature map, the initial fusion feature map, and the tampering probability map to obtain semantic features, and determines the tampering region and tampering method of the image to be detected based on the semantic features.
[0100] In specific implementation, after completing the pixel-level prediction of the tampered area, this embodiment further introduces an interpretability generation system based on multi-agent collaborative reasoning, which is used to transform the intermediate reasoning results of the model into a structured, human-readable image tampering analysis report.
[0101] Figure 4 This is a schematic diagram of the structure of a multi-agent interpretable generation system according to an embodiment of this disclosure. Figure 4 As shown, the interpretability generation system based on multi-agent collaborative reasoning consists of three parts: a perceptual agent, a reasoning agent, and a report generation agent. The three work together through a standardized interface.
[0102] The first agent can be a perceptual agent. The perceptual agent receives key intermediate features output by the heterogeneous resolution perceptual space-frequency fusion network during the forward inference process. These key intermediate features include at least one of the following: a multi-scale frequency domain feature map (DCT), an initial noise feature map (noise map), an initial fusion feature map (attention map), and a tampering probability map. The perceptual agent accesses these numerical features through a function call interface and inputs these key features into a pre-trained visual language model (VLM). The VLM performs semantic parsing of the key features based on prior knowledge to generate semantic features describing the image in natural language, and determines the tampered region and tampering method of the image to be detected based on these semantic features. For example, when the DCT coefficients of a certain region show obvious... When the noise is distributed in blocks, the VLM outputs "This area has significant DCT block effects, suspected to have been compressed or spliced by JPEG"; when the noise response intensity is significantly lower than the surrounding background in a local area, the output is "The noise statistics in the central area are abnormal, possibly due to erasure or content restoration".
[0103] Step 1044: Use the second intelligent agent to verify the tampered area and the tampering method to obtain the verification result.
[0104] In practice, the second intelligent agent can be an inference agent. The inference agent receives semantic features and the image to be detected from the perceptual agent, and automatically invokes a series of image analysis tools to perform causal inference and counterfactual verification based on the assumptions about the tampered region and the tampering method. Specifically, the operations that the inference agent can perform include: simulating operations such as cropping, rotating, affine transformation, and spectral superposition on the suspected tampered region, and comparing the consistency of features before and after the operation. For example, when the perceptual agent points out that "the textures of the upper left and lower right corners are highly similar," the inference agent will attempt to copy and paste the upper left region to the lower right corner, and then calculate the residuals between the synthesized image and the image to be detected in the noise domain or frequency domain. If the residuals are significantly reduced, the hypothesis of "copy-and-move forgery" is supported. The inference results are returned to the perceptual agent in a structured form (e.g., confidence score, operation log, comparison image) to correct or strengthen the semantic features.
[0105] Step 1045: In response to determining that the verification result is successful, an image detection report is generated using a third intelligent agent based on the tampered area and the tampering method.
[0106] In specific implementation, the third agent can be a report-generating agent. The report-generating agent integrates the semantic features from the perceptual agent and the verification results from the reasoning agent to automatically generate a visualized image detection report containing the following elements: (1) a tampered area location map, marking suspicious areas in the form of heatmaps or bounding boxes; (2) a tampering type judgment, such as "splicing and synthesizing", "copy-and-move forgery" or "erasure and repair"; (3) supporting evidence chains, including multimodal clues such as DCT anomalies, noise discontinuities, and attention-focused areas; and (4) a reasoning logic description, briefly explaining the process of deducing conclusions from feature observation. The image detection report is output in PDF or HTML format and can be selectively stored in an evidence database for manual review or judicial acceptance.
[0107] Through the aforementioned multi-agent collaborative mechanism, this embodiment not only achieves high-precision tamper detection, but also constructs a trusted reasoning chain from low-level features to high-level interpretation, significantly improving the transparency and auditability of the system in high-risk application scenarios such as digital forensics and news authenticity verification.
[0108] Based on the tampering probability map and intermediate feature map, the system automatically generates an interpretable image detection report through a multi-agent collaborative reasoning mechanism. The image detection report includes precise localization of the tampered area, forgery type identification, and a multimodal evidence chain supporting the conclusion. The entire detection process and interpretability generation share the same multimodal reasoning framework, forming an inseparable unified solution. This solution can efficiently and robustly handle various typical image tampering operations such as copy-paste, image stitching, and content erasure, achieving high-precision detection and transparent, reliable decision output.
[0109] The above scheme involves resizing the target fusion feature map to obtain a target fusion feature map of a preset size, which is consistent with the size of the image to be detected. Prediction processing is performed on the preset-size target fusion feature map to obtain the tampering probability of each pixel block. A tampering probability map is generated based on the tampering probability of each pixel block and used as the image tampering detection result. A first intelligent agent performs semantic parsing processing on the multi-scale frequency domain feature map, initial noise feature map, initial fusion feature map, and tampering probability map to obtain semantic features. Based on the semantic features, the tampered region and tampering method of the image to be detected are determined. A second intelligent agent verifies the tampered region and tampering method to obtain a verification result. If the verification result is successful, a third intelligent agent generates an image detection report based on the tampered region and tampering method. In this way, the image detection report provides users with intuitive evidence, making the image tampering detection results more convincing and conducive to application and promotion in fields with high requirements for result interpretability.
[0110] Through the above embodiments, an image to be detected is acquired, and a multi-scale frequency domain feature map is obtained by transforming the image to be detected, which can capture the frequency domain features in the image to be detected. An initial noise feature map is obtained by extracting noise features from the image to be detected, which can capture the noise features in the image to be detected. A multi-scale spatial feature map is obtained by resolving the resolution of the image to be detected. The multi-scale spatial feature map can reflect the spatial variation features at different resolution levels. Similarly, a multi-scale noise feature map is obtained by resolving the initial noise feature map. This multi-scale noise feature map can reflect the noise variation features at different resolution levels. Resolving the image resolution can effectively avoid the problem of missed or false detections caused by the limitations of single-scale features. Adding the multi-scale frequency domain feature map to the multi-scale spatial feature map yields an enhanced spatial feature map. The enhanced spatial feature map integrates information from both the frequency and spatial domains, which can enhance the expressive power of the features on the image content. Finally, the enhanced spatial feature map and the multi-scale noise feature map are fused to obtain an initial fused feature map. This initial fused feature map integrates multiple types of feature information, which can more accurately reflect the true state of the image to be detected, reduce the influence of external factors on the image tampering detection results, and improve robustness. The initial fused feature map is aggregated to obtain the target fused feature map. A tampering probability map is determined based on the target fused feature map, and this map is used as the image tampering detection result. An image detection report is generated based on the multi-scale frequency domain feature map, the initial noise feature map, the initial fused feature map, and the tampering probability map. This image detection report provides users with intuitive evidence, making the image tampering detection results more convincing and facilitating its application in fields where interpretability is crucial.
[0111] It should be noted that the method of this disclosure embodiment can be executed by a single device, such as a computer or server. The method of this embodiment can also be applied to a distributed scenario, where multiple devices cooperate to complete the task. In such a distributed scenario, one of these devices may execute only one or more steps of the method of this disclosure embodiment, and the multiple devices will interact with each other to complete the method described.
[0112] It should be noted that the above description describes some embodiments of this disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in a different order than that shown in the above embodiments and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0113] Based on the same inventive concept, corresponding to any of the above embodiments, this disclosure also provides an image detection device.
[0114] refer to Figure 5 The image detection device includes: The preprocessing module 301 is configured to acquire the image to be detected, perform transformation processing on the image to be detected to obtain a multi-scale frequency domain feature map, and perform extraction processing on the image to be detected to obtain an initial noise feature map; The resolution processing module 302 is configured to perform resolution processing on the image to be detected to obtain a multi-scale spatial feature map, and to perform resolution processing on the initial noise feature map to obtain a multi-scale noise feature map. The fusion processing module 303 is configured to add the multi-scale frequency domain feature map to the multi-scale spatial feature map to obtain an enhanced spatial feature map, perform fusion processing on the enhanced spatial feature map and the multi-scale noise feature map to obtain an initial fusion feature map, and perform aggregation processing on the initial fusion feature map to obtain a target fusion feature map; The image detection module 304 is configured to determine a tampering probability map based on the target fusion feature map, use the tampering probability map as the image tampering detection result, and generate an image detection report based on the multi-scale frequency domain feature map, the initial noise feature map, the initial fusion feature map, and the tampering probability map.
[0115] In some embodiments, the preprocessing module 301 includes: An image segmentation unit is configured to convert the image to be detected into a brightness image and divide the brightness image into multiple pixel blocks using a discrete cosine transform encoder. The transformation processing unit is configured to perform a two-dimensional discrete cosine transform on each pixel block to obtain frequency domain coefficients, and to map the frequency domain coefficients to the feature space to obtain pixel block features. The pooling processing unit is configured to reorganize the pixel block features of all pixel blocks to obtain an initial frequency domain feature map, and to perform pooling processing on the initial frequency domain feature map to obtain a multi-scale frequency domain feature map. The processing unit is configured to input the image to be detected into a differentiable convolutional layer, and use the differentiable convolutional layer to process the image to be detected to obtain an initial noise feature map.
[0116] In some embodiments, the resolution processing module 302 includes: The first resolution processing unit is configured to input the image to be detected into a first neural network, generate a preset number of spatial feature maps using a preset number of high-resolution channels in the first neural network, and use the preset number of spatial feature maps as multi-scale spatial feature maps. The second resolution processing unit is configured to input the initial noise feature map into the second neural network, generate a preset number of noise feature maps using a preset number of high-resolution channels in the second neural network, and use the preset number of noise feature maps as multi-scale noise feature maps.
[0117] In some embodiments, the fusion processing module 303 includes: The normalization processing unit is configured to perform instance normalization processing on the multi-scale frequency domain feature map to obtain a normalized frequency domain feature map. The mapping unit is configured to map the normalized frequency domain feature map to the feature space corresponding to the multi-scale spatial feature map to obtain the mapped frequency domain feature map. The spatial attention map determination unit is configured to use the mapped frequency domain feature map as a query parameter, the multi-scale spatial feature map as a key parameter and a value parameter, and determine the spatial attention map based on the query parameter and the key parameter. An enhanced spatial feature map generation unit is configured to generate an enhanced spatial feature map based on the spatial attention map and the multi-scale spatial feature map.
[0118] In some embodiments, the fusion processing module 303 includes: The stitching processing unit is configured to stitch the enhanced spatial feature map and the multi-scale noise feature map to obtain a stitched feature map; The position attention parameter determination unit is configured to determine the feature correlation between the first position and the second position in the spliced feature map, normalize the feature correlation to obtain attention weights, and determine the position attention parameters based on the attention weights. The channel attention parameter determination unit is configured to determine channel descriptors from the stitched feature map by global average pooling, determine channel weights based on the channel descriptors, and determine channel attention parameters based on the channel weights. The initial fused feature map determination unit is configured to determine the initial fused feature map based on the position attention parameters and the channel attention parameters.
[0119] In some embodiments, the fusion processing module 303 includes: The mapping unit is configured to map the number of channels of the initial fused feature map to a preset dimension to obtain lateral features; The target fusion feature map determination unit is configured to determine the lowest resolution feature based on the lateral feature, determine the highest resolution feature based on the lowest resolution feature, and use the highest resolution feature as the target fusion feature map.
[0120] In some embodiments, the image detection module 304 includes: The size transformation unit is configured to perform size transformation processing on the target fusion feature map to obtain a target fusion feature map of a preset size, wherein the target fusion feature map of the preset size is consistent with the size of the image to be detected; The prediction processing unit is configured to perform prediction processing on the target fusion feature map of the preset size to obtain the tampering probability of each pixel block, generate a tampering probability map based on the tampering probability of each pixel block, and use the tampering probability map as the image tampering detection result; The semantic parsing unit is configured to use a first agent to perform semantic parsing processing on the multi-scale frequency domain feature map, the initial noise feature map, the initial fusion feature map, and the tampering probability map to obtain semantic features, and to determine the tampering region and tampering method of the image to be detected based on the semantic features; The verification processing unit is configured to use a second intelligent agent to perform verification processing on the tampered area and the tampering method to obtain a verification result; The report generation unit is configured to generate an image detection report using a third agent based on the tampered area and the tampering method in response to determining that the verification result is a successful verification.
[0121] For ease of description, the above apparatus is described in terms of its functions, divided into various modules. Of course, in implementing this disclosure, the functions of each module can be implemented in one or more software and / or hardware.
[0122] The apparatus of the above embodiments is used to implement the corresponding image detection method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0123] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this disclosure also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the image detection method described in any of the above embodiments.
[0124] Figure 6 This embodiment illustrates a more specific hardware structure of an electronic device. The device may include a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, memory 1020, input / output interface 1030, and communication interface 1040 are interconnected internally via the bus 1050.
[0125] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0126] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.
[0127] The input / output interface 1030 is used to connect input / output modules to realize information input and output. Input / output modules can be configured as components within the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touchscreens, microphones, various sensors, etc., while output devices may include displays, speakers, vibrators, indicator lights, etc.
[0128] The communication interface 1040 is used to connect the communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB (Universal Serial Bus), network cable, etc.) or wireless means (such as mobile network, WIFI (Wireless Fidelity), Bluetooth, etc.).
[0129] Bus 1050 includes a pathway for transmitting information between various components of the device, such as processor 1010, memory 1020, input / output interface 1030, and communication interface 1040.
[0130] It should be noted that although the above-described device only shows the processor 1010, memory 1020, input / output interface 1030, communication interface 1040, and bus 1050, in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the embodiments of this specification, and not necessarily all the components shown in the figures.
[0131] The electronic devices described above are used to implement the corresponding image detection methods in any of the foregoing embodiments and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0132] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this disclosure also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to perform the image detection method as described in any of the above embodiments.
[0133] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.
[0134] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to execute the image detection method as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0135] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this application also provides a computer program product, including computer program instructions. When the computer program instructions are run on a computer, the computer performs the image detection method as described in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0136] It is understood that before using the technical solutions of the various embodiments in this disclosure, users will be informed of the type, scope of use, and usage scenarios of the personal information involved in an appropriate manner, and user authorization will be obtained.
[0137] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose, based on the prompt message, whether to provide personal information to the software or hardware such as electronic devices, applications, servers, or storage media performing the operations of this disclosed technical solution.
[0138] As an optional but not limited implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0139] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0140] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of this disclosure is limited to these examples; within the framework of this disclosure, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of the embodiments of this disclosure as described above, which are not provided in detail for the sake of brevity.
[0141] Additionally, to simplify the description and discussion, and to avoid obscuring the embodiments of this disclosure, the provided drawings may or may not show well-known power / ground connections to integrated circuit (IC) chips and other components. Furthermore, the apparatus may be shown in block diagram form to avoid obscuring the embodiments of this disclosure, and this also takes into account the fact that the details of implementation of these block diagram apparatuses are highly dependent on the platform on which the embodiments of this disclosure will be implemented (i.e., these details should be fully understood by those skilled in the art). While specific details (e.g., circuits) have been set forth to describe exemplary embodiments of this disclosure, it will be apparent to those skilled in the art that the embodiments of this disclosure can be implemented without these specific details or with variations thereof. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0142] Although this disclosure has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.
[0143] This disclosure is intended to cover all such substitutions, modifications, and variations that fall within the broad scope of this disclosure. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the protection scope of this disclosure.
Claims
1. An image detection method, characterized in that, The method includes: The image to be detected is acquired, and the image to be detected is transformed to obtain a multi-scale frequency domain feature map. The image to be detected is then extracted to obtain an initial noise feature map. The image to be detected is processed to obtain a multi-scale spatial feature map, and the initial noise feature map is processed to obtain a multi-scale noise feature map. The multi-scale frequency domain feature map is added to the multi-scale spatial feature map to obtain an enhanced spatial feature map. The enhanced spatial feature map and the multi-scale noise feature map are fused to obtain an initial fused feature map. The initial fused feature map is then aggregated to obtain a target fused feature map. Based on the target fusion feature map, a tampering probability map is determined, and the tampering probability map is used as the image tampering detection result. An image detection report is generated based on the multi-scale frequency domain feature map, the initial noise feature map, the initial fusion feature map, and the tampering probability map.
2. The method according to claim 1, characterized in that, The process of transforming the image to be detected to obtain a multi-scale frequency domain feature map and extracting noise feature maps from the image to be detected includes: The image to be detected is converted into a brightness image, and the brightness image is divided into multiple pixel blocks using a discrete cosine transform encoder. A two-dimensional discrete cosine transform is performed on each pixel block to obtain frequency domain coefficients, and the frequency domain coefficients are mapped to the feature space to obtain pixel block features; The pixel block features of all pixel blocks are recombined to obtain an initial frequency domain feature map, and the initial frequency domain feature map is then pooled to obtain a multi-scale frequency domain feature map. The image to be detected is input into a differentiable convolutional layer, and the image to be detected is processed by the differentiable convolutional layer to obtain an initial noise feature map.
3. The method according to claim 1, characterized in that, The process of performing resolution processing on the image to be detected to obtain a multi-scale spatial feature map, and performing resolution processing on the initial noise feature map to obtain a multi-scale noise feature map, includes: The image to be detected is input into a first neural network, and a preset number of spatial feature maps are generated using a preset number of high-resolution channels in the first neural network. The preset number of spatial feature maps are then used as multi-scale spatial feature maps. The initial noise feature map is input into the second neural network, and a preset number of noise feature maps are generated using a preset number of high-resolution channels in the second neural network. The preset number of noise feature maps are then used as multi-scale noise feature maps.
4. The method according to claim 1, characterized in that, The step of adding the multi-scale frequency domain feature map to the multi-scale spatial feature map to obtain the enhanced spatial feature map includes: The multi-scale frequency domain feature map is subjected to instance normalization to obtain a normalized frequency domain feature map; The normalized frequency domain feature map is mapped to the feature space corresponding to the multi-scale spatial feature map to obtain the mapped frequency domain feature map. The mapped frequency domain feature map is used as a query parameter, and the multi-scale spatial feature map is used as a key parameter and a value parameter. A spatial attention map is determined based on the query parameter and the key parameter. An enhanced spatial feature map is generated based on the spatial attention map and the multi-scale spatial feature map.
5. The method according to claim 1, characterized in that, The process of fusing the enhanced spatial feature map and the multi-scale noise feature map to obtain the initial fused feature map includes: The enhanced spatial feature map and the multi-scale noise feature map are stitched together to obtain a stitched feature map; The feature correlation between the first position and the second position in the spliced feature map is determined, the feature correlation is normalized to obtain the attention weight, and the position attention parameter is determined based on the attention weight. Channel descriptors are determined from the stitched feature map using global average pooling, channel weights are determined based on the channel descriptors, and channel attention parameters are determined based on the channel weights. The initial fused feature map is determined based on the location attention parameters and the channel attention parameters.
6. The method according to claim 1, characterized in that, The process of aggregating the initial fused feature map to obtain the target fused feature map includes: The number of channels in the initial fused feature map is mapped to a preset dimension to obtain lateral features; The lowest resolution feature is determined based on the horizontal feature, the highest resolution feature is determined based on the lowest resolution feature, and the highest resolution feature is used as the target fusion feature map.
7. The method according to claim 1, characterized in that, The step of determining a tampering probability map based on the target fusion feature map, using the tampering probability map as the image tampering detection result, and generating an image detection report based on the multi-scale frequency domain feature map, the initial noise feature map, the initial fusion feature map, and the tampering probability map includes: The target fusion feature map is subjected to a size transformation to obtain a target fusion feature map of a preset size, wherein the target fusion feature map of the preset size is consistent with the size of the image to be detected; The target fusion feature map of the preset size is subjected to prediction processing to obtain the tampering probability of each pixel block. A tampering probability map is generated based on the tampering probability of each pixel block, and the tampering probability map is used as the image tampering detection result. The first intelligent agent performs semantic parsing on the multi-scale frequency domain feature map, the initial noise feature map, the initial fusion feature map, and the tampering probability map to obtain semantic features, and determines the tampering region and tampering method of the image to be detected based on the semantic features; The verification result is obtained by using a second intelligent agent to verify the tampered area and the tampering method; In response to determining that the verification result is successful, an image detection report is generated using a third intelligent agent based on the tampered area and the tampering method.
8. An image detection device, characterized in that, include: The preprocessing module is configured to acquire the image to be detected, perform transformation processing on the image to be detected to obtain a multi-scale frequency domain feature map, and extract the image to be detected to obtain an initial noise feature map; The resolution processing module is configured to perform resolution processing on the image to be detected to obtain a multi-scale spatial feature map, and to perform resolution processing on the initial noise feature map to obtain a multi-scale noise feature map. The fusion processing module is configured to add the multi-scale frequency domain feature map to the multi-scale spatial feature map to obtain an enhanced spatial feature map, perform fusion processing on the enhanced spatial feature map and the multi-scale noise feature map to obtain an initial fusion feature map, and perform aggregation processing on the initial fusion feature map to obtain a target fusion feature map; The image detection module is configured to determine a tampering probability map based on the target fusion feature map, use the tampering probability map as the image tampering detection result, and generate an image detection report based on the multi-scale frequency domain feature map, the initial noise feature map, the initial fusion feature map, and the tampering probability map.
9. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor, when executing the program, implements the method as claimed in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores computer instructions for causing a computer to perform the method according to any one of claims 1 to 7.