Deep counterfeit image detection method based on multi-level feature space attention mechanism

Through the combination of multi-level feature space attention mechanism and feature pyramid network, the problem of detail information being ignored in depth forged image detection in the prior art is solved, and a more accurate depth forged image detection effect is achieved.

CN120339770APending Publication Date: 2025-07-18THE FIRST RES INST OF MIN OF PUBLIC SECURITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510422913.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing neural networks tend to ignore some detailed information in the detection of deep fake images, resulting in poor detection results.

Method used

A multi-level feature spatial attention mechanism is adopted to extract multi-layer feature maps of different scales through feature pyramid networks, and the spatial attention mechanism is used to strengthen effective local features, and classify and predict them in combination with leveling and splicing operations.

Benefits of technology

The effect of deep fake image detection is improved, especially the recovery of missing information in blurred images, reducing unimportant information interference from network operations, and improving the processing capability of areas related to classification tasks in the image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339770A_ABST
    Figure CN120339770A_ABST
Patent Text Reader

Abstract

The invention discloses a deep counterfeit image detection method based on a multi-level feature space attention mechanism, and the method comprises the steps: carrying out the multi-layer feature extraction of different scales of an input image, then obtaining the fusion features of a deep layer and a shallow layer in an input blurred image, carrying out the average pooling operation of a feature map of the last layer, and carrying out the detection of a deep counterfeit image. Obtaining a feature map of the backbone network for prediction; and finally, converting each feature map into a one-dimensional vector by utilizing a leveling operation and a splicing operation, and sending the one-dimensional vector into a classification layer to obtain a final prediction value, namely a prediction result of whether the input image is a deep forged image or not. According to the method, multi-level features are effectively utilized, and effective local features are specifically enhanced by a space attention mechanism, so that the purpose of distinguishing true and false faces is achieved, and the detection effect of a deep counterfeit image is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and specifically relates to a deepfake image detection method based on a multi-level feature space attention mechanism. Background Art

[0002] Deepfake image detection is an important technology for dealing with AI-Generated Content (AIGC), aiming to identify fake images or videos synthesized through deep learning technologies (such as GAN, diffusion models, etc.). Deepfake image detection is widely used in various fields, such as social media and content moderation (platforms (such as Meta, Twitter) use detection technologies to automatically mark or delete fake content), justice and legal forensics (identifying forged evidence to assist court forensics (such as authenticity identification of Deepfake videos)), financial security (preventing face recognition systems from being attacked by fake images / videos (such as bank identity authentication)), film and television and copyright protection (detecting copyright infringement of unauthorized AI-generated content (such as AI-generated star portraits)).

[0003] Existing neural networks generally process the input image layer by layer, through a series of convolutional operations, and use the feature map of the last layer to predict the classification value. Although this method eliminates some redundant information, it will cause some detailed information that has a significant impact on the result to be ignored, resulting in poor detection effects for deepfake images. Summary of the Invention

[0004] Aiming at the deficiencies of the existing technology, the present invention aims to provide a deepfake image detection method based on a multi-level feature space attention mechanism, which can distinguish true and false human faces by effectively using multi-level features and specifically strengthening effective local features through the spatial attention mechanism, and optimize the detection effect of deepfake images.

[0005] To achieve the above object, the present invention adopts the following technical solutions:

[0006] The deepfake image detection method based on a multi-level feature space attention mechanism specifically includes the following steps:

[0007] S1. Use a feature pyramid network to extract multi-level feature maps of different scales from the input image;

[0008] S2. Respectively convert the multiple feature maps of different scales obtained in step S1 into the same number of channels through 1×1 convolution. Then, take the feature map with the largest size as the first layer and the feature map with the smallest size as the last layer. For each layer of feature map, use bilinear interpolation to double the size of the next layer of feature map and fuse it with the current layer of feature map using element-wise addition to obtain the feature map for prediction of the current layer. The feature map of the last layer is directly used as the feature map for prediction of the current layer;

[0009] S3. Perform average pooling operation on the feature map of the last layer to obtain the feature map for prediction of the backbone network;

[0010] S4. Finally, use flattening operation and concatenation operation to convert the multi-layer feature maps for prediction obtained in step S2 and the feature map for prediction of the backbone network obtained in step S3 into one-dimensional vectors, and then send them into the classification layer to obtain the final prediction value, that is, the prediction result of whether the input image is a deepfake image.

[0011] Further, for each layer of feature map for prediction obtained in step S2, use the feature space attention mechanism to perform the following processing:

[0012] For a feature map with height H×width W×channel dimension C, divide it into G groups according to the channel dimension. Each group has a vector representation X = {x1,...x m} in space, where x i , i = 1, 2,..., m represents the features of each group; First, obtain the global semantic vector g learned by the entire group through global average pooling F gp (·) as shown in the following formula:

[0013]

[0014] Then, perform dot product operation on the obtained global semantic vector g and each feature x i in the group to generate the key coefficient corresponding to each feature;

[0015] To avoid the influence caused by the bias size of the coefficients between different samples, perform normalization again. Use the Sigmoid function to scale the weights generated in space according to the key coefficient of each feature to obtain the enhanced feature as shown in the following formula:

[0016]

[0017] g·x i represents the key coefficient of feature x i , is the result feature group of the group it belongs to, and the result feature groups of G groups form the enhanced feature map;

[0018] In step S4, the multi-layer enhanced feature maps corresponding to the multi-layer feature maps for prediction obtained in step S2 and the feature maps for prediction of the backbone network obtained in step S3 are converted into one-dimensional vectors by using a flattening operation and a splicing operation, and then fed into the classification layer to obtain the final predicted value, that is, the result of whether the input image is a deepfake image.

[0019] Further, in step S1, EfficientNet-B4 is used to extract feature maps of multiple different scales from the input image.

[0020] Further, in step S1, the different scales are 112×112×24, 56×56×32, 28×28×56, 14×14×160, and 7×7×448 respectively.

[0021] The present invention also provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the above method is implemented.

[0022] The present invention also provides a computer device, including a processor and a memory, where the memory is used to store a computer program; when the processor executes the computer program, the above method is implemented.

[0023] The beneficial effects of the present invention are as follows:

[0024] 1. For blurred images, the information missing in feature extraction may be mined or even restored under large-scale feature maps. For example, the tampered face contour or facial features. The present invention makes full use of multi-scale features for classification prediction, and combines the spatial attention mechanism to specifically strengthen the effective local features, which can effectively improve the effect of deepfake image detection.

[0025] 2. Using the present invention, attention can avoid the interference of unimportant information, concentrate on processing the regions most relevant to the classification task in the image, and reduce network operations. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 is the overall implementation framework diagram of the method of the embodiment of the present invention;

[0027] Figure 2 is the flow chart of the feature space attention mechanism in the method of the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0028] The following will further describe the present invention with reference to the drawings. It should be noted that this embodiment is based on the present technical solution and gives detailed implementation manners and specific operation processes, but the protection scope of the present invention is not limited to this embodiment.

[0029] This embodiment provides a deepfake image detection method based on a multi-level feature space attention mechanism. Considering both local and global aspects, feature maps of various sizes are obtained through levels of different depths, and then features with deep and shallow representations are fused to obtain more accurate and comprehensive results. Among them, shallow features are more refined, mainly local pixel-level information, while deep features contain global information and can obtain context information and accurate abstract semantic information. As Figure 1 shown, the specific steps are as follows:

[0030] S1. Perform five-layer feature extraction on the input image.

[0031] The method of this embodiment uses a Feature Pyramid Network (FPN) to make multi-level predictions by fully utilizing multi-scale information and can better learn features of different scales existing in the input blurred image. Since each scale in the blurred image contains information of different scales, using FPN is beneficial for extracting various semantic features of the input image. In this embodiment, five different-scale feature maps are extracted, and their sizes are 112×112×24, 56×56×32, 28×28×56, 14×14×160, and 7×7×448 respectively. Each feature map represents the features of the input image at a different scale. Specifically, EfficientNet-B4 can be used to perform five different-scale feature extractions on the input image.

[0032] S2. To obtain the features of the deep and shallow layers fused in the input blurred image, first, the five different-scale feature maps extracted in step S1 are respectively converted to the same number of channels through 1×1 convolution. Then, taking the feature map with the largest size as the first layer and the feature map with the smallest size as the last layer, for each layer of the feature map, the feature map of the next layer is enlarged twice by bilinear interpolation and then fused with the current layer of the feature map using element-wise addition to obtain the feature map for prediction of the current layer. The feature map of the last layer (i.e., the feature map with the smallest size) is directly used as the feature map for prediction of the current layer. In this way, the feature map of each layer for prediction can contain both global and local features. The feature map of each layer for prediction fuses various information of different scales, can enrich the semantic information of each layer, and can better learn the features of different scales existing in the blurred image.

[0033] S3. To obtain more accurate information, the method of this embodiment retains the prediction features of the backbone network. The feature map of the last layer is subjected to average pooling operation to obtain the feature map for prediction of the backbone network. Therefore, a total of six feature maps for prediction can be obtained by adding step S2 and step S3 together.

[0034] S4. Use the leveling operation and the splicing operation to convert the five layers of feature maps for prediction obtained in step S2 and the feature maps for prediction of the backbone network obtained in step S3 into one-dimensional vectors, and then send them into the classification layer to obtain the final predicted value, that is, the prediction result of whether the input image is a deepfake image.

[0035] Furthermore, the feature map with complete semantics is composed of multiple sub-features, and these sub-features often exist in the form of channel groups. The set of sub-features of all channel groups constitutes a complete layer of feature maps. In the traditional feature extraction process, the same method is used to process different sub-features, and the sub-features of each channel group are affected by background noise. Therefore, the semantics of the sub-features of each channel group are not clear, resulting in misclassification of the input image. For low-resolution blurred images, the above problem is more serious because the foreground and background in blurred images are not clearly defined, and the wrong influence brought by background noise may be amplified during the feature extraction process. To solve this problem, in the method of this embodiment, for the five layers of feature maps for prediction obtained in step S2, the Feature Spatial Attention (FSA) mechanism is used to generate respective spatial attention masks for the sub-features of each channel group of the feature map, so that the model pays more attention to foreground information and ignores background noise.

[0036] The structure of the feature space attention mechanism is as Figure 2 shown. For a feature map of H×W×C (H: height, W: width, C: channel dimension), it is divided into G groups according to the channel dimension, and each group has a vector representation X = {x1,..., x m} in space, where x i , i = 1, 2,..., m represents the features of each group. Assuming that during the learning stage of the model, specific semantic information of a certain region will be gradually captured, then in this region, features with a larger vector length and similar vector directions can be obtained, while other parts are zero vectors that do not represent any information. To suppress the noise introduced by compression processing during transmission and highlight important semantic feature regions, this embodiment selects to better learn important feature regions with the overall information in the space of the entire group. Taking one of the groups as an example, first, the global semantic vector g learned by the entire group is obtained through global average pooling F gp (·) as shown in the following formula:

[0037]

[0038] Then, the dot product operation is performed on the obtained global semantic vector g and each feature x i in the group to generate key coefficients corresponding to each feature.

[0039] To avoid the influence caused by the bias magnitude of coefficients between different samples, normalization is performed. The weights generated in space are scaled according to the key coefficients of each feature through the Sigmoid function to obtain enhanced features. As shown in the following formula:

[0040]

[0041] g·x i represents the key coefficient of feature x i and is the result feature group of the corresponding group. The result feature groups of group G form the enhanced feature map.

[0042] In step S4, the five enhanced feature maps corresponding to the five feature maps for prediction obtained in step S2 and the feature map for prediction of the backbone network obtained in step S3 are converted into one-dimensional vectors by using the flattening operation and the splicing operation, and then sent into the classification layer to obtain the final prediction value, that is, the result of whether the input image is a deepfake image.

[0043] For those skilled in the art, various corresponding changes and deformations can be given according to the above technical solutions and concepts, and all these changes and deformations should be included in the protection scope of the claims of the present invention.

Claims

1. A method for detecting deepfake images based on a multi-level feature space attention mechanism, characterized in that Specifically, it includes the following steps: S1. Use a Feature Pyramid Network to extract multi-scale feature maps of the input image; S2. Respectively convert the multi-scale feature maps obtained in step S1 into the same number of channels through 1×1 convolution. Then, take the feature map with the largest size as the first layer and the feature map with the smallest size as the last layer. For each layer of feature map, use bilinear interpolation to double the size of the next layer of feature map and fuse it with the current layer of feature map using element-wise addition to obtain the feature map for prediction of the current layer. The feature map of the last layer is directly used as the feature map for prediction of the current layer; S3. Perform average pooling operation on the feature map of the last layer to obtain the feature map for prediction of the backbone network; S4. Finally, use flattening operation and concatenation operation to convert the multi-layer feature maps for prediction obtained in step S2 and the feature map for prediction of the backbone network obtained in step S3 into one-dimensional vectors, and then send them into the classification layer to obtain the final prediction value, that is, the prediction result of whether the input image is a deepfake image.

2. The method according to claim 1, wherein For each layer of feature map for prediction obtained in step S2, the following processing is performed using a feature space attention mechanism: For a feature map with a height H × width W × channel dimension C, it is divided into G groups according to the channel dimension, and each group has a vector representation X = {x1, … x m} in space, where x i , i = 1, 2, …, m represents the features of each group; first, the global semantic vector g learned by the entire group is obtained through global average pooling F gp (·) as shown in the following formula: Then, perform a dot product operation on the obtained global semantic vector g and each feature x in the group i to generate key coefficients corresponding to each feature; To avoid the influence caused by the bias magnitude of coefficients between different samples, normalization is performed again. The Sigmoid function is used to scale each weight generated in the space according to the key coefficient of each feature, resulting in enhanced features As shown in the following formula: g·x i represents the key coefficient of feature x i , which is the result feature group of the corresponding group. The result feature groups of group G form an enhanced feature map; In step S4, the multi-layer enhanced feature maps corresponding to the multi-layer feature maps for prediction obtained in step S2 and the feature map for prediction of the backbone network obtained in step S3 are converted into one-dimensional vectors using flattening operation and concatenation operation, and then sent into the classification layer to obtain the final prediction value, that is, the result of whether the input image is a deepfake image.

3. The method according to claim 1, characterized in that, In step S1, EfficientNet-B4 is used to extract feature maps of various different scales of the input image.

4. The method according to claim 1, wherein In step S1, the different scales are respectively 112×112×24, 56×56×32, 28×28×56, 14×14×160, 7×7×448.

5. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described in any one of claims 1-4 is implemented.

6. A computer device, characterized in that, It includes a processor and a memory. The memory is used to store a computer program; when the processor executes the computer program, the method described in any one of claims 1-4 is implemented.