A method and device for detecting facial images generated by local GAN ​​with adaptive frequency perception

Through the adaptive frequency perception module and the improved detection network structure, the generalization and robustness problems of the local GAN ​​generated face detection method are solved, and efficient detection on different data sets and anti-interference ability against Gaussian noise are achieved.

CN116959063BActive Publication Date: 2025-09-12NANJING UNIV OF INFORMATION SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310743069.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-21
Publication Date
2025-09-12
Estimated Expiration
2043-06-21

AI Technical Summary

Technical Problem

The existing technology lacks a local GAN-generated face detection method with good generalization performance, especially poor detection effect on cross-datasets and insufficient robustness to low-intensity Gaussian noise.

Method used

A local GAN-generated face image detection model is adopted with an adaptive frequency perception module and a detection network. The irrelevant frequency information is filtered out by the adaptive frequency perception module. Combined with the deletion of some residual blocks in the Xception network and the addition of an ECA attention module, the detection accuracy and robustness of face images generated by local GAN ​​are improved.

Benefits of technology

While achieving excellent detection performance on the in-library test set, it also has good generalization across database datasets and maintains high detection accuracy under low-intensity Gaussian noise, improving the detection accuracy and robustness of local GAN-generated face images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116959063B_ABST
    Figure CN116959063B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for detecting facial images generated by a local GAN ​​using adaptive frequency perception, comprising: obtaining a facial image; detecting the facial image using a trained local GAN-generated facial image detection model to obtain a local GAN-generated facial image detection result; the local GAN-generated facial image detection model includes an adaptive frequency perception module and a detection network; wherein the adaptive frequency perception module includes a two-dimensional discrete cosine transform submodule, a learnable filter bank, and a two-dimensional inverse discrete cosine transform submodule arranged in sequence, the learnable filter bank including a high-pass filter and a learnable filter; and the detection network is obtained by deleting the fourth to eleventh residual blocks in an Xception network and adding an ECA attention module to the first to third and twelfth residual blocks. The present invention can accurately detect facial images generated by local GAN ​​and has good generalization and robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and device for detecting facial images generated by a local GAN ​​with adaptive frequency perception, and belongs to the technical field of facial image detection. Background Art

[0002] With the advent of the digital age, identity authentication technologies based on biometric information such as faces, fingerprints, and voiceprints have become increasingly popular. Due to the advantages of faces, such as low acquisition cost, rich information, and no physical contact, identity authentication services generally use facial images as identification information. While this has brought many benefits to people's lives, it also comes with many risks. With the emergence and development of face generation technology based on generative adversarial networks (GANs), "face scanning" technology has become less secure. Therefore, research on GAN-generated face detection technology has important application value and practical significance.

[0003] As one of the hot research directions in the field of blind digital image forensics, many detection methods have been proposed in recent years to combat malicious GAN-generated faces. In the task of detecting global GAN-generated face images, Liu et al. [Liu ZZ, Qi XJ, Torr PH S. Global texture enhancement for fake face detection in the wild [C]. Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2020: 8060-8069.] used the Gram matrix to suppress image content and combined it with a ResNet network to extract global texture features, thereby distinguishing real faces from GAN-generated faces. Wang et al. [Wang J, Zeng K, Ma B, et al. GAN-generated fake face detection via two-stream CNN with PRNU in the wild [J]. Multimedia Tools and Applications, 2022, 81(29): 42527-42545.] used a two-stream network to extract the RGB spatial features and PRNU features of the image based on the unique trace of photo response non-uniformity (PRNU) in natural images, and then fused the two features and input them into the classifier to distinguish between real faces and GAN-generated faces. However, since these methods are all based on the assumption that the entire face is generated by GAN, they are not very effective in detecting local GAN-generated face images.

[0004] Chen et al. [Chen BJ, Ju XW, Xiao B, et al. Locally GAN-generated face detection based on an improved Xception [J]. Information Sciences, 2021, 572: 16-28.] proposed a detection method for local GAN-generated faces. This method deletes some network layers in the Xception network and adds a Squeeze-and-excitation (SE) module to improve the network's feature extraction performance. In addition, by introducing an Inception module with dilated convolution and a feature pyramid network to extract multi-scale and multi-level features for final classification, this method achieved relatively good detection results in simple scenarios. However, this method lacks generalization on cross-datasets and is not very effective in actual use.

[0005] Based on the above analysis, there is currently a lack of local GAN-generated face detection methods with good generalization performance on the market. Summary of the Invention

[0006] To address the problems existing in the prior art, the present invention proposes an adaptive frequency-aware method and device for detecting local GAN-generated facial images. By optimizing the model's network structure, the model is able to better capture high-frequency artifacts in local GAN-generated facial images, thereby accurately detecting local GAN-generated facial images. The method of the present invention has superior generalization and robustness to the prior art.

[0007] In order to solve the above technical problems, the present invention adopts the following technical means:

[0008] In a first aspect, the present invention proposes an adaptive frequency-aware local GAN-generated face image detection method, comprising the following steps:

[0009] Get face image;

[0010] Detecting the face image using the trained local GAN-generated face image detection model to obtain a local GAN-generated face image detection result;

[0011] The local GAN ​​generated face image detection model includes an adaptive frequency perception module and a detection network; wherein, the adaptive frequency perception module includes a two-dimensional discrete cosine transform submodule, a learnable filter group and a two-dimensional discrete cosine inverse transform submodule arranged in sequence, and the learnable filter group includes a high-pass filter and a learnable filter; the detection network is obtained by deleting the fourth to eleventh residual blocks in the Xception network and adding an ECA attention module to the first to third and twelfth residual blocks.

[0012] In combination with the first aspect, further, the face image is detected using the trained local GAN ​​generated face image detection model to obtain a local GAN ​​generated face image detection result, including:

[0013] extracting a frequency-aware image component of the facial image from the facial image using the adaptive frequency-aware module;

[0014] The detection network is used to perform feature extraction on the frequency perception image component of the facial image and local GAN ​​generated facial image judgment to obtain a local GAN ​​generated facial image detection result.

[0015] In combination with the first aspect, further, extracting a frequency-aware image component of the facial image from the facial image using the adaptive frequency-aware module includes:

[0016] The two-dimensional discrete cosine transform submodule in the adaptive frequency perception module is used to perform a two-dimensional discrete cosine transform on the face image to obtain a frequency spectrum F of the face image:

[0017]

[0018] Where F(u,v) represents the spectral coefficient of row u and column v of the spectrum F after two-dimensional discrete cosine transform, c(u) and c(v) represent the compensation coefficients, R(i,j) represents the pixel value of row i and column j of the face image, and the values ​​of i, j, u, and v all range from 0 to M-1, where M is the side length of the face image.

[0019] The spectrum F is filtered using the learnable filter bank in the adaptive frequency perception module to obtain the frequency component F′:

[0020] F′=P*F

[0021] P(u,v)=H(u,v)+K(u,v)

[0022]

[0023] K(u,v)=σ(w u,v )

[0024] Where P is the mask matrix P of the learnable filter group, P(u,v) represents the weight value of the u-th row and the v-th column in the mask matrix P of the learnable filter group, H(u,v) represents the weight value of the u-th row and the v-th column in the mask matrix H of the high-pass filter, K(u,v) represents the weight value of the u-th row and the v-th column in the mask matrix K of the learnable filter, w u,v represents the learnable weight value in K(u,v), and σ() is the restriction function;

[0025] The frequency component F′ is converted back to the RGB spatial domain using the two-dimensional inverse discrete cosine transform submodule in the adaptive frequency perception module to obtain the frequency perception image component Y of the face image:

[0026]

[0027] Among them, Y(i,j) represents the pixel value of row i and column j in the frequency perception image component Y, and F′(u,v) represents the spectral coefficient of row u and column v in the frequency component F′.

[0028] In combination with the first aspect, further, the detection network includes a first convolutional layer, a second convolutional layer, first to fourth residual blocks, a first depth-separable convolutional layer, a second depth-separable convolutional layer, a global average pooling layer and a fully connected layer arranged in sequence; wherein, the first convolutional layer, the second convolutional layer, the first depth-separable convolutional layer and the second depth-separable convolutional layer are respectively followed by a ReLU activation layer; each residual block contains 2 depth-separable convolutional layers, 1 maximum pooling layer and 1 ECA attention module, and the ECA attention module is arranged in sequence with 1 global average pooling layer, 1 one-dimensional convolutional layer and a Sigmoid activation layer.

[0029] In combination with the first aspect, further, the first convolutional layer, the second convolutional layer, the first to the fourth residual blocks, the first depth-separable convolutional layer, and the second depth-separable convolutional layer in the detection network are used to perform convolution calculations on the frequency-perceived image components in sequence to obtain a feature map; the feature map output by the second depth-separable convolutional layer is pooled using the global average pooling in the detection network to obtain a pooled feature map; the fully connected layer in the detection network is used to perform a fully connected calculation on the pooled feature map to obtain a one-dimensional tensor containing two numerical values, and the subscripts of the two numerical values ​​are 0 and 1 respectively; the sizes of the two numerical values ​​are compared, and the subscript of the larger numerical value is used as the output of the detection network; when the detection network output is 0, it indicates that the face image is a face image generated by a local GAN, and when the detection network output is 1, it indicates that the face image is a real face image.

[0030] In combination with the first aspect, further, the training process of the local GAN ​​to generate the face image detection model is as follows:

[0031] (1) obtaining a face image library, wherein the face image library includes local GAN-generated face images and their label values, and real face images and their label values;

[0032] (2) Divide the face images in the face image library into training images and test images;

[0033] (3) Initialize the network parameters of the local GAN-generated face image detection model;

[0034] (4) Inputting the training image into the local GAN ​​generated face image detection model to obtain the local GAN ​​generated face image detection result corresponding to the training image;

[0035] (5) Generate face image detection results and label values ​​of training images based on the local GAN ​​of the training image, and calculate the cross entropy loss;

[0036] (6) When the cross entropy loss is lower than the threshold, the training is completed and the trained local GAN ​​generated face image detection model is obtained, otherwise return to step (4).

[0037] In combination with the first aspect, further, the calculation formula of the cross entropy loss is as follows:

[0038]

[0039] Among them, L cross Indicates the cross entropy loss value between the detection result and the label value output by the local GAN ​​generated face image detection model at the current moment, y R is the label value of the Rth face image, p R The detection result of the Rth face image output by the local GAN-generated face image detection model at the current moment, where n is the total number of face images.

[0040] In a second aspect, the present invention proposes an adaptive frequency-aware local GAN-generated face image detection device, comprising:

[0041] An image acquisition module, used to acquire a face image;

[0042] An image detection module is used to detect the face image using the trained local GAN-generated face image detection model to obtain a local GAN-generated face image detection result;

[0043] In the image detection module, the local GAN-generated face image detection model includes an adaptive frequency perception module and a detection network; wherein the adaptive frequency perception module includes a two-dimensional discrete cosine transform submodule, a learnable filter group and a two-dimensional discrete cosine inverse transform submodule arranged in sequence, and the learnable filter group includes a high-pass filter and a learnable filter; the detection network is obtained by deleting the fourth to eleventh residual blocks in the Xception network and adding an ECA attention module to the first to third and twelfth residual blocks.

[0044] In conjunction with the second aspect, further, the image detection module is specifically configured to:

[0045] extracting a frequency-aware image component of the facial image from the facial image using the adaptive frequency-aware module;

[0046] The detection network is used to perform feature extraction on the frequency perception image component of the facial image and local GAN ​​generated facial image judgment to obtain a local GAN ​​generated facial image detection result.

[0047] In combination with the second aspect, further, the image detection module extracts a frequency-perceived image component of the face image from the face image using the adaptive frequency-perceived module, including:

[0048] The two-dimensional discrete cosine transform submodule in the adaptive frequency perception module is used to perform a two-dimensional discrete cosine transform on the face image to obtain a frequency spectrum F of the face image:

[0049]

[0050] Where F(u,v) represents the spectral coefficient of row u and column v of the spectrum F after two-dimensional discrete cosine transform, c(u) and c(v) represent the compensation coefficients, R(i,j) represents the pixel value of row i and column j of the face image, and the values ​​of i, j, u, and v all range from 0 to M-1, where M is the side length of the face image.

[0051] The spectrum F is filtered using the learnable filter in the adaptive frequency perception module to obtain the frequency component F′:

[0052] F′=P*F

[0053] P(u,v)=H(u,v)+K(u,v)

[0054]

[0055] K(u,v)=σ(w u,v )

[0056] Where P is the mask matrix P of the learnable filter group, P(u,v) represents the weight value of the u-th row and the v-th column in the mask matrix P of the learnable filter group, H(u,v) represents the weight value of the u-th row and the v-th column in the mask matrix H of the high-pass filter, K(u,v) represents the weight value of the u-th row and the v-th column in the mask matrix K of the learnable filter, w u,v represents the learnable weight value in K(u,v), and σ() is the restriction function;

[0057] The frequency component F′ is converted back to the RGB spatial domain using the two-dimensional inverse discrete cosine transform submodule in the adaptive frequency perception module to obtain the frequency perception image component Y of the face image:

[0058]

[0059] Among them, Y(i,j) represents the pixel value of row i and column j in the frequency perception image component Y, and F′(u,v) represents the spectral coefficient of row u and column v in the frequency component F′.

[0060] The following advantages can be obtained by adopting the above technical means:

[0061] This invention proposes an adaptive frequency-aware method and device for detecting facial images generated by local GANs. This method utilizes a local GAN-generated facial image detection model, including an adaptive frequency-aware module and a detection network, to detect facial images and determine whether an input facial image is generated by a local GAN. To address the drawback of GAN upsampling, which can cause high-frequency artifacts in the generated image, this invention utilizes an adaptive frequency-aware module to adaptively filter out irrelevant frequency information in the image, thereby highlighting high-frequency artifacts in local GAN-generated facial images. To better capture high-frequency artifacts in local GAN-generated facial images, this invention removes the fourth through eleventh residual blocks in the Xception network, shallowing the network to focus the detection network on local high-frequency texture information. Furthermore, an ECA attention module is added to the first through third and twelfth residual blocks to focus the detection network on important features, thereby improving the detection accuracy of local GAN-generated facial images. This invention achieves excellent detection performance on a test set within a library while also generalizing well to cross-library datasets and exhibiting robustness to low-intensity Gaussian noise. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 A schematic diagram of the steps of a method for generating facial image detection using an adaptive frequency-aware local GAN ​​according to the present invention;

[0063] Figure 2 A schematic diagram of a local GAN ​​generating a face image detection model in an embodiment of the present invention;

[0064] Figure 3 Schematic diagram of the structure of the adaptive frequency sensing module in an embodiment of the present invention;

[0065] Figure 4 A schematic diagram of the structure of a detection network in an embodiment of the present invention;

[0066] Figure 5 Schematic diagram of the training process of generating a face image detection model using a local GAN ​​in an embodiment of the present invention;

[0067] Figure 6Schematic diagram of facial image generation by local GAN ​​using the four deep restoration methods GC, CA, DF, and RFR in an embodiment of the present invention. DETAILED DESCRIPTION

[0068] The technical solution of the present invention will be further described below with reference to the accompanying drawings:

[0069] The technical solution of the present invention is described in detail below through the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations on the technical solution of the present invention. In the absence of conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.

[0070] Example 1:

[0071] This embodiment introduces an adaptive frequency-aware local GAN ​​generated face image detection method, such as Figure 1 As shown, it mainly includes the following steps:

[0072] Step A: Acquire a face image. In this application, the face image is preferably square.

[0073] Step B: Use the trained local GAN ​​generated face image detection model to detect the face image, determine whether the face image is a local GAN ​​generated face image, and obtain the local GAN ​​generated face image detection result.

[0074] The present invention constructs an end-to-end trained convolutional neural network to complete the detection of local GAN-generated facial images. The input of the convolutional neural network is the facial image to be detected, and the output of the convolutional neural network is the detection result. In the present invention, the convolutional neural network is called a local GAN-generated facial image detection model.

[0075] In the embodiment of the present invention, the local GAN ​​generates a face image detection model using GFDNet, which includes an adaptive frequency perception module and a detection network, such as Figure 2 shown.

[0076] The specific operations of step B are:

[0077] Step B01: Input a facial image into a local GAN ​​to generate a facial image detection model, and use an adaptive frequency perception module to filter the facial image to obtain a frequency perception image component corresponding to the facial image.

[0078] Step B02: Input the frequency perception image component into the detection network, use the detection network to extract features of the frequency perception image component and judge the local GAN ​​generated face image, and obtain the local GAN ​​generated face image detection result.

[0079] In the embodiment of the present invention, the structure of the adaptive frequency sensing module is as follows: Figure 3 As shown, the adaptive frequency perception module includes a two-dimensional discrete cosine transform (DCT) submodule, a learnable filter group and a two-dimensional inverse discrete cosine transform (inverse DCT) submodule arranged in sequence; wherein the learnable filter group includes a high-pass filter and a learnable filter.

[0080] The present invention improves upon the Xception network to obtain a detection network. Specifically, the present invention removes the fourth to eleventh residual blocks in the Xception network and adds an Efficient Channel Attention (ECA) module to the first to third and twelfth residual blocks.

[0081] The structure of the detection network is as follows Figure 4 As shown, the detection network includes a first convolutional layer, a second convolutional layer, four residual blocks with similar structures, a first depth-separable convolutional layer, a second depth-separable convolutional layer, a global average pooling layer, and a fully connected layer. The first convolutional layer, the second convolutional layer, the first depth-separable convolutional layer, and the second depth-separable convolutional layer are followed by a ReLU activation layer respectively. Each residual block contains two depth-separable convolutional layers, a maximum pooling layer, and an ECA attention module. The second depth-separable convolutional layer in the first residual block is preceded by a ReLU activation layer, and each depth-separable convolutional layer in the second to fourth residual blocks is preceded by a ReLU activation layer. LU activation layer; in the residual block, a global average pooling layer, a one-dimensional convolution layer and a Sigmoid activation layer are set in the ECA attention module in sequence. The input of the ECA attention module is multiplied by the output of the last Sigmoid activation layer of the module to enhance the features at the channel level; the input of each residual block is first calculated by a convolution layer with a stride of 2, and then added to the output of the ECA attention module in the residual block to achieve downsampling; the output of the last residual block in the detection network is used as the input of the first depth-separable convolution layer, and then through the second depth-separable convolution layer, the global average pooling layer and the fully connected layer to obtain the final detection result.

[0082] In order to obtain accurate and reliable detection results, before the local GAN ​​generated face image detection model is put into use, it is necessary to first train the local GAN ​​generated face image detection model. The model training method in the present invention is as follows: Figure 5 As shown:

[0083] Step S1: Obtain a facial image library for model training. Local GAN-generated facial images are generated using GAN-based deep inpainting techniques. These local GAN-generated facial images are combined with their corresponding real-world facial images to establish a facial image library. In an embodiment of the present invention, the facial image library also includes a label value for each image, with the label values ​​corresponding to the local GAN-generated facial images and the real-world facial images being 0 and 1, respectively. In this embodiment of the present invention, all facial images in the facial image library are square images.

[0084] The present invention divides facial images in a facial image library into training images and test images according to a preset ratio. Both the training images and the test images contain local GAN-generated facial images and real facial images.

[0085] Step S2: Initialize network parameters. In this embodiment of the present invention, the learning rate α is set to 5.0. -4 , set the weight value to be adjusted once every batchsize=32 training samples.

[0086] Step S3: Input the training image into the local GAN-generated face image detection model to obtain the detection result at the current moment. The specific operations are as follows:

[0087] S31. Input the training image into the adaptive frequency perception module, and use the two-dimensional discrete cosine transform submodule in the adaptive frequency perception module to perform a two-dimensional discrete cosine transform on the input face image R to obtain the frequency spectrum F of the input face image R. The formula is as follows:

[0088]

[0089] Among them, F(u,v) represents the spectral coefficient of row u and column v of the spectrum F after two-dimensional discrete cosine transform, c(u) and c(v) represent the compensation coefficients, R(i,j) represents the pixel value of row i and column j of the input face image R, and the value range of i, j, u and v is 0 to M-1, where M is the side length of the face image.

[0090] The calculation formulas for c(u) and c(v) are as follows:

[0091]

[0092] Step S32: Input the frequency spectrum F of the input face image R into the learnable filter bank, and obtain the frequency component F′ after processing by the high-pass filter and the learnable filter.

[0093] In the learnable filter bank, the mask matrix H of the high-pass filter is calculated as follows:

[0094]

[0095] Among them, H(u,v) represents the weight value of the u-th row and v-th column in the mask matrix H of the high-pass filter.

[0096] In the learnable filter bank, the mask matrix K of the learnable filter is calculated as follows:

[0097] K(u,v)=σ(w u,v ) (4)

[0098] Among them, K(u,v) represents the weight value of the u-th row and v-th column in the mask matrix K of the learnable filter, w u,v is the learnable weight value in K(u,v), w u,v The initial value of satisfies the normal distribution with mean 0 and standard deviation 0.1. σ() is the limiting function, and the formula is as follows:

[0099]

[0100] Combining formulas (3) and (4), the mask matrix P of the learnable filter bank in the adaptive frequency perception module is:

[0101] P(u,v)=H(u,v)+K(u,v)(6)

[0102] Among them, P(u,v) represents the weight value of the u-th row and v-th column in the mask matrix P of the learnable filter group.

[0103] Multiply the mask matrix P by the spectrum F to obtain the frequency component F′, as follows:

[0104] F′=P*F (7)

[0105] Step S33: Use the two-dimensional inverse discrete cosine transform submodule to convert the frequency component F′ back to the RGB spatial domain to obtain the frequency perception image component Y of the input face image R:

[0106]

[0107] Among them, Y(i,j) represents the pixel value of row i and column j in the frequency perception image component Y, and F′(u,v) represents the spectral coefficient of row u and column v in the frequency component F′.

[0108] Step S34: Input the frequency perception image component Y into the detection network, and use the detection network to extract features. First, use the first convolutional layer to perform convolution calculation on the input image and pass it through the ReLU activation layer to obtain the output feature map of the first convolutional layer; use the second convolutional layer to perform convolution calculation on the output feature map of the first convolutional layer and pass it through the ReLU activation layer to obtain the output feature map of the second convolutional layer; use the first residual block to perform convolution calculation on the output feature map of the second convolutional layer to obtain the feature map of the first residual block; use the second residual block to perform convolution calculation on the feature map of the first residual block to obtain the feature map of the second residual block; use the third residual block to perform convolution calculation on the feature map of the second residual block to obtain the feature map of the third residual block; use the fourth residual block to perform convolution calculation on the feature map of the third residual block Perform convolution calculation to obtain the feature map of the fourth residual block; use the first depth-separable convolution layer to convolve the feature map of the fourth residual block and calculate it through the ReLU activation layer to obtain the output feature map of the first depth-separable convolution layer; use the second depth-separable convolution layer to convolve the output feature map of the first depth-separable convolution layer and calculate it through the ReLU activation layer to obtain the output feature map of the second depth-separable convolution layer; use global average pooling to pool the output feature map of the second depth-separable convolution layer to obtain the pooled feature map; use a fully connected layer of size 2×2048 to perform full connection calculation on the pooled feature map to obtain a one-dimensional tensor T containing two values:

[0109] T=[t0,t1] (9)

[0110] Among them, t0 and t1 are two values ​​output by the detection network, and the subscripts of t0 and t1 are 0 and 1 respectively. The subscript of the largest value of the two values ​​of the one-dimensional tensor T is selected as the output Y of the network i , if Y i = 0, indicating that the input face image R is a local GAN ​​generated face image, if Y i =1, indicating that the input face image R is a real face image.

[0111] Step S4: Calculate the cross entropy loss based on the detection result output from step 34 and the label value of the input face image R in the local GAN-generated face image library. The calculation formula is as follows:

[0112]

[0113] Among them, L cross Represents the cross entropy loss value between the detection result output by the local GAN-generated face image detection model and the true label value at the current moment, y R is the true label value of the Rth face image, p RThe detection result of the Rth face image output by the local GAN-generated face image detection model at the current moment, where n is the total number of face images.

[0114] S5, when L cross When L is greater than or equal to the preset threshold a, the network parameters of the local GAN ​​generated face image detection model are adjusted, and the model is trained again after returning to step S3. cross When it is less than the threshold a, the model training is completed and the model performance is tested using the test image.

[0115] In step B of the present invention, the trained local GAN-generated face image detection model is used to detect the face image. The specific detection steps are the same as step S3 in the embodiment of the present invention, and the local GAN-generated face image detection result is finally output.

[0116] In order to verify the effectiveness of the method of the present invention, the embodiment of the present invention uses the public celebA-HQ dataset and four GAN-based deep restoration techniques to construct a variety of local GAN-generated face images, and then establish a face image library for the experiment. The four GAN-based deep restoration techniques are Free-form image inpainting with gated convolution (GC), Generative image inpainting with contextual attention (CA), Deep fusion network for image completion (DF), and Recurrent feature reasoning for image inpainting (RFR). The images generated by the four deep restoration methods are as follows: Figure 6 As shown in the figure, each group of images is, from left to right, a real face image, a binary image mask, a face image with part of the area missing, and a face image generated by a local GAN. The original image is a real face image in celebA-HQ, and the white area in the binary image mask is the area deducted from the real face image. Then, the four deep restoration techniques mentioned above are used to repair the face image with the missing area to obtain the face image generated by the local GAN.

[0117] To verify that the method of the present invention achieves excellent detection accuracy on the in-library test set and also achieves good detection accuracy on multiple cross-library test sets, this embodiment of the present invention selects five advanced algorithms as control models: Xception (Xception: Deep learning with depthwise separable convolutions), EfficientNet_b0 (Efficientnet: Rethinking model scaling for convolutional neural networks), GramNet (Global texture enhancement for fake face detection in the wild), DCFDNet (A robust GAN-generated face detection method based on dual-color spaces and an improved Xception), and LGFDNet (Locally GAN-generated face detection based on an improved Xception). To ensure fairness, the five models are trained using the same training images. Among them, EfficientNet_b0 follows the instructions in its paper and resizes the training images to 224×224 before input. The detection accuracy (%) of Xception, EfficientNet_b0, GramNet, DCFDNet, LGFDNet and the method of the present invention on the in-library test set and the cross-library test set is shown in Table 1. It can be seen from the data in the table that the method of the present invention has achieved the best detection accuracy on both the in-library test set and the cross-library test set.

[0118] Table 1

[0119]

[0120] Furthermore, to demonstrate the robustness of the method in the face of Gaussian noise attacks, the embodiment of the present invention adds Gaussian noise to the in-library test set (GC), considering a standard deviation of 0.5-2.5 and an interval of 0.5. Table 2 shows the detection accuracy (%) of Xception, EfficientNet_b0, GramNet, DCFDNet, LGFDNet, and the method of the present invention under the same Gaussian noise intensity. As can be seen from the data in the table, the present invention can still maintain good performance in the face of Gaussian noise attacks.

[0121] Table 2

[0122]

[0123] Example 2:

[0124] Based on the same inventive concept as Example 1, this example introduces an adaptive frequency-aware local GAN-generated facial image detection device, which primarily includes an image acquisition module and an image detection module. The image acquisition module is used to acquire a facial image; the image detection module is used to detect the facial image using a trained local GAN-generated facial image detection model, thereby obtaining a local GAN-generated facial image detection result.

[0125] In an embodiment of the present invention, the local GAN-generated face image detection model includes an adaptive frequency perception module and a detection network.

[0126] The adaptive frequency perception module includes a two-dimensional discrete cosine transform submodule, a learnable filter bank and a two-dimensional inverse discrete cosine transform submodule arranged in sequence, wherein the learnable filter bank includes a high-pass filter and a learnable filter.

[0127] The detection network is obtained by deleting the fourth to eleventh residual blocks in the Xception network and adding ECA attention modules to the first to third and twelfth residual blocks. The detection network includes the first convolutional layer, the second convolutional layer, four residual blocks with similar structures, the first depth-separable convolutional layer, the second depth-separable convolutional layer, a global average pooling layer and a fully connected layer. Among them, the first convolutional layer, the second convolutional layer, the first depth-separable convolutional layer and the second depth-separable convolutional layer are respectively followed by a ReLU activation layer. Each residual block contains two depth-separable convolutional layers, a maximum pooling layer and an ECA attention module. The second depth-separable convolutional layer in the first residual block is preceded by a ReLU activation layer, and each depth-separable convolution layer in the second to fourth residual blocks is preceded by a ReLU activation layer. U activation layer; in the residual block, a global average pooling layer, a one-dimensional convolution layer and a Sigmoid activation layer are set in the ECA attention module in sequence. The input of the ECA attention module is multiplied by the output of the last Sigmoid activation layer of the module to enhance the features at the channel level; the input of each residual block is first calculated by a convolution layer with a stride of 2, and then added to the output of the ECA attention module in the residual block to achieve downsampling; the output of the last residual block in the detection network is used as the input of the first depth-separable convolution layer, and then through the second depth-separable convolution layer, the global average pooling layer and the fully connected layer to obtain the final detection result.

[0128] The image detection module uses an adaptive frequency perception module to extract a frequency perception image component of the face image from the face image; uses a detection network to perform feature extraction and local GAN-generated face image judgment on the frequency perception image component of the face image, and obtains a local GAN-generated face image detection result.

[0129] The local GAN-generated face image detection device of the present invention further includes a model training module for training the local GAN-generated face image detection model of the present invention. The operation of the model training module is the same as steps S1-S5 in Example 1.

[0130] According to Examples 1 and 2, the present invention solves the problem of insufficient generalization of existing local GAN-generated face image detection methods. While ensuring excellent detection performance on the in-library test set, it also has good generalization on the cross-library dataset, achieving accurate and reliable local GAN-generated face image detection effects. At the same time, the present invention can still maintain the detection accuracy of the material number under low-intensity Gaussian noise interference, and has good robustness.

[0131] The present invention mainly optimizes the model based on the characteristics of local GAN ​​generated facial images to improve the accuracy of the local GAN ​​generated facial image detection model. Specifically: to address the defect that the upsampling operation in GAN will cause high-frequency artifacts in the generated image, the present invention uses an adaptive frequency perception module to adaptively filter out irrelevant frequency information in the image, thereby highlighting the high-frequency artifacts in the local GAN ​​generated facial images; in order to better capture the high-frequency artifacts in the local GAN ​​generated facial images, the present invention deletes the fourth to eleventh residual blocks in the Xception network, shallows the network to make the network pay more attention to local high-frequency texture information, and adds ECA attention modules to the first to third and twelfth residual blocks to make the network pay more attention to important features.

[0132] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. An adaptive frequency-aware local GAN-generated face image detection method, characterized in that: The steps include: Get face image; Detecting the face image using the trained local GAN-generated face image detection model to obtain a local GAN-generated face image detection result; The local GAN ​​generated face image detection model includes an adaptive frequency perception module and a detection network; wherein, the adaptive frequency perception module includes a two-dimensional discrete cosine transform submodule, a learnable filter group and a two-dimensional discrete cosine inverse transform submodule arranged in sequence, and the learnable filter group includes a high-pass filter and a learnable filter; the detection network is obtained by deleting the fourth to eleventh residual blocks in the Xception network and adding an ECA attention module to the first to third and twelfth residual blocks.

2. The local GAN ​​generated face image detection method according to claim 1, characterized in that The face image is detected using the trained local GAN ​​generated face image detection model to obtain a local GAN ​​generated face image detection result, including: extracting a frequency-aware image component of the facial image from the facial image using the adaptive frequency-aware module; The detection network is used to perform feature extraction on the frequency perception image component of the facial image and local GAN ​​generated facial image judgment to obtain a local GAN ​​generated facial image detection result.

3. The local GAN ​​generated face image detection method according to claim 2, characterized in that Extracting a frequency-aware image component of the facial image from the facial image using the adaptive frequency-aware module includes: The two-dimensional discrete cosine transform submodule in the adaptive frequency perception module is used to perform a two-dimensional discrete cosine transform on the face image to obtain a frequency spectrum F of the face image: Where F(u,v) represents the spectral coefficient of row u and column v of the spectrum F after two-dimensional discrete cosine transform, c(u) and c(v) represent the compensation coefficients, R(i,j) represents the pixel value of row i and column j of the face image, and the values ​​of i, j, u, and v all range from 0 to M-1, where M is the side length of the face image. The spectrum F is filtered using the learnable filter bank in the adaptive frequency perception module to obtain the frequency component F′: F′=P*F P(u,v)=H(u,v)+K(u,v) K(u,v)=σ(w u,v ) Where P is the mask matrix P of the learnable filter group, P(u,v) represents the weight value of the u-th row and the v-th column in the mask matrix P of the learnable filter group, H(u,v) represents the weight value of the u-th row and the v-th column in the mask matrix H of the high-pass filter, K(u,v) represents the weight value of the u-th row and the v-th column in the mask matrix K of the learnable filter, w u,v represents the learnable weight value in K(u,v), and σ() is the restriction function; The frequency component F′ is converted back to the RGB spatial domain using the two-dimensional inverse discrete cosine transform submodule in the adaptive frequency perception module to obtain the frequency perception image component Y of the face image: Among them, Y(i,j) represents the pixel value of row i and column j in the frequency perception image component Y, and F′(u,v) represents the spectral coefficient of row u and column v in the frequency component F′.

4. The local GAN ​​generated face image detection method according to claim 1, characterized in that The detection network includes a first convolutional layer, a second convolutional layer, first to fourth residual blocks, a first depth-separable convolutional layer, a second depth-separable convolutional layer, a global average pooling layer and a fully connected layer, which are arranged in sequence; wherein, the first convolutional layer, the second convolutional layer, the first depth-separable convolutional layer and the second depth-separable convolutional layer are respectively followed by a ReLU activation layer; each residual block contains two depth-separable convolutional layers, a maximum pooling layer and an ECA attention module, and the ECA attention module is sequentially arranged with a global average pooling layer, a one-dimensional convolutional layer and a Sigmoid activation layer.

5. The local GAN ​​generated face image detection method according to claim 3 or 4, characterized in that: Using the first convolutional layer, the second convolutional layer, the first to fourth residual blocks, the first depth-separable convolutional layer, and the second depth-separable convolutional layer in the detection network to sequentially perform convolution calculations on the frequency-perceived image components to obtain a feature map; using the global average pooling in the detection network to pool the feature map output by the second depth-separable convolutional layer to obtain a pooled feature map; A fully connected layer in the detection network is used to perform a fully connected calculation on the pooled feature map to obtain a one-dimensional tensor containing two numerical values, where the subscripts of the two numerical values ​​are 0 and 1, respectively; the two numerical values ​​are compared, and the subscript of the larger numerical value is used as the output of the detection network; when the detection network output is 0, it indicates that the facial image is a facial image generated by a local GAN, and when the detection network output is 1, it indicates that the facial image is a real facial image.

6. The local GAN ​​generated face image detection method according to claim 1, characterized in that The training process of the local GAN-generated face image detection model is as follows: (1) obtaining a face image library, wherein the face image library includes local GAN-generated face images and their label values, and real face images and their label values; (2) Divide the face images in the face image library into training images and test images; (3) Initialize the network parameters of the local GAN-generated face image detection model; (4) Inputting the training image into the local GAN ​​generated face image detection model to obtain the local GAN ​​generated face image detection result corresponding to the training image; (5) Generate face image detection results and label values ​​of training images based on the local GAN ​​of the training image, and calculate the cross entropy loss; (6) When the cross entropy loss is lower than the threshold, the training is completed and the trained local GAN ​​generated face image detection model is obtained, otherwise return to step (4).

7. The local GAN ​​generated face image detection method according to claim 6, characterized in that The calculation formula of the cross entropy loss is as follows: Among them, L cross Indicates the cross entropy loss value between the detection result and the label value output by the local GAN ​​generated face image detection model at the current moment, y R is the label value of the Rth face image, p R The detection result of the Rth face image output by the local GAN-generated face image detection model at the current moment, where n is the total number of face images.

8. An adaptive frequency-aware local GAN-generated face image detection device, characterized in that: include: An image acquisition module, used to acquire a face image; An image detection module is used to detect the face image using the trained local GAN-generated face image detection model to obtain a local GAN-generated face image detection result; In the image detection module, the local GAN-generated face image detection model includes an adaptive frequency perception module and a detection network; wherein the adaptive frequency perception module includes a two-dimensional discrete cosine transform submodule, a learnable filter group and a two-dimensional discrete cosine inverse transform submodule arranged in sequence, and the learnable filter group includes a high-pass filter and a learnable filter; the detection network is obtained by deleting the fourth to eleventh residual blocks in the Xception network and adding an ECA attention module to the first to third and twelfth residual blocks.

9. The local GAN ​​generated face image detection device according to claim 8, characterized in that The image detection module is specifically used to: extracting a frequency-aware image component of the facial image from the facial image using the adaptive frequency-aware module; The detection network is used to perform feature extraction on the frequency perception image component of the facial image and local GAN ​​generated facial image judgment to obtain a local GAN ​​generated facial image detection result.

10. The local GAN ​​generated face image detection device according to claim 8, characterized in that The image detection module extracts the frequency perception image component of the face image from the face image using the adaptive frequency perception module, including: The two-dimensional discrete cosine transform submodule in the adaptive frequency perception module is used to perform a two-dimensional discrete cosine transform on the face image to obtain a frequency spectrum F of the face image: Where F(u,v) represents the spectral coefficient of row u and column v of the spectrum F after two-dimensional discrete cosine transform, c(u) and c(v) represent the compensation coefficients, R(i,j) represents the pixel value of row i and column j of the face image, and the values ​​of i, j, u, and v all range from 0 to M-1, where M is the side length of the face image. The spectrum F is filtered using the learnable filter in the adaptive frequency perception module to obtain the frequency component F′: F′=P*F P(u,v)=H(u,v)+K(u,v) K(u,v)=σ(w u,v ) Where P is the mask matrix P of the learnable filter group, P(u,v) represents the weight value of the u-th row and the v-th column in the mask matrix P of the learnable filter group, H(u,v) represents the weight value of the u-th row and the v-th column in the mask matrix H of the high-pass filter, K(u,v) represents the weight value of the u-th row and the v-th column in the mask matrix K of the learnable filter, w u,v represents the learnable weight value in K(u,v), and σ() is the restriction function; The frequency component F′ is converted back to the RGB spatial domain using the two-dimensional inverse discrete cosine transform submodule in the adaptive frequency perception module to obtain the frequency perception image component Y of the face image: Among them, Y(i,j) represents the pixel value of row i and column j in the frequency perception image component Y, and F′(u,v) represents the spectral coefficient of row u and column v in the frequency component F′.

Citation Information

Patent Citations

  • Novel HOG feature Uyghur face image recognition algorithm

    CN108038464A

  • Face counterfeit image detection method and device, terminal and storage medium

    CN115601820A