A Face Liveness Detection Method for Image Presentation Attacks

By generating an adversarial network GAN model and a deep anomaly detection learning framework for multi-frequency domain features, the problems of low accuracy and high error detection rate of image demonstration attacks in face live detection are solved, and efficient and accurate live detection are achieved.

CN116824664BActive Publication Date: 2025-07-29BEIJING SHISHI MANAGEMENT CONSULTING CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310700231.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-13
Publication Date
2025-07-29
Estimated Expiration
2043-06-13

AI Technical Summary

Technical Problem

When facing image demonstration attacks, the existing face live detection methods have problems such as low detection accuracy and high error detection rate. In particular, the GAN-based methods are prone to collapse during the training process and the reconstruction of abnormal data images is more realistic, resulting in unstable live detection results.

Method used

A deep anomaly detection learning framework based on the generative adversarial network GAN model and multi-frequency domain features is adopted, including three independent generator networks and discriminator networks. By pre-processing and feature learning of the image, combining the prediction scores generated by the abnormal discriminator network with the threshold comparison, we determine whether the image is a real living face.

Benefits of technology

It improves the accuracy and efficiency of facial live detection, reduces the amount of calculation, can effectively identify various demonstration attack forms, reduces the error detection rate and optimization time, and improves the accuracy of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116824664B_ABST
    Figure CN116824664B_ABST
Patent Text Reader

Abstract

The present invention discloses a face liveness detection method for image presentation attacks, including: building a deep anomaly detection learning framework, where the deep anomaly detection learning framework includes three independent generator networks and discriminator networks respectively connected to the generator networks; preprocessing the original image to be detected to obtain a first preprocessed image and a second preprocessed image of the original image to be detected; respectively inputting the original image to be detected, the first preprocessed image, and the second preprocessed image into three independent generator networks in the trained deep anomaly detection learning framework for feature learning, image reconstruction, and anomaly discrimination, obtaining latent vectors and reconstructed images generated by the generator networks and prediction scores generated by the discriminator networks; comparing the prediction scores generated by the discriminator network with an anomaly score threshold. The present invention solves the problem of difficult feature extraction in traditional anomaly detection and greatly improves the accuracy of face liveness detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image anomaly detection, and particularly relates to a face liveness detection method for image presentation attacks. Background Art

[0002] With the continuous development of computer technology and the continuous improvement of the informatization level, face recognition technology has been widely applied in people's daily lives, such as in access control systems, encryption of electronic devices such as mobile phones and computers. In the task of identity recognition, face recognition has the advantages of convenience, low cost, and imperceptibility compared with biometric recognition such as iris recognition and fingerprint recognition, and has become one of the most widely used biometric systems. However, the face recognition system has potential security risks and is relatively vulnerable to attacks. In order to prevent criminals from attacking the face recognition system, it is necessary to perform face liveness detection to ensure that the first is real face information rather than photo or video replay.

[0003] Saket Sathe, Jinghui Chen et al. proposed an improved method RandNet based on AE (AutoEncoder) in "Jinghui Chen, Saket Sathe, Charu Aggarwal, and Deepak Turaga. 2017. Outlier detection with autoencoder ensembles. In SDM. 90–98.", which trained a group of independent AEs, and each AE had some randomly selected constant dropout connections. An adaptive sampling strategy was used by exponentially increasing the sample size of the mini-batch, thereby further enhancing the performance of the AE in the anomaly detection task. Although the autoencoder is a simple and effective anomaly detection architecture, this method does not perform well in the actual face liveness detection task because the noise of the training data images will degrade the performance of the model, resulting in relatively average face liveness detection results.

[0004] The literature "Goodfellow, Ian, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. "Generative adversarial networks." Communications of the ACM 63, no. 11 (2020): 139-144." records an anomaly detection based on GAN, which quickly became a popular deep anomaly detection method after its early use. This method usually uses a generator G to learn the latent feature space so that the latent space can well capture the normal features behind the given data, and then defines a certain form of residual between the real instance and the generated instance as the anomaly score. This method aims to learn a specific feature representation and then optimize specifically for this feature representation. One-class classification is called the problem of learning a description of a set of data instances to detect whether a new instance conforms to the training data. In practical applications, the training process of this method is prone to collapse, and when the learning ability of the generator is too strong, the reconstruction of abnormal data images can also be relatively real, thus deceiving the discriminator and resulting in a high false detection rate in the live detection results. Therefore, this method framework has disadvantages such as unstable representation ability and high training difficulty for the recognition and classification of face images. Summary of the Invention

[0005] In order to solve the above problems existing in the prior art, the present invention provides a face liveness detection method for image presentation attacks. The technical problems to be solved by the present invention are realized through the following technical solutions:

[0006] The present invention provides a face liveness detection method for image presentation attacks, including:

[0007] S1: Build a deep anomaly detection learning framework based on the generative adversarial network GAN model and multi-frequency domain features. The deep anomaly detection learning framework includes three independent generator networks and a discriminator network respectively connected to the generator networks;

[0008] S2: Preprocess the original image to be detected to obtain a first preprocessed image and a second preprocessed image of the original image to be detected;

[0009] S3: Input the original image to be detected, the first preprocessed image, and the second preprocessed image into three independent generator networks in the trained deep anomaly detection learning framework respectively for feature learning, image reconstruction, and anomaly discrimination, and obtain the latent vectors and reconstructed images generated by the generator networks and the prediction scores generated by the discriminator network;

[0010] S4: Compare the prediction scores generated by the discriminator network with the anomaly score threshold to determine whether the original image to be detected is a genuine face liveness image.

[0011] In one embodiment of the present invention, the generator network includes a cascaded first encoder and a first decoder, wherein,

[0012] The first encoder includes a first convolutional layer, a first batch normalization layer, a first activation layer, a second convolutional layer, a second batch normalization layer, a second activation layer, a third convolutional layer, a third batch normalization layer, and a third activation layer connected in sequence;

[0013] The first decoder includes a first deconvolutional layer, a fourth batch normalization layer, a fourth activation layer, a second deconvolutional layer, a fifth batch normalization layer, a fifth activation layer, and a third deconvolutional layer connected in sequence.

[0014] In one embodiment of the present invention, the discriminator network includes a fourth convolutional layer, a sixth activation layer, a first Dropout layer, a fifth convolutional layer, a seventh activation layer, a second Dropout layer, a fully connected layer, and a Sigmod layer connected in sequence.

[0015] In one embodiment of the present invention, a second encoder is further included behind the generator network. The input of the second encoder is the image generated by the generator network. The second encoder includes a sixth convolutional layer, a sixth batch normalization layer, an eighth activation layer, a seventh convolutional layer, a seventh batch normalization layer, a ninth activation layer, an eighth convolutional layer, an eighth batch normalization layer, and a tenth activation layer connected in sequence.

[0016] In one embodiment of the present invention, the S2 includes:

[0017] Obtain the original image to be detected, and perform object detection and key point localization of the face on the original image to be detected to reduce image noise and enhance image quality;

[0018] Perform alignment and central cropping operations on the image to be detected after face object detection, and only retain the face area of the image;

[0019] Perform LBP and bilateral filtering processing on the cropped face area respectively to obtain a first preprocessed image and a second preprocessed image.

[0020] In one embodiment of the present invention, before step S3, it further includes:

[0021] Train the deep anomaly detection learning framework using a training data set composed of a large number of normal face liveness images to obtain a trained deep anomaly detection learning framework.

[0022] In one embodiment of the present invention, training the deep anomaly detection learning framework by using a training data set composed of a large number of normal face live images includes:

[0023] Select a training data set composed of a large number of normal face live images, and perform LBP filtering and bilateral filtering on each training image in the training data set to obtain the first filtered image and the second filtered image of each training image to form an image group including three training images;

[0024] Input each group of image groups into three independent generator networks of the deep anomaly detection learning framework batch by batch, and respectively obtain the latent vector and the reconstructed image by the generator network, obtain the prediction score by the discriminator network, and obtain the high-dimensional feature vector

[0025] extracted for the generated image by the second encoder;

[0026] Define the objective function during training as: enc L = L con + L adv ,

[0027]

[0028]

[0029]

[0030] wherein, L enc represents the model reconstruction loss, which is the second norm between the latent space vectors obtained by the generator network and the encoder network, L con represents the context loss, which is the element-wise first norm between the original image and the generated image obtained by the generator network, and L adv represents the adversarial training loss, which is the prediction scores given by the discriminator for the original image and the generated image;

[0031] Set the anomaly scoring function during training Use the anomaly scoring function as the standard for the model to input the anomaly score during the training phase, and adaptively learn the anomaly score threshold;

[0032] Compare the scores output by the deep anomaly detection learning framework with the anomaly score threshold obtained through adaptive learning. If it is greater than the threshold, it is determined as an abnormal image; otherwise, it is determined as a normal image. Subsequently, compare it with the label information of the current training image. If the two are the same, it indicates that the training network has the best learning effect; otherwise, continuously optimize the network parameters of the deep anomaly detection learning framework according to the objective function to enhance the learning effect of the training network.

[0033] In an embodiment of the present invention, step S3 further includes:

[0034] Use a test data set composed of a large number of normal face liveness images and abnormal images to test the trained deep anomaly detection learning framework to determine the effectiveness of the trained deep anomaly detection learning framework.

[0035] In an embodiment of the present invention, using a test data set composed of a large number of normal face liveness images and abnormal images to test the trained deep anomaly detection learning framework includes:

[0036] Define the anomaly scoring function B(x) during the test process:

[0037]

[0038] Wherein, w1 and w2 respectively represent weight parameters;

[0039] Obtain a large number of normal face liveness images and face images of abnormal presentation attacks, and construct a test data set;

[0040] Input the test images in the test data set into the trained deep anomaly detection learning framework, and respectively obtain different latent space vectors and generated images through the generator network and the discriminator network, and obtain scores through the discriminator network;

[0041] Calculate the anomaly score of the input image according to the anomaly scoring function, and compare the calculated anomaly score with the adaptively generated threshold: if the calculated anomaly score is greater than the threshold, it is determined as an abnormal image; otherwise, it is determined as a normal image.

[0042] Compared with the prior art, the beneficial effects of the present invention are:

[0043] 1. The present invention provides a detection method for demonstration attacks in a face recognition system. Starting from the concept of anomaly detection, the face liveness detection problem is regarded as a binary classification problem. At the same time, in response to problems such as difficulty in obtaining abnormal samples in real application scenarios, a detection method improved and designed based on the generative adversarial network (GAN) model is proposed to improve performance such as detection efficiency and accuracy.

[0044] 2. The anomaly detection network framework of the present invention based on the GAN model overcomes both the problem of insufficient representation ability of manual features and the problem of difficult feature extraction in traditional anomaly detection, greatly improving the accuracy of face liveness detection. In anomaly detection, the classification method combined with the final determination method greatly reduces the computational amount compared with traditional classification methods such as SVM, improving the efficiency of face liveness detection.

[0045] 3. The present invention targets demonstration attack methods and takes into account various demonstration attack forms, such as printing attacks, electronic screen demonstration attacks, and other attack situations. It overcomes the problems of poor performance of traditional face liveness detection algorithms in the face of diverse attack types, such as long optimization time and low accuracy caused by optimizing all inputs with the same set of prediction functions.

[0046] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings

[0047] Figure 1 is a schematic flowchart of a face liveness detection method for image demonstration attacks provided by an embodiment of the present invention;

[0048] Figure 2 is a schematic diagram of the processing process of a deep anomaly detection learning framework provided by an embodiment of the present invention;

[0049] Figure 3 is a schematic diagram of an image preprocessing process provided by an embodiment of the present invention;

[0050] Figure 4 is a comparison schematic diagram of an original image and a preprocessed image provided by an embodiment of the present invention. Detailed Embodiments

[0051] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the intended invention purpose, the following provides a detailed description of a face liveness detection method for image demonstration attacks proposed according to the present invention with reference to the accompanying drawings and specific embodiments.

[0052] The foregoing and other technical contents, features and effects of the present invention will be clearly presented in the following detailed description of the specific embodiments in conjunction with the accompanying drawings. Through the description of the specific embodiments, a more in-depth and specific understanding of the technical means and effects adopted by the present invention to achieve the predetermined purpose can be obtained. However, the accompanying drawings are only for reference and illustration, and are not used to limit the technical solutions of the present invention.

[0053] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant are intended to cover non-exclusive inclusion, so that an article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the article or device including the said element.

[0054] Based on the generative adversarial network GAN model and multi-frequency domain features, the embodiment of the present invention constructs a deep anomaly detection learning framework; the training image and the filtered image are respectively input into the deep learning network after the preprocessing process, and the network parameters are continuously optimized in combination with the label information until the learning effect is the best, so as to achieve fast and accurate detection of whether the face image is a live body.

[0055] Specifically, please refer to Figure 1 , Figure 1 which is a schematic flow chart of a face liveness detection method for image presentation attacks provided by an embodiment of the present invention. The face liveness detection method includes:

[0056] S1: Construct a deep anomaly detection learning framework based on the generative adversarial network GAN model and multi-frequency domain features. Please refer to Figure 2 , Figure 2 which is a schematic diagram of the processing process of a deep anomaly detection learning framework provided by an embodiment of the present invention. The deep anomaly detection learning framework includes three independent generator networks and discriminator networks respectively connected to the generator networks.

[0057] Specifically, the generator network includes a cascaded first encoder and first decoder. Among them, the first encoder is used to map the input image to a vector in the latent space, and the first decoder is used to map the latent vector back to the image space.

[0058] The first encoder in this embodiment consists of multiple convolutional layers, which map an input image of size (H, W, C) to a 100-dimensional vector in the latent space. Specifically, it includes a first convolutional layer, a first batch normalization layer, a first activation layer, a second convolutional layer, a second batch normalization layer, a second activation layer, a third convolutional layer, a third batch normalization layer, and a third activation layer that are connected in sequence. Each convolutional layer is used to perform a convolutional operation on the output result of the previous layer, enhance and extract data features, and reduce noise. Each batch normalization layer is used to normalize the output of the convolutional layer. Each activation layer in the first encoder uses the LeakReLU function as the activation function to perform a non-linear transformation on the output of the batch normalization layer.

[0059] The first decoder consists of multiple transposed convolutional layers, which map the latent vector back to the image space and output a reconstructed image of size (H, W, C). Specifically, it includes a first transposed convolutional layer, a fourth batch normalization layer, a fourth activation layer, a second transposed convolutional layer, a fifth batch normalization layer, a fifth activation layer, and a third transposed convolutional layer that are connected in sequence. Each transposed convolutional layer in the first decoder upsamples the latent vector obtained in the first encoder step by step, and finally outputs an image of the same size as the original input image in the third transposed convolutional layer. Each batch normalization layer is used to normalize the output of the transposed convolutional layer. Each activation layer in the first decoder uses the LeakReLU function as the activation function to perform a non-linear transformation on the output of the batch normalization layer.

[0060] Furthermore, the deep anomaly detection learning framework of this embodiment includes a discriminator network, which takes the outputs of the three generator networks as inputs at the same time. It uses a structure similar to that of the encoder to downsample the images output by the generator networks step by step, and finally obtains a score to evaluate the authenticity of the generated images. The discriminator network includes a fourth convolutional layer, a sixth activation layer, a first Dropout layer, a fifth convolutional layer, a seventh activation layer, a second Dropout layer, a fully connected layer, and a Sigmod layer that are connected in sequence. Each convolutional layer of the discriminator network is used for feature extraction. The activation layer also uses the LeakReLU function as the activation function to perform a non-linear transformation on the output of the convolutional layer. Each Dropout layer discards some neurons with a certain probability to prevent the model network from overfitting. The fully connected layer maps the outputs of the previous layers to a real number, representing the probability that the input image is a real image. The Sigmod layer compresses the output of the fully connected layer to between 0 and 1.

[0061] In addition, a second encoder is further included behind the generator network, that is, a second encoder is included between the generator network and the discriminator network. The input of the second encoder is the image generated by the generator network, and the output is a latent space vector of 100 dimensions. The structure of the second encoder is the same as that of the first encoder in the generator network. Specifically, the second encoder includes a sixth convolutional layer, a sixth batch normalization layer, an eighth activation layer, a seventh convolutional layer, a seventh batch normalization layer, a ninth activation layer, an eighth convolutional layer, an eighth batch normalization layer, and a tenth activation layer connected in sequence. Each convolutional layer in the second encoder is used to perform a convolutional operation on the output result of the previous layer respectively, enhance and extract data features, and reduce noise; the batch normalization layer is used to perform normalization processing on the output of the convolutional layer; the activation layer in the second encoder uses the LeakReLU function as the activation function to perform a non-linear transformation on the output of the batch normalization layer.

[0062] It should be noted that during the actual training or use process, the loss value of each image is calculated according to the set loss function. The smaller the loss value, the better the learning effect of the training network; otherwise, the network parameters need to be fine-tuned multiple times to enhance the learning effect of the training network. The specific training process will be described in detail below.

[0063] S2: Preprocess the original image to be detected to obtain the first preprocessed image and the second preprocessed image of the original image to be detected.

[0064] Please refer to Figure 3 , Figure 3 which is a schematic diagram of an image preprocessing process provided by an embodiment of the present invention. Step S2 specifically includes:

[0065] S2.1: Obtain the original image to be detected, and use the RetinaFace method to perform object detection and key point localization of the face on the original image to be detected, reduce image noise, and enhance image quality;

[0066] S2.2: Perform alignment and central cropping operations on the image to be detected after face object detection, reduce the interference of irrelevant factors in the image, and only retain the area of the face part to shorten the running time and improve efficiency;

[0067] S2.3: Perform LBP and bilateral filtering processing on the cropped face area respectively to obtain the first preprocessed image and the second preprocessed image. Please refer to Figure 4 , Figure 4 which is a comparison schematic diagram of the original image and the preprocessed image provided by an embodiment of the present invention. Among them, (a) is the original image, (b) is the image after bilateral filtering, that is, the first preprocessed image, and (c) is the image after LBP filtering, that is, the second preprocessed image.

[0068] S3: Input the to-be-detected original image, the first preprocessed image, and the second preprocessed image into three independent generator networks in the trained deep anomaly detection learning framework respectively for feature learning, image reconstruction, and anomaly discrimination, to obtain the latent vectors and reconstructed images generated by the generator networks and the prediction scores generated by the discriminator network.

[0069] In this embodiment, before step S3, it further includes: training the deep anomaly detection learning framework with a training data set composed of a large number of normal face liveness images to obtain a trained deep anomaly detection learning framework. That is to say, before using the constructed deep anomaly detection learning framework to detect the to-be-detected original image, it is necessary to train the deep anomaly detection learning framework with the training data set to continuously improve the accuracy of its detection, and finally obtain a trained deep anomaly detection learning framework that meets the requirements. The specific training process is as follows:

[0070] Select a training data set composed of a large number of normal face liveness images. In this embodiment, it is set that the training data set χ = {x1, x2,..., x n ,..., x N} composed of N training images, where x n is the nth training image in the training data set. Perform LBP filtering and bilateral filtering on each training image in the training data set to obtain the first filtered image and the second filtered image of each training image, and form an image group including three training images; the labels of the images in the training data set are all set to 0, and 0 represents a normal image.

[0071] Next, as Figure 2 shown, input each group of image groups into the three independent generator networks in the deep anomaly detection learning framework batch by batch for operations such as feature learning, image reconstruction, and anomaly discrimination. Respectively obtain the latent vector and the reconstructed image from the generator network, obtain the prediction score from the discriminator network, and obtain the high-dimensional feature vector

[0072] extracted for the generated image by the second encoder.

[0073] Define the objective function in the training process as: enc + L con + L adv ,

[0074]

[0075]

[0076]

[0077] Among them, L enc represents the model reconstruction loss, which is the two-norm between the latent space vectors obtained by the generator network and the encoder network. L con represents the context loss, which is the element-wise one-norm between the original image and the generated image obtained by the generator network. L adv represents the adversarial training loss, which is the prediction scores given by the discriminator for the original image and the generated image;

[0078] Set the anomaly scoring function during the training process Use the anomaly scoring function as the standard for the input anomaly score of the model during the training phase, and adaptively learn the anomaly score threshold.

[0079] Subsequently, compare the score output by the deep anomaly detection learning framework with the adaptively learned anomaly score threshold. If it is greater than the threshold, it is determined as an abnormal image; otherwise, it is determined as a normal image. Subsequently, compare it with the label information of the current training image. If the two are the same, it means that the training network has the best learning effect; otherwise, continuously optimize the network parameters according to the objective function to enhance the learning effect of the training network. The normal image mentioned in this embodiment is the real face image collected by the recognition camera during the face recognition process, and the abnormal image can be an abnormal image such as a printed face photo, etc.

[0080] Furthermore, after the training process is completed, step S3 further includes:

[0081] Use the test data set composed of a large number of normal face liveness images and abnormal images to test the trained deep anomaly detection learning framework to determine the effectiveness of the trained deep anomaly detection learning framework.

[0082] Specifically,

[0083] First, define the anomaly scoring function B(x) during the test process:

[0084]

[0085] Among them, w1 and w2 respectively represent weight parameters;

[0086] Obtain a large number of normal face liveness images and abnormal face images of presentation attacks, and construct a test data set;

[0087] Input the test images in the test dataset into the trained deep anomaly detection learning framework. Different latent space vectors and generated images are obtained through the generator network and the discriminator network respectively, and scores are obtained through the discriminator network.

[0088] Calculate the anomaly score of the input image according to the anomaly scoring function, and compare the calculated anomaly score with the adaptively generated threshold: if the calculated anomaly score is greater than the threshold, it is determined as an abnormal image; otherwise, it is determined as a normal image. Then, verify the effect of the trained deep anomaly detection learning framework according to the labels in the test images.

[0089] Further, after the training and testing of the deep anomaly detection learning framework are completed, input the original image to be detected, the first preprocessed image, and the second preprocessed image into three independent generator networks of the trained deep anomaly detection learning framework for feature learning, image reconstruction, and anomaly discrimination, and obtain the latent vectors and reconstructed images generated by the generator network and the prediction scores generated by the discriminator network.

[0090] S4: Compare the prediction scores generated by the discriminator network with the anomaly score threshold to determine whether the original image to be detected is a real live face image.

[0091] The effectiveness of the face liveness detection method according to the embodiments of the present invention is further verified through experiments below.

[0092] (I) Test environment:

[0093] This invention conducts experiments on a neural network model built based on Python + PyTorch with a central processing unit of Intel(R) Core(TM) i9 10900K 3.7GHz, 64G of memory, a GPU NVIDIA GeForce RTX3090 24G, and an Ubuntu22.0.4.2 LTS operating system.

[0094] (II) Test content:

[0095] Experiment 1: Use the method according to the embodiments of the present invention for face liveness detection in the Multispectral dataset. The test results are as follows in the table:

[0096] Table 1 Performances of various methods on the Multispectral dataset

[0097] F1Score AUROC Accuracy Precision Recall CFlow 0.9152 0.7308 0.8451 0.8471 <![CDATA 0.9953 > FastFlow 0.9092 <![CDATA 0.7636 > 0.8342 0.8345 0.9986 GANomaly 0.9132 0.7213 0.8404 0.8404 0.9132 STFPM 0.9152 0.7016 0.8451 0.8477 0.9943 The method of the present invention 0.9132 0.8115 0.8404 <![CDATA 0.8404 > 1

[0098] Among them, accuracy refers to the ratio of the number of samples correctly classified by the classifier to the total number of samples in a given test dataset, that is, the probability of correct prediction; precision represents the proportion of actual positive samples among the samples predicted as positive; recall represents the proportion of actual positive samples determined as positive samples; F1Score comprehensively considers both precision and recall, achieving a balance between the two, that is, it is necessary to "seek precision" and also "seek comprehensiveness"; AUROC represents the area under the curve in the Receiver Operating Characteristic image, which is suitable for measuring the situation where the number of classes in a classification task has a large deviation. As can be seen from Table 1, the AUROC index of the method of the present invention ranks first in this dataset, and this index can better reflect the quality of a classifier than other indicators. Especially in the case of an imbalance in the ratio of positive and negative samples, the ROC index of the method of the present invention reaches 0.8115, exceeding the second-place FastFlow by about 5 percentage points; in other indicators such as F1Score and Recall, the method also achieves the first or second place.

[0099] Experiment 2: In the NUAA dataset, the method of the embodiment of the present invention is used for face liveness detection, and the test results are as follows in the table:

[0100] Table 2 Performance of each method on the NUAA dataset

[0101] F1Score AUROC Accuracy Precision Recall CFlow 0.9442 0.7826 0.8943 0.8943 1 FastFlow 0.9514 0.8362 0.9086 0.9073 1 GANomaly 0.9674 0.9067 0.9398 0.9369 1 STFPM 0.9611 0.8717 0.9310 0.9686 0.9537 The method of the present invention 0.9950 0.9988 0.9910 0.9901 1

[0102] As can be seen from Table 2, the ROC of the method of the present invention reaches an astonishing 0.9988, far leading other methods, and is the best in other indicators, which is sufficient to prove its superiority.

[0103] The present invention provides a detection method for face recognition system presentation attacks. Starting from the concept of anomaly detection, the face liveness detection problem is regarded as a binary classification problem, and at the same time, problems such as difficulty in obtaining abnormal samples faced in real application scenarios are addressed. It is a detection method improved and designed based on the generative adversarial network GAN model, which improves performance such as detection efficiency and accuracy.

[0104] The anomaly detection network framework of the present invention based on the generative adversarial network (GAN) model overcomes both the problem of insufficient representation ability of manual features and the problem of difficult feature extraction in traditional anomaly detection, greatly improving the accuracy of face liveness detection. In anomaly detection, the classification method combined with the final determination method significantly reduces the computational complexity compared to traditional classification methods such as SVM, improving the efficiency of face liveness detection. In addition, the present invention addresses the demonstration attack mode and takes into account various forms of demonstration attacks, such as printing attacks, electronic screen demonstration attacks, and other attack situations, overcoming the problems of poor performance of traditional face liveness detection algorithms in the face of diverse attack types, long optimization time, and low accuracy caused by optimizing all inputs with the same set of prediction functions.

[0105] In several embodiments provided by the present invention, it should be understood that the devices and methods disclosed by the present invention can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed.

[0106] In addition, in each embodiment of the present invention, the functional modules can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module. The above integrated modules can be implemented in the form of hardware or in the form of a combination of hardware and software functional modules.

[0107] Another embodiment of the present invention provides a storage medium storing a computer program for executing the steps of the face liveness detection method for image demonstration attacks described in the above embodiments. On the other hand, the present invention provides an electronic device including a memory and a processor. The memory stores a computer program, and when the processor calls the computer program in the memory, it implements the steps of the face liveness detection method for image demonstration attacks as described in the above embodiments. Specifically, the above integrated modules implemented in the form of software functional modules can be stored in a computer-readable storage medium. The above software functional modules stored in a storage medium include several instructions for causing an electronic device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute some steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, and other media that can store program codes.

[0108] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can be made, and all should be regarded as falling within the protection scope of the present invention.

Claims

1. A face liveness detection method for image presentation attacks, characterized in that, Including: S1: Build a deep anomaly detection learning framework based on a Generative Adversarial Network (GAN) model and multi-frequency domain features. The deep anomaly detection learning framework includes three independent generator networks and discriminator networks respectively connected to the generator networks; S2: Preprocess the original image to be detected to obtain a first preprocessed image and a second preprocessed image of the original image to be detected; S3: Input the original image to be detected, the first preprocessed image, and the second preprocessed image into three independent generator networks of the trained deep anomaly detection learning framework respectively for feature learning, image reconstruction, and anomaly discrimination, to obtain the latent vectors and reconstructed images generated by the generator networks and the prediction scores generated by the discriminator networks; S4: Compare the prediction scores generated by the discriminator network with an anomaly score threshold to determine whether the original image to be detected is a real live face image; Before step S3, it further includes: Use a training data set composed of a large number of normal live face images to train the deep anomaly detection learning framework to obtain a trained deep anomaly detection learning framework; Among them, using a training data set composed of a large number of normal live face images to train the deep anomaly detection learning framework includes: Select a training data set composed of a large number of normal live face images, and perform LBP filtering and bilateral filtering on each training image in the training data set to obtain the first filtered image of each training image and the second filtered image to form an image group including three training images; ​ Input each group of image groups into three independent generator networks of the deep anomaly detection learning framework according to batches, and respectively obtain latent vectors from the generator networks and reconstructed images Obtain prediction scores from the discriminator network and high-dimensional feature vectors extracted for the generated images by the second encoder Define the objective function during training as: L = L enc + L con + L adv , Among them, L enc represents the model reconstruction loss, which is the two-norm between the latent space vectors obtained by the generator network and the encoder network. L con represents the context loss, which is the element-wise one-norm between the original image and the generated image obtained by the generator network. L adv represents the adversarial training loss, which is the prediction scores given by the discriminator for the original image and the generated image; Set the anomaly scoring function during the training process Use the anomaly scoring function as the standard for the model to input anomaly scores during the training phase, and adaptively learn the anomaly score threshold; Compare the scores output by the deep anomaly detection learning framework with the anomaly score threshold obtained by adaptive learning. If it is greater than the threshold, it is determined as an abnormal image; otherwise, it is determined as a normal image. Then compare it with the label information of the current training image. If the two are the same, it means that the training network has the best learning effect; otherwise, continuously optimize the network parameters of the deep anomaly detection learning framework according to the objective function to enhance the learning effect of the training network.

2. The face liveness detection method for image presentation attack according to claim 1, characterized in that, The generator network includes a cascaded first encoder and first decoder, where The first encoder includes a first convolutional layer, a first batch normalization layer, a first activation layer, a second convolutional layer, a second batch normalization layer, a second activation layer, a third convolutional layer, a third batch normalization layer, and a third activation layer connected in sequence; The first decoder includes a first deconvolutional layer, a fourth batch normalization layer, a fourth activation layer, a second deconvolutional layer, a fifth batch normalization layer, a fifth activation layer, and a third deconvolutional layer connected in sequence.

3. The method for face liveness detection against image presentation attacks according to claim 2, characterized in that, The discriminator network includes a fourth convolutional layer, a sixth activation layer, a first Dropout layer, a fifth convolutional layer, a seventh activation layer, a second Dropout layer, a fully connected layer, and a Sigmod layer connected in sequence.

4. The face liveness detection method for image presentation attacks according to claim 1, characterized in that, Behind the generator network, there is also a second encoder. The input of the second encoder is the image generated by the generator network. The second encoder includes a sixth convolutional layer, a sixth batch normalization layer, an eighth activation layer, a seventh convolutional layer, a seventh batch normalization layer, a ninth activation layer, an eighth convolutional layer, an eighth batch normalization layer, and a tenth activation layer connected in sequence.

5. The face liveness detection method for image presentation attack according to claim 1, wherein S2 includes: Obtain the original image to be detected, and perform object detection and key point localization of the face on the original image to be detected; Perform alignment and central cropping operations on the image to be detected after face target detection, and only retain the face region of the image; Perform LBP and bilateral filtering processing on the cropped face regions respectively to obtain a first preprocessed image and a second preprocessed image.

6. The face liveness detection method for image presentation attacks according to claim 5, characterized in that Step S3 further includes: Use a test data set composed of a large number of normal face liveness images and abnormal images to test the trained deep anomaly detection learning framework to determine the effectiveness of the trained deep anomaly detection learning framework.

7. The face liveness detection method for image presentation attack according to claim 6, characterized in that, Using a test data set composed of a large number of normal face liveness images and abnormal images to test the trained deep anomaly detection learning framework includes: Define the anomaly scoring function B(x) during the test process: Among them, w1 and w2 respectively represent weight parameters; Obtain a large number of normal face liveness images and face images of abnormal presentation attacks to construct a test data set; Input the test images in the test data set into the trained deep anomaly detection learning framework, and obtain different latent space vectors and generated images through the generator network and the discriminator network respectively, and obtain scores through the discriminator network; Calculate the anomaly score of the input image according to the anomaly scoring function, and compare the calculated anomaly score with the adaptively generated threshold: if the calculated anomaly score is greater than the threshold, it is determined as an abnormal image, otherwise it is determined as a normal image.

Citation Information

Patent Citations

  • Face liveness detection method and apparatus, and electronic device

    CA3043230A1

  • Single-mode face living body detection method based on multi-mode face training

    CN113705400A