Face living body detection method and system based on double-flow convolutional neural network

By converting the face feature images into HSV and YUV color spaces, and using the VGG16 convolutional neural network model to extract features for fusion, the problem of insufficient detection capabilities in the existing technology in complex live detection environments is solved, and high accuracy and cross-data set detection capabilities are improved.

CN120198974APending Publication Date: 2025-06-24JIANGSU UNIV OF SCI & TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510081824.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The prior art is difficult to effectively detect video face swaps and AI generation of fake faces in complex live detection environments, and it lacks detection capabilities in complex data sets and cross-data sets.

Method used

The face live detection method based on the dual-stream convolutional neural network is adopted. By converting the face feature images into HSV and YUV color spaces, and using the VGG16 convolutional neural network model to extract features for fusing, and inputting the full connection layer for classification.

Benefits of technology

It improves the accuracy of face live detection, enhances the ability to detect complex data sets and cross-data sets, and effectively improves the detection ability of complex forged face images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198974A_ABST
    Figure CN120198974A_ABST
Patent Text Reader

Abstract

The invention provides a face in-vivo detection method and system based on a double-flow convolutional neural network, and the method comprises the steps: converting an original face feature image into two different color space images after preprocessing the face feature image, and constructing a training set through employing the preprocessed face feature image; constructing a human face in-vivo detection FAS model by using a double-flow convolutional neural network model framework, designing a double-flow convolutional layer, extracting image features in parallel by using two different color spaces, and classifying the double features; using the constructed training set to train a face in-vivo detection FAS model to obtain an optimal face in-vivo detection FAS model; and finally, calling the optimal human face in-vivo detection FAS model to carry out classification detection on the human face feature image data. According to the face in-vivo detection FAS model provided by the invention, the features in two color spaces are comprehensively considered, the classification credibility is increased, and extremely high accuracy is obtained on complex data set and cross-data set detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image recognition, and particularly to a face liveness detection method and system based on a two-stream convolutional neural network. Background Art

[0002] With the wide application of face recognition technology in recent years, liveness detection, as a part of the face recognition system, has gradually shown its importance. Liveness detection is a technology that uses image processing methods to detect the physiological characteristics of the tested person and determine whether the tested person actually exists. Its purpose is to eliminate fraud behaviors that may be disguised by false information in the application of face recognition.

[0003] Early liveness detection methods mainly included two types: facial movements and color textures, and mainly relied on traditional machine learning methods to implement. By combining local binary pattern (LBP) with histogram of oriented gradients (HOG), they relied on pre-designed feature selection operators and feature extraction algorithms during liveness detection. The learning process was complex and the algorithm flexibility was poor. In recent years, deep learning network models have begun to be widely used in the field of liveness detection. For example, using the CNN network as the basic network, the face depth map is used to judge liveness and non-liveness.

[0004] Although deep learning has solved some problems in the training process of traditional machine learning, at the same time, deep learning technology has inevitably created a more complex liveness detection environment. New forgery attack methods such as video face swapping and AI-generated fake faces emerge in an endless stream, and traditional machine learning and deep learning are lacking in the detection ability of complex datasets and cross-dataset detection ability. Summary of the Invention

[0005] Object of the Invention: The first object of the present invention is to provide a face liveness detection method based on a two-stream convolutional neural network applicable to a complex liveness detection environment; the second object is to provide a face liveness detection system corresponding to the above face liveness detection method.

[0006] Technical Solution: A face liveness detection method based on a two-stream convolutional neural network includes the following steps:

[0007] S1. Preprocess the face feature image, convert the original face feature image into two different color space images, and construct a training set using the preprocessed face feature image;

[0008] S2. Construct a face liveness detection FAS model based on the VGG16 convolutional neural network model: use the VGG16 convolutional neural network model to extract the features of two different color space images respectively, perform feature fusion and then input them into the fully connected layer for classification;

[0009] S3. Use the training set to train the face liveness detection FAS model to obtain the optimal face liveness detection FAS model;

[0010] S4. Call the optimal face liveness detection FAS model to classify and detect the face feature image data.

[0011] Specifically, in step S1, the preprocessing of the face feature image includes:

[0012] S11. Crop all face feature images to the same size and normalize the pixel values of all face feature images;

[0013] S12. Convert the normalized face feature images into HSV color space images and YUV color space images respectively;

[0014] S13. Convert the labels of the face feature images into one-hot encoding and use the one-hot encoding to shuffle the original order of the face feature images.

[0015] Specifically, in step S12, the HSV color space image is obtained through the following steps:

[0016] Calculate the values of H, S, and V:

[0017] V = max(R, G, B)

[0018]

[0019] In the formula: R, G, and B are the intensities of the three channels of the image in the RGB color space, and H, S, and V are the intensities of the three channels of the image in the HSV color space;

[0020] Normalize the values of H, S, and V to the range of 0 to 255:

[0021] 0 ≤ H < 360 → H' = H ÷ 2, 0 ≤ H' < 180

[0022] 0 < S < 1 → S' = S × 255, 0 < S' < 255

[0023] 0 < V < 1 → V' = V × 255, 0 < V' < 255

[0024] In the formula: H', S', and V' are the intensities of the three channels of the image in the HSV color space after normalization.

[0025] Specifically, in step S12, the YUV color space image is obtained through the following steps:

[0026] Calculate the values of Y, U, and V:

[0027] Y = 0.299R + 0.587G + 0.114B

[0028] U = 0.492(B - Y)

[0029] V = 0.877(R - Y)

[0030] Where: R, G, and B are the intensities of the three channels of the image in the RGB color space, and Y, U, and V are the intensities of the three channels of the image in the YUV color space.

[0031] Specifically, in step S13, converting the label of the face feature image into a one-hot encoding is specifically as follows:

[0032]

[0033] Where: onehot(x) is the one-hot encoding function, is the indicator function, x is the classification variable, C i is the class value, i = 1, 2, 3…, k, and k is the number of class values.

[0034] Specifically, step S2 includes:

[0035] Input the face feature images converted to the HSV color space image and the YUV color space image into the VGG16 convolutional neural network model respectively to extract features, perform feature fusion and then dimensionality reduction processing, and then classify through three fully connected layers and output the classification result.

[0036] Specifically, step S3 includes:

[0037] Randomly divide an independent validation set from the training set according to a set ratio. After each round of training, use the model obtained in this round of training to test on the validation set to obtain the accuracy of the model obtained in this round of training on the validation set. If the accuracy of the models obtained in the subsequent three rounds of training does not improve, end the training and retain the model of this round of training as the optimal face liveness detection FAS model.

[0038] Specifically, in step S3, the calculation formula for accuracy is:

[0039]

[0040] Where: Accuracy is the accuracy, TP is the positive example data correctly classified by the classifier, TN is the negative example data correctly classified by the classifier, FP is the negative example data mislabeled as a positive example, and FN is the positive example data mislabeled as a negative example.

[0041] Specifically, step S4 includes:

[0042] Directly call the optimal face anti-spoofing (FAS) model to detect a single face facial information image or a face facial information dataset;

[0043] Output the face facial video frame by frame as images, and input them into the optimal face anti-spoofing (FAS) model for detection;

[0044] Connect to an image acquisition device, and call the optimal face anti-spoofing (FAS) model to detect the face facial information images collected in real time.

[0045] The present invention also provides a face anti-spoofing (FAS) detection system based on a two-stream convolutional neural network, including:

[0046] An image preprocessing module: used to preprocess the face feature image, convert the original face feature image into two different color space images, and construct a training set using the preprocessed face feature image;

[0047] An FAS model design module: used to construct a face anti-spoofing (FAS) model using the two-stream convolutional neural network model framework, design a two-stream convolutional layer for parallel extraction of image features using two different color spaces, and after the dual feature information extracted by the two-stream convolutional layer is transmitted to the pooling layer for processing, it is then classified through a fully connected layer;

[0048] An FAS model training module: used to train the face anti-spoofing (FAS) model using the training set to obtain the optimal face anti-spoofing (FAS) model;

[0049] An anti-spoofing detection module: used to call the optimal face anti-spoofing (FAS) model to classify and detect the face feature image data.

[0050] Advantageous effects: Compared with the prior art, the remarkable effect of the present invention is that: by converting the face feature image from the RGB color space to two different three-dimensional color spaces, namely HSV and YUV, and accordingly designing parallel two-stream convolutional layers to fuse the two different image feature information, in this way, the features in the two color spaces are comprehensively considered, increasing the classification credibility, and achieving extremely high accuracy in complex datasets and cross-dataset detections; the present invention randomly sorts the images in the training set through one-hot encoding, reducing the fitting problem that may be caused by adjacent images. Description of the Drawings

[0051] Figure 1 is the method flow chart of the present invention.

[0052] Figure 2 is the structural schematic diagram of the face anti-spoofing (FAS) model of the present invention.

[0053] Figure 3It is a schematic diagram of the training set images in different color spaces and different color channels in Embodiment 2 of the present invention. Detailed implementation manners

[0054] The following further illustrates a preferred solution of the present invention with reference to the accompanying drawings.

[0055] Embodiment 1

[0056] Please refer to Figure 1 As shown, this embodiment provides a face liveness detection method based on a two-stream convolutional neural network, including the following steps:

[0057] S1. Preprocess the face feature images, convert the original face feature images into two different color space images, and construct a training set using the preprocessed face feature images.

[0058] The preprocessing specifically includes:

[0059] S11. Crop all the face feature images to the same size. In this embodiment, the face feature images are cropped to a size of 128, and divide the pixel values of all the face feature images by 255.0 to normalize the pixel values to between [0, 1]. The normalization operation effectively avoids the incorrect training of the F model due to different pixel values of the input images in the training set, and at the same time, the normalization operation significantly improves the speed of the subsequent model training process.

[0060] S12. Convert the normalized face feature images into HSV color space images and YUV color space images respectively.

[0061] The HSV color space image is obtained through the following steps:

[0062] Calculate the values of H, S, and V:

[0063] V = max(R, G, B)

[0064]

[0065] In the formula: R, G, and B are the intensities of the three channels of the image in the RGB color space, and H, S, and V are the intensities of the three channels of the image in the HSV color space;

[0066] Normalize the values of H, S, and V to the range of 0 to 255:

[0067] 0 ≤ H < 360 → H' = H ÷ 2, 0 ≤ H' < 180

[0068] 0 < S < 1 → S' = S × 255, 0 < S' < 255

[0069] 0 < V < 1 → V' = V × 255, 0 < V' < 255

[0070] Where: H′, S′, and V′ are the intensities of the image in the three channels of the HSV color space after normalization.

[0071] Specifically, in step S12, the YUV color space image is obtained through the following steps:

[0072] Calculate the values of Y, U, and V:

[0073] Y = 0.299R + 0.587G + 0.114B

[0074] U = 0.492(B - Y)

[0075] V = 0.877(R - Y)

[0076] Where: R, G, and B are the intensities of the image in the three channels of the RGB color space, and Y, U, and V are the intensities of the image in the three channels of the YUV color space.

[0077] S13. Convert the label of the face feature image into a one-hot encoding, and use the one-hot encoding to shuffle the original order of the face feature images.

[0078] The mathematical principle of converting the label of the face feature image into a one-hot encoding is: Suppose there is a categorical variable x, and x can take k different categories, then the one-hot encoding can be defined as a vector function:

[0079]

[0080] Where: onehot(x) is the one-hot encoding function, is the indicator function, x is the categorical variable, C i is the category value, i = 1, 2, 3…, k, and k is the number of category values.

[0081] Use the one-hot encoding to randomly shuffle the order of the images in the dataset, and assign the originally adjacent pictures to new positions that have no association in the original dataset, effectively avoiding the dataset fitting that may occur during the training process.

[0082] S2. Use the two-stream convolutional neural network model framework to construct a face liveness detection FAS model, use the VGG16 convolutional neural network model to extract the features of two different color space images respectively, perform feature fusion and then input them into the fully connected layer for classification.

[0083] Please refer to Figure 2As shown, the dual-stream convolutional neural network model is modified according to the image size and format of the input face liveness detection FAS model, including the number of network layers and the connection relationship between the convolutional layer and the pooling layer. The feature information of the HSV color space image and the YUV color space image is extracted through the VGG16 convolutional neural network model respectively. After the feature information is fused, it is passed to the pooling layer for dimensionality reduction processing, and then classified through three fully connected layers to output the classification result. Finally, the classification probability of real and fake faces is output by the softmax function.

[0084] The VGG16 convolutional neural network consists of 16 layers of networks and is a typical convolutional neural network. This model is composed of multiple consecutive small convolutional kernels (3x3), pooling layers, and finally three fully connected layers. After each convolutional layer and fully connected layer, the rectified linear unit activation function (ReLU) is adopted. In this embodiment, the face liveness detection FAS model is a network structure of VGG16+HSV+YUV, that is, on the basis of the VGG16 convolutional neural network, the HSV and YUV color space images are used as inputs respectively. After the feature information is extracted and fused, a single output result is obtained, which is a dual-stream convolutional neural network with multiple inputs and a single output. The total number of parameters of the VGG16 convolutional neural network is 14,714,688, while the total number of parameters of the face liveness detection FAS model provided in this embodiment is 31,544,450, which is a more complex neural network structure. The specific layers and corresponding parameters are shown in Table 1 below.

[0085] Table 1: VGG16+HSV+YUV Network Structure

[0086]

[0087]

[0088] S3. Use the training set to train the face liveness detection FAS model to obtain the optimal face liveness detection FAS model.

[0089] Randomly divide an independent validation set from the training set according to a set ratio. After each round of training, use the model obtained in this round of training to test on the validation set to obtain the accuracy of the model obtained in this round of training on the validation set. If the accuracy of the models obtained in the subsequent three rounds of training has not improved, end the training and retain the model of this round of training as the optimal face liveness detection FAS model.

[0090] Specifically, in step S3, the calculation formula for the accuracy is:

[0091]

[0092] Where: Accuracy is the accuracy rate, TP is the positive example data correctly classified by the classifier, TN is the negative example data correctly classified by the classifier, FP is the negative example data mislabeled as a positive example, and FN is the positive example data mislabeled as a negative example.

[0093] S4. Invoke the optimal face liveness detection FAS model to perform classification detection on the face feature image data.

[0094] In this embodiment, the optimal face liveness detection FAS model can be invoked or deployed in the following three ways:

[0095] Directly invoke the optimal face liveness detection FAS model to detect a single face facial information image or a face facial information data set;

[0096] Output the face facial video frame by frame as an image and input it into the optimal face liveness detection FAS model for detection;

[0097] Connect to an image acquisition device and invoke the optimal face liveness detection FAS model to detect the face facial information image collected in real time.

[0098] Embodiment 2

[0099] In this embodiment, the face liveness detection method based on the two-stream convolutional neural network provided in Embodiment 1 is used to perform liveness detection instances on the NUAA data set and the AIFAKE data set respectively to verify the implementation effect of the present invention.

[0100] The NUAA data set is a publicly available data set made by Nanjing University of Aeronautics and Astronautics, which is divided into two major categories: real faces and fake faces. Among them, both real faces and fake faces are composed of 15 small categories, representing the real faces and fake faces of 15 subjects respectively. The data set includes 12,146 face images, among which there are 5,105 real facial images and 7,041 fake facial images. The data set takes into account the influence of different illuminations and different poses on liveness detection, and relevant images are taken for different poses, high exposure and insufficient illumination of each subject. The AIFAKE data set is a self-made data set, which is divided into two major categories: real faces and fake faces. Among them, the fake faces are composed of 3 small categories, which are divided into simple fake faces, medium fake faces, and difficult fake faces according to the degree of difficulty. All fake faces are generated by the GAN network. The data set includes 2,040 face images, among which there are 240 simple fake face images, 480 medium fake face images, 240 difficult fake face images and 1,080 real face images. When making the data set, the fraud images forged by the new network are considered, so all the fake face attack methods in the data set are attacks forged by the GAN network. The above data set information is summarized in Table 2 below:

[0101] Table 2: Sample Information of NUAA Data Set and AIFAKE Data Set

[0102]

[0103] The NUAA dataset and the AIFAKE dataset are respectively input into the face anti-spoofing (FAS) model for training. Please refer to Figure 2 As shown, the figure shows the original images of two sample images in the RGB color space and the images in the HSV color space, the YUV color space, the H, S, V channels of the HSV color space, and the Y, U, V channels of the YUV color space after being processed by the model. It should be noted that the images are all images from the public dataset or images generated or processed by artificial intelligence.

[0104] Evaluate the classification results of the datasets on the face anti-spoofing (FAS) model. Use the accuracy classifier to count the percentage of correctly classified data, the true positive rate (TPR) to count the percentage of correctly identified positive example data in the actual positive example data, and the false positive rate (FPR) to count the percentage of negative example data that is mispredicted as positive as the evaluation criteria.

[0105]

[0106]

[0107] In the formula: Accuracy is the accuracy rate, TP is the positive example data correctly classified by the classifier, TN is the negative example data correctly classified by the classifier, FP is the negative example data mislabeled as positive, FN is the positive example data mislabeled as negative, TPR is the true positive rate, and FPR is the false positive rate.

[0108] Design a cross-dataset detection experiment, that is, use the model trained with the NUAA dataset to detect the AIFAKE dataset; use the model trained with the AIFAKE dataset to detect the NUAA dataset.

[0109] Compare the detection results of the face anti-spoofing (FAS) model provided by the present invention with those of three conventional models on the above datasets. The results are shown in Tables 3 to 5.

[0110] Table 3: Cross-dataset detection results

[0111]

[0112] Table 4: Detection results of the NUAA dataset

[0113]

[0114] Table 5: Detection results of the AIFAKE dataset

[0115]

[0116] By comparing with other methods, it can be verified that the method proposed by the present invention can effectively improve the accuracy of live detection, effectively improve the cross-dataset detection accuracy, and effectively improve the detection ability of complex forged face images compared with the existing models.

[0117] Example 3

[0118] This example provides a face liveness detection system corresponding to the face liveness detection method described in Example 1, including the following modules:

[0119] Image preprocessing module: used to preprocess the face feature image, convert the original face feature image into two different color space images, and construct a training set using the preprocessed face feature image;

[0120] FAS model design module: used to construct a face liveness detection FAS model using the dual-stream convolutional neural network model framework, design a dual-stream convolutional layer for parallel extraction of image features using two different color spaces, and after the dual feature information extracted by the dual-stream convolutional layer is transmitted to the pooling layer for processing, it is then classified through the fully connected layer;

[0121] FAS model training module: used to train the face liveness detection FAS model using the training set to obtain the optimal face liveness detection FAS model;

[0122] Liveness detection module: used to call the optimal face liveness detection FAS model to classify and detect the face feature image data.

Claims

1. A method for face liveness detection based on a two-stream convolutional neural network, characterized in that: The following steps are involved: S1. Preprocessing the facial feature image, converting the original facial feature image into two different color space images, and using the preprocessed facial feature image to construct a training set; S2. Construct a face liveness detection FAS model based on the VGG16 convolutional neural network model: Use the VGG16 convolutional neural network model to extract two different color space image features, fuse the features, and then input them into the fully connected layer for classification; S3, using the training set to train the face liveness detection FAS model to obtain the optimal face liveness detection FAS model; S4. Call the optimal face liveness detection FAS model to classify and detect facial feature image data.

2. The method for face liveness detection according to claim 1, characterized in that: In the step S1, preprocessing the facial feature image includes: S11, cropping all facial feature images to the same size, and normalizing the pixel values ​​of all facial feature images; S12, converting the normalized facial feature image into an HSV color space image and a YUV color space image respectively; S13, converting the labels of the facial feature images into one-hot encoding, and using the one-hot encoding to disrupt the original order of the facial feature images.

3. The method for face liveness detection according to claim 2, characterized in that: In step S12, the HSV color space image is obtained by the following steps: Calculate the values ​​of H, S, and V: V = max(R, G, B) Where: R, G, B are the intensities of the three channels of the image in the RGB color space, H, S, V are the intensities of the three channels of the image in the HSV color space; Normalize the values ​​of H, S, and V to between 0 and 255: 0≤H<360→H′=H÷2,0≤H′<180 0<S<1→S′=S×255,0 <S′<255 0 <V<1 → v′=V×255.0 <V′<255 Where: H′, S′, and v′ are the intensities of the three channels of the image in the HSV color space after normalization.

4. The method for face liveness detection according to claim 2, characterized in that: In step S12, the YUV color space image is obtained by the following steps: Calculate the values ​​of Y, U, and V: Y=0.299R+0.587G+0.114B U=0.492(BY) V=0.877(RY) Where: R, G, B are the intensities of the three channels of the image in the RGB color space, and Y, U, V are the intensities of the three channels of the image in the YUV color space.

5. The method for face liveness detection according to claim 2, characterized in that: In step S13, the label of the facial feature image is converted into a one-hot encoding as follows: Where: onehot(x) is the one-hot encoding function, is the indicator function, x is a categorical variable, C i is the category value, i=1,2,3...,k, k is the number of category values.

6. The method for face liveness detection according to claim 2, characterized in that: The step S2 comprises: The facial feature images converted into HSV color space images and YUV color space images are respectively input into the VGG16 convolutional neural network model to extract features, and then the dimensionality reduction processing is performed after feature fusion, and then classified through three fully connected layers, and the classification results are output.

7. The method for face liveness detection according to claim 1, characterized in that: The step S3 comprises: An independent validation set is randomly divided from the training set according to the set ratio. After each round of training, the model obtained in this training is tested on the validation set to obtain the accuracy of the model obtained in this round of training on the validation set. If the accuracy of the model obtained in the next three rounds of training does not improve, the training is terminated and the model of this round of training is retained as the optimal face liveness detection FAS model.

8. The method for face liveness detection according to claim 7, characterized in that: In step S3, the calculation formula of accuracy is: Where: Accuracy is the accuracy rate, TP is the positive data correctly classified by the classifier, TN is the negative data correctly classified by the classifier, FP is the negative data incorrectly marked as positive, and FN is the positive data incorrectly marked as negative.

9. The method for face liveness detection according to claim 1, characterized in that: The step S4 comprises: Directly call the optimal face liveness detection FAS model to detect a single face image or a face dataset; Output the facial video frame by frame as images and input them into the optimal face liveness detection FAS model for detection; Connect to the image acquisition device and call the optimal face liveness detection FAS model to detect the facial information images collected in real time.

10. A face liveness detection system based on a two-stream convolutional neural network, characterized in that: include: Image preprocessing module: used to preprocess the facial feature image, convert the original facial feature image into two different color space images, and use the preprocessed facial feature image to construct a training set; FAS model design module: used to build a face liveness detection FAS model using a two-stream convolutional neural network model framework. A two-stream convolutional layer is designed to extract image features in parallel using two different color spaces. The dual feature information extracted by the two-stream convolutional layer is passed to the pooling layer for processing and then classified through the fully connected layer. FAS model training module: used to train the face liveness detection FAS model using the training set to obtain the optimal face liveness detection FAS model; Liveness detection module: used to call the optimal face liveness detection FAS model to classify and detect facial feature image data.

Citation Information

Cited By

  • Method, device and equipment for determining face authentic identification model, medium and product

    CN121214516A