A multi-classifier for face liveness detection based on a convolutional self-attention strategy
The convolutional self-attention strategy in the multi-classifier system addresses the limitations of binary detection in face spoofing by effectively classifying diverse attacks through enhanced feature extraction and reduced computational needs.
Patent Information
- Application Number
- CN202210778413.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-04
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-07-04
AI Technical Summary
The existing face spoof attack detection methods cannot effectively distinguish specific attack methods, making it difficult to adapt to diverse attack methods, and traditional CNN models cannot effectively extract long-distance features, resulting in difficulty in fine-grained classification.
A multi-classifier for face live detection based on convolutional self-attention strategy is used to extract features through convolutional operations in the data augmentation and improved Transformer model, and fine-grained classification is performed in combination with MLP structure, and local and long-distance features are extracted using the convolutional self-attention mechanism.
It has achieved effective detection of specific types of face deception attacks on the basis of ensuring high accuracy, with good generalization and computing power efficiency.
Smart Images

Figure CN114913589B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of face recognition, and particularly to a multi-classifier for face liveness detection based on a convolutional self-attention strategy. Background Art
[0002] With the development of face recognition related technologies, face recognition technology has been widely applied in fields such as security and payment. Along with the wide application of face recognition technology, its security has become increasingly important. In fact, face recognition systems may still be deceived by false face fingerprint information when facing specific attacks, resulting in serious consequences. For example, if a face payment system is attacked, it will cause serious economic losses to relevant users. Therefore, face liveness detection technology, as a pre-module of face recognition systems, plays an important role in protecting its security.
[0003] Early deception attacks were often of two types: print attacks and replay attacks, that is, using a printed photo of a specific face and playing a video of a specific face on an electronic screen to achieve the purpose of the attack. For these two attack methods, some models detect face deception behaviors by detecting temporal features, such as asking users to perform actions like blinking or turning their heads; some other models analyze the specific features of deception attacks, manually design feature descriptors, and use classifiers such as SVM for the extracted features to distinguish between real and fake faces.
[0004] With the development of related technologies, there are more and more attack methods for face deception attacks. For example, attackers begin to attack by 3D printing specific masks or applying specific makeup to the face. At the same time, due to the excellent performance of deep neural networks in the field of computer vision, people begin to introduce deep neural networks into the field of face liveness detection. Deep neural network models automatically extract the features of deception attacks through means such as convolution and attention mechanisms, and the features they extract show extremely high accuracy. These methods often take a convolutional neural network (CNN) as the main body and design corresponding structures. For example, some methods design a two-stream network to simultaneously extract spatial domain features and temporal rPPG features from deception samples, thereby achieving effective detection of deception attacks.
[0005] The existing technology often regards face spoofing attack detection as a binary classification problem, that is, the relevant models only focus on whether it is a spoofed face, rather than the specific types of spoofing attacks. In the early stage, due to the relatively scarce attack means, it was reasonable to regard it as a binary classification problem. However, with the development of attack technologies, there are more and more attack methods and purposes. Different attack methods often reflect the attacker's technical ability and financial support. At the same time, distinguishing different attack methods is also beneficial for the system to adopt more reasonable processing strategies. Since the previous liveness detection methods regarded detection as a binary classification problem, they could not effectively distinguish specific attack methods, making it difficult for them to adapt to more and more attack methods. Relevant personnel can only blindly and inefficiently use the same strategy to handle all face spoofing attacks.
[0006] At the same time, due to the traditional CNN model's inability to effectively extract long-distance features, most of the previous face liveness detection models could not effectively utilize long-distance features, so it was difficult for them to achieve effective fine-grained classification. To solve this problem, we introduced the self-attention mechanism into face liveness detection, thereby achieving effective detection of specific types of spoofing attacks on the basis of accurately identifying real and fake faces.
[0007] In addition, there are also a small number of face spoofing attack detection methods based on self-attention. However, these methods often use linear mapping to achieve the purpose of relevant feature extraction. This feature extraction method cannot make good use of local features in the image and requires a large amount of computing power, making its application value in the field of face liveness detection not high. Summary of the Invention
[0008] Aiming at the problem that the existing face spoofing attack detection methods cannot effectively distinguish specific attack methods, making it difficult for them to adapt to more and more attack methods, the present invention proposes a multi-classifier for face liveness detection based on a convolutional self-attention strategy.
[0009] The method of the present invention is as follows:
[0010] Step 1: Augment the input image, transfer the input image to different styles, and at the same time, maximize the retention of the original key information of the image;
[0011] Step 2: Convert the calculation methods of Query, Key, and Value in the Transformer model into convolutional operations, and perform separate convolutional processing on Key and then superimpose them. Import the augmented image in Step 1 into the Transformer model for processing;
[0012] Step 3: Reduce the dimension of the processed feature map in Step 2 using an MLP structure, and determine the confidence through a softmax layer. The value with the highest confidence is the determined classification result:
[0013] result = Max(Softmax(Linear(Linear(Linear(youtputfinal)))))。
[0014] Preferably, the data augmentation of the image in Step 1 includes the following sub-steps: Sub-step 1-1: Process the input image using a convolution with a kernel size of 3 and a stride of 1:
[0015] y feature1 = Conv2d(xinput, stride = 1);
[0016] Sub-step 1-2: Process the output of the first step using a convolution with a kernel size of 3 and a stride of 1, and put it into the first stream: y feature2 = Conv2d(Conv2d(y feature1 , stride = 1), stride = 1);
[0017] Sub-step 1-3: Process the output of the first step using a convolution with a kernel size of 3 and a stride of 2, and put it into the second stream: y feature3 = Conv2d(Conv2d(y feature1 , stride = 2), stride = 1);
[0018] Sub-step 1-4: Process the output of the first step using a convolution with a kernel size of 3 and a stride of 4, and put it into the third stream: y feature4 = Conv2d(Conv2d(y feature1 , stride = 4), stride = 1);
[0019] Sub-step 1-5: For the feature maps in the second and third streams, use bilinear interpolation to increase the scale to the same as the input image, and concatenate the outputs of the three streams together:
[0020] y featureCat = Concat(y feature2 , Concat(BilinearUpsampling(y feature3 ), BilinearUpsampling(y feature4 )));
[0021] Sub-step 1-6: Use a convolution with a kernel size of 3 and a stride of 1 to reduce the dimension of the feature map output in the previous step by half:
[0022] y final = Conv2d(Conv2d(y featureCat , stride = 1), stride = 1).
[0023] Use a three-stream network to extract different types of style features in the image, and then fuse these different features to obtain images with different style features. Each parallel branch stream is designed based on the spatial pyramid characteristics of the convolution operation, converting the image into different scales to obtain features of different dimensions from low-dimensional to high-dimensional. Among them, low-dimensional features are generally features such as texture and object edges, middle-dimensional features are generally color block features such as lips and eyes, and high-dimensional features are generally abstract features such as identity information. These three major types of features correspond to the three streams of style transfer respectively, so that it can be well transferred to different styles on the basis of retaining the key information of the original image, thus achieving the purpose of augmentation.
[0024] Preferably, three self-attention blocks are used in step two, and the calculation methods of Query, Key, and Value of the self-attention block are changed from linear mapping to convolution operation. The detailed steps are as follows:
[0025] Sub-step 2-1, use a convolution with a kernel size of 3 and a stride of 2 to process the input image to obtain Key:
[0026] Key = Conv(x input , stride = 2);
[0027] Sub-step 2-2, use a convolution with a kernel size of 3 and a stride of 2 to process the input image to obtain Value:
[0028] Value = Conv(x input , stride = 2);
[0029] Sub-step 2-3, use a convolution with a kernel size of 3 and a stride of 1 to process the input image to obtain Query:
[0030] Query = Conv(x input , stride = 1);
[0031] Sub-step 2-4, perform a dot product on Key and Query to obtain a weight matrix:
[0032]
[0033] Sub-step 2-5, stack the weight matrix with the Key matrix, and then use a convolution with a kernel size of 3 and a stride of 1 to process the stacked matrix to obtain an attention matrix:
[0034] Attention = Conv(Concat(WeightMatrix, Key), stride = 1);
[0035] Sub-step 2.6: Multiply the attention matrix with Value to obtain the output feature map:
[0036]
[0037] In this way, the ability of convolution to extract local features can be utilized to effectively extract local features in the image. At the same time, due to the existence of the subsequent self-attention mechanism, it can continue to retain the ability to extract long-distance features. In addition, since the relative relationship between Keys in the previous attention mechanism was not well utilized, we perform convolution processing on Key alone and then stack them, so that the relative relationship between Keys can be fully utilized in the entire attention mechanism.
[0038] In Step 3, we use a simple MLP structure to reduce the dimension of the features we obtained, so as to obtain the confidence of each category we need. Among them, MLP uses three layers. The input dimension of the first layer is the same as the output dimension of the last self-attention block. In the second layer, we reduce its dimension to half of the first layer. The last layer reduces the dimension to the same as the number of categories we need. Through a softmax layer, the confidence of each category can be obtained, where the first category is the real face, and the other categories are the fine-grained categories of spoofed faces.
[0039] The substantial features of the present invention are as follows: Face liveness detection is no longer limited to binary classification tasks. We only need to perform one detection to obtain the specific attack category. It has good generalization ability on the basis of ensuring sufficient fine-grained detection ability of face spoofing. Compared with other face spoofing attack detection methods based on self-attention mechanism, it requires less data set quantity and computing power. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 The overall structure diagram of the present invention.
[0041] Figure 2 The data augmentation structure diagram of the present invention;
[0042] Figure 3 The data augmentation schematic diagram of the present invention;
[0043] Figure 4 The convolutional self-attention structure diagram of the present invention;
[0044] Figure 5 The convolutional self-attention schematic diagram of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0045] The technical solution of the present invention will be further specifically described below through specific embodiments in conjunction with the accompanying drawings.
[0046] Embodiment 1
[0047] As Figure 1 shown, our face liveness detection method is mainly divided into three main parts
[0048] The first part is the data augmentation module, which consists of a multi-stream convolutional neural network. Based on the spatial pyramid hierarchical characteristics of the convolutional neural network, the mutual relationship between different streams in this module is designed. Based on these different features, we can transfer the input image to different styles while maximizing the retention of the original key information of the image. Thus, the purpose of data augmentation is achieved.
[0049] The second part is the convolutional self-attention spoofing feature extraction module. This part is improved based on the Transformer model in the field of computer vision. Convolutional operations are introduced into the Transformer model, and certain improvements are made to the self-attention mechanism in the Transformer, making it more suitable for face liveness detection tasks. Since the original Transformer network uses linear mapping as the main means of feature extraction, but linear mapping has the following two disadvantages in face liveness detection tasks: 1. Linear mapping cannot make good use of local features in the image, 2. Linear mapping requires a large amount of computing power. Therefore, in this part, we use convolutional operations to replace the original linear mapping operations. At the same time, since the relative relationship between different blocks in the original self-attention mechanism is not well utilized, we also make certain modifications to the operation of the self-attention mechanism. The convolutional self-attention mechanism will be repeated three rounds in this module, and the calculation methods in the three rounds are exactly the same. The input of the first round is the augmented image, and the inputs of the latter two rounds are the outputs of the previous round.
[0050] In the third part, we introduce a multi-layer perceptron. We use the feature map output by the second part as the input of this multi-layer perceptron, and further reduce the dimension of the features extracted in the previous step to obtain the confidence of each specific classification for us.
[0051] As Figure 2 and Figure 3As shown, in the first - part data augmentation module, we use a three - stream network to extract different types of style features in the image, and then fuse these different features to obtain images with different style features. Each parallel branch stream is designed based on the spatial pyramid property of the convolution operation, converting the image into different scales to obtain features of different dimensions from low - dimension to high - dimension. Among them, low - dimension features are generally features such as texture and object edges, middle - dimension features are generally features such as color patches of lips and eyes, and high - dimension features are generally abstract features such as identity information. These three major types of features respectively correspond to the three streams of style transfer, so that it can well transfer the original image to different styles while retaining the key information of the original image, thus achieving the purpose of augmentation. The overall steps are as follows:
[0052] 1) Process the input image using a convolution with a kernel size of 3 and a stride of 1
[0053] y feature1 = Conv2d(x input , stride = 1)
[0054] 2) Process the output of the first step using a convolution with a kernel size of 3 and a stride of 1, and put it into the first stream
[0055] y feature2 = Conv2d(Conv2d(y feature1 , stride = 1), stride = 1)
[0056] 3) Process the output of the first step using a convolution with a kernel size of 3 and a stride of 2, and put it into the second stream
[0057] y feature3 = Conv2d(Conv2d(y feature1 , stride = 2), stride = 1)
[0058] 4) Process the output of the first step using a convolution with a kernel size of 3 and a stride of 4, and put it into the third stream
[0059] y feature4 = Conv2d(Conv2d(y feature1 , stride = 4), stride = 1)
[0060] 5) For the feature maps in the second and third streams, use bilinear interpolation to increase the scale to the same as the input image, and concatenate the outputs of the three streams together
[0061] y featureCat
[0062] = Concat(yfeature2 , Concat(BilinearUpsampling(y feature3 ), BilinearUpsampling(y feature4 )))
[0063] 6) Use a convolution with a kernel size of 3 and a stride of 1 to reduce the dimension of the feature map output in the previous step by half
[0064] y final = Conv2d(Conv2d(y featureCat , stride = 1), stride = 1)
[0065] In the second part, we use three self-attention blocks. The inside of each self-attention block is as Figure 4 and Figure 5 shown. Compared with the traditional self-attention mechanism, we change the calculation method of Query, Key, and Value from a linear mapping to a convolution operation. This can utilize the ability of convolution to extract local features and effectively extract local features in the image. At the same time, due to the existence of the subsequent self-attention mechanism, it can continue to retain the ability to extract long-distance features. In addition, since the relative relationship between Keys in the previous attention mechanism was not well utilized, we perform a convolution operation on Keys separately and then stack them, so that the relative relationship between Keys can be fully utilized in the entire attention mechanism. The overall steps are as follows:
[0066] 1) Use a convolution with a kernel size of 3 and a stride of 2 to process the input image to obtain Key
[0067] Key = Conv(x input , stride = 2)
[0068] 2) Use a convolution with a kernel size of 3 and a stride of 2 to process the input image to obtain Value
[0069] Value = Conv(x input , stride = 2)
[0070] 3) Use a convolution with a kernel size of 3 and a stride of 1 to process the input image to obtain Query
[0071] Query = Conv(x input , stride = 1)
[0072] 4) Perform a dot product on Key and Query to obtain the weight matrix
[0073]
[0074] 5) Superimpose the weight matrix and the Key matrix, and then process the superimposed matrix using a convolution with a kernel size of 3 and a stride of 1 to obtain the attention matrix
[0075] Attention = Conv(Concat(WeightMatrix, Key), stride = 1)
[0076] 6) Multiply the attention matrix with Value to obtain the output feature map
[0077]
[0078] In the third part, we use a simple MLP structure to reduce the dimension of the features we obtained, so as to obtain the confidence of each category we need. The MLP uses three layers. The input dimension of the first layer is the same as the output dimension of the last self-attention block. In the second layer, we reduce its dimension to half of the first layer. The last layer reduces the dimension to the same as the number of categories we need. Then, through a softmax layer, we can obtain the confidence of each category. The first category is the real face, and the other categories are the fine-grained categories of spoofing faces. The category with the highest confidence is the prediction result of the model.
[0079] result = Max(Softmax(Linear(Linear(Linear(y outputfinal ))))
Claims
1. A multi-classifier for face liveness detection based on a convolutional self-attention strategy, characterized in that, It includes the following steps: Step 1: Augment the input image, migrate the input image to different styles, and at the same time maximize the retention of the original key information of the image; The following sub-steps are included in Step 1: Sub-step 1-1: Process the input image using a convolution with a kernel size of 3 and a stride of 1: y feature1 = Conv2d(x input , stride=1); Sub-step 1-2: Process the output of the first step using a convolution with a kernel size of 3 and a stride of 1, and put it into the first stream: y feature2 = Conv2d(Conv2y(y feature1 , stride=1), stride=1); Sub-step 1-3: Process the output of the first step using a convolution with a kernel size of 3 and a stride of 2, and put it into the second stream: y feature3 = Conv2d(Conv2d(y feature1 , stride=2), stride=1); Sub-step 1.4, process the output of the first step using a convolution with a kernel size of 3 and a stride of 4, and put it into the third stream: y feature4 = Conv2d(Conv2d(y feature1 , stride = 4), stride = 1); Sub-step 1-5: For the feature maps in the second and third streams, use bilinear interpolation to upscale the scale to the same as the input image, and splice the outputs of the three streams together: y featureCat = Concat(y feature2 , Concat(BilinearUpsampling(y feature3 ), BilinearUpsampling(y feature4 ))); Sub-step 1-6, reduce the dimension of the feature map output in the previous step by half using a convolution with a kernel size of 3 and a stride of 1: y final = Conv2d(Conv2d(y featureCat , stride=1), stride=1); Step 2: Convert the calculation methods of Query, Key, and Value in the Transformer model into convolution operations, and perform convolution processing on Key alone and then stack it. Import the augmented image in Step 1 into the Transformer model for processing; In Step 2, the following sub-steps are included: Sub-step 2-1, process the input image using a convolution with a kernel size of 3 and a stride of 2 to obtain Key: Key = Conv(x input , stride = 2); Sub-step 22, process the input image using a convolution with a kernel size of 3 and a stride of 2 to obtain Value: Value = Conv(x input , stride = 2); Sub-step 2-3: Process the input image using a convolution with a kernel size of 3 and a stride of 1 to obtain Query: Query = Conv(x input , stride = 1); Sub-step 2-4: Perform a dot product of Key and Query to obtain a weight matrix: Sub-step 2-5: Stack the weight matrix with the Key matrix, and then process the stacked matrix using a convolution with a kernel size of 3 and a stride of 1 to obtain an attention matrix: Attention = Conv(Concat(WeightMatrix, Key), stride = 1); Sub-step 2-6: Perform a dot product of the attention matrix and Value to obtain the output feature map: Step 3: Downscale the processed feature map in Step 2 in an MLP structure, and judge the confidence through a softmax layer. The value with the highest confidence is the determined classification result: result = Max(Softmax(Linera(Linear(Linear(y outputfinal )))))。
Citation Information
Patent Citations
Face recognition method based on convolutional neural network and attention model
CN111582044A
Face deception detection method based on feature space constraint
CN113221655A