Face liveness detection method based on MobileFaceNet
By improving MobileFaceNet to build a lightweight face liveness detection model, the problem of excessive memory requirements on embedded devices is solved, and efficient and low-latency face liveness detection is achieved, which is suitable for embedded devices with limited resources.
Patent Information
- Application Number
- CN202211499885.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-28
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2042-11-28
AI Technical Summary
Existing face liveness detection models have excessive memory requirements when deployed on embedded devices and cannot be effectively applied to embedded devices with limited resources.
By improving MobileFaceNet, a lightweight face liveness detection model is constructed, including a backbone feature extraction network and a classifier network. A lightweight network structure and improvements to the convolution block are adopted to reduce the number of bottles and the number of channels in the convolution block. The model parameters are exported using a TXT file, and forward reasoning is performed using the C language.
It achieves efficient and low-latency face liveness detection on embedded devices, significantly reducing the model size to 51KB while maintaining high detection accuracy and completing the detection task without additional hardware acceleration.
Smart Images

Figure CN115798058B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of face recognition, and in particular to a face liveness detection method based on MobileFaceNet. Background Art
[0002] The rapid development of biometric technology has greatly facilitated people's lives. Common biometric technologies include facial recognition, fingerprint recognition, and iris recognition. These technologies are often used in scenarios such as facial payment, fingerprint unlocking, and face unlocking, which have a significant impact on people's privacy and property security. Common attacks in the facial recognition field include photo attacks, video attacks, and 3D headgear attacks. To protect people's privacy and property security, technology is needed to prevent these attacks. Liveness detection is a technology that determines whether the identified object is alive. It is often used in conjunction with biometric technology. Performing liveness detection before recognition can effectively prevent attacks.
[0003] In related technologies, face liveness detection uses a more lightweight network based on MobileFaceNet, referred to as MobileASFNet. With a model size of only 4MB, this MobileFaceNet is well-suited for deployment on embedded or mobile devices with limited computing resources. However, the application targets smaller embedded devices, requiring a network with memory requirements of only a few hundred KB. Therefore, MobileFaceNet needs to be modified to achieve a more streamlined and efficient architecture. Summary of the Invention
[0004] This embodiment of the present application provides a method for face liveness detection based on MobileFaceNet. This method addresses the issues of overly large network models and embedded device deployment. The method comprises the following steps:
[0005] The dataset containing real and fake face photos is divided into training, test, and validation sets according to the proportions, and the faces of the photos are annotated to obtain cropped face images.
[0006] Construct a backbone feature extraction network for the face liveness detection model and iteratively train it using a dataset. The backbone feature extraction network is used to extract image features and includes three common convolutional blocks (CW), one deep convolutional block (DW), nine bottleneck layer (BK) modules, and two point-by-point convolutional blocks (PW).
[0007] Constructing a classifier network of the face liveness detection model and performing iterative training using the dataset and the trained backbone feature extraction network; the classifier network is used to make judgments based on the extracted image features to obtain liveness probability and non-liveness probability;
[0008] The model parameters of the face liveness detection model are exported and deployed to an embedded device. The network structure and convolution operations between network layers are executed using C language, and a forward reasoning process is performed to obtain the face liveness detection model.
[0009] Specifically, the data set containing real face and fake face photos is divided into a training set, a test set, and a validation set according to a certain ratio, and the faces of the photos are annotated to obtain cropped face images, including:
[0010] The dataset is divided into a training set, a validation set, and a test set in a ratio of 7:1:2;
[0011] Identify the face images in the dataset and return the coordinate information of the face frame in the image;
[0012] The coordinates of the face frame are doubled in all directions according to the coordinate information, and the face is cropped and extracted according to the face frame to obtain a cropped face image.
[0013] Specifically, the backbone feature extraction network inputs a face image of size 96*96; the output is a 128-dimensional feature vector extracted from the face image;
[0014] The first common convolution block and the depth convolution block are connected before the nine bottleneck layer modules, and the convolution kernel size of the first common convolution block is 3*3, the step size s=3, and the number of output channels is 2; the convolution kernel size of the depth convolution block is 3*3, the step size s=1, and the number of output channels is 2;
[0015] The nine BK modules are divided into three groups. The number of channels in the first two groups is 4. A direct connection shortcut structure is set between the BK modules in each group. The step size of the first BK module in each group is s=2, and the step size of the second and third BK modules is s=1. The number of channels in the third group is 8. A shortcut structure is set between the BK modules in the group. The step size of the first BK module is s=2, and the step size of the second and third BK modules is s=1.
[0016] After the BK module, there are a first point-by-point convolution block, a second ordinary convolution block, a third ordinary convolution block and a second point-by-point convolution block in cascade sequence; the convolution kernel of the first point-by-point convolution block is 1*1, the step size s=1, and the number of output channels is 8; the convolution kernel of the second ordinary convolution block is 3*3, the step size s=2, and the number of output channels is 8; the convolution kernel of the third ordinary convolution block is 3*3, the step size s=1, and the number of output channels is 8; the convolution kernel of the point-by-point convolution block is 1*1, the step size s=2, and the number of output channels is 128.
[0017] Specifically, the training cycle of the backbone feature extraction network is 90 epochs, the batch size is set to 128, and the initial learning rate of the SGD optimizer is 0.01;
[0018] When epochs reaches 30, the learning rate is adjusted to 0.001, and when epochs reaches 60, the learning rate is adjusted to 0.0001.
[0019] Specifically, the input of the classifier network is the 128-dimensional feature vector output by the backbone feature extraction network, including the fourth ordinary convolution and the fifth ordinary convolution; the convolution kernel of the fourth ordinary convolution is 3*3, the step size s=1, and the number of output channels is 32; the convolution kernel of the fifth ordinary convolution is 3*3, the step size s=1, the number of output channels is 2, and the two output channels represent the probability of liveness and the probability of non-liveness, respectively.
[0020] Specifically, the training cycle of the classifier network is 90 epochs, the batch size is set to 128, and the initial learning rate of the SGD optimizer is 0.01;
[0021] When epochs reaches 30, the learning rate is adjusted to 0.001, and when epochs reaches 60, the learning rate is adjusted to 0.0001.
[0022] Specifically, the label of the living face image is set to 1, and the label of the prosthetic face image is set to 0;
[0023] Sending the label and the face image into the face liveness detection model to respectively train the backbone feature extraction network and the classifier network, and adjusting the model batch size, learning rate and gradient parameters according to the liveness probability and the non-liveness probability;
[0024] The iterative training is stopped when the live probability minus the non-live probability output by the classifier network is greater than 0.2, or the number of training times reaches a set threshold.
[0025] Specifically, the first loss function L1 of the backbone feature extraction network is expressed as:
[0026]
[0027] Among them, the front part of L1 is SoftMax Loss, the back part is Triplet Loss, D(a,n) represents the distance between the sample to be predicted and the negative sample, D(a,p) represents the distance between the sample to be predicted and the positive sample, m is 0.2, and σ is 0.3;
[0028] The second loss function L2 of the classifier network is expressed as:
[0029]
[0030] Among them, y i is the label value, a one-hot encoding consisting of 0 and 1, y i ' is the predicted value, y i '∈(0,1).
[0031] Specifically, the method comprises: exporting the model parameters of the face liveness detection model and deploying them to an embedded device; executing the network structure and convolution operations between network layers through C language; performing the forward reasoning process to obtain the face liveness detection model; and
[0032] Save the model's weight parameters in the form of a TXT file;
[0033] The saved weight parameter file is read using C language, imported into the embedded end, and the network structure of the face liveness detection model and the convolution operation between the network layers are constructed using C language, and the forward reasoning process is performed to obtain the face liveness detection model.
[0034] The beneficial effects brought about by the technical solution provided by the embodiment of the present application include at least: improving the backbone network based on the original MobileFaceNet, on the one hand reducing the number of Bottlenecks, and on the other hand reducing the number of channels of each convolution block network and the number of Bottleneck channels. For the original 6*6 global depth convolution block, this solution splits it into two 3*3 depth convolutions, reducing the amount of calculation and parameters, while also improving the feature extraction capability. The improved face liveness detection model has greatly reduced memory, and when deployed on embedded devices, it uses TXT text export and C language format to advance forward, without the assistance of additional hardware acceleration, it can complete face liveness detection under low latency requirements while maintaining high accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 This is the network structure diagram of the standard MobileFaceNet model;
[0036] Figure 2 This is a flowchart of the face liveness detection method based on MobileFaceNet provided in an embodiment of the present application;
[0037] Figure 3 This is a network structure diagram of the backbone feature extraction network in the face liveness detection model provided in an embodiment of the present application;
[0038] Figure 4 This is a network structure diagram of the classifier network in the face liveness detection model provided in an embodiment of the present application. DETAILED DESCRIPTION
[0039] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0040] In this document, "plurality" refers to two or more. "And / or" describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. The character " / " generally indicates an "or" relationship between the associated objects.
[0041] With the development of facial recognition technology, facial liveness detection technology has been applied in various scenarios, such as facial attendance, facial gate access, and facial payment. With the widespread application of facial liveness detection technology, it is no longer limited to high-performance servers, but is increasingly being applied to embedded systems. However, due to the limited resources of embedded devices, existing mainstream large models cannot be ported to embedded systems. Therefore, the development of an efficient, highly accurate facial liveness detection method that can be applied to embedded devices is of great significance for practical engineering applications.
[0042] The solution provided in this application uses a backbone network + classifier model to perform face liveness detection. The backbone network is used for feature extraction, and the classifier is used to classify the extracted features. The backbone network is a lighter network improved based on MobileFaceNet, referred to as MobileASFNet. The MobileFaceNet model is mainly used for facial feature extraction. The model size is only 4M, which is very suitable for deployment in embedded devices or mobile devices with limited computing resources. However, this patent is aimed at smaller-scale embedded devices, and it is necessary to design a network with a memory requirement of several hundred KB, so MobileFaceNet needs to be modified.
[0043] like Figure 1 Figure 2 shows the network structure of the standard MobileFaceNet model. The structural connection order includes a normal convolution block, a depthwise convolution block, 15 consecutive BK modules (bottleneck layer, abbreviated as BK), a point-by-point convolution block, a depthwise convolution, and a point-by-point convolution. Both the normal and depthwise convolution blocks use 3*3 convolution kernels, and the 15 BK modules have 64-channel outputs, 128-channel outputs, and 512-channel outputs. The normal and depthwise convolutions connected to the output of the BK module both have 512-channel outputs. Some adjacent BK modules are connected by shortcuts. As can be seen from the figure, the standard MobileFaceNet model has a large number of embedded BK modules, and the number of channels in each convolution layer is also set to a large number. Therefore, the network has a large number of parameters and computational complexity. The actual lightweight network size is 4M.
[0044] Figure 2 This is a flowchart of the face liveness detection method based on MobileFaceNet provided in an embodiment of the present application. It is an improved detection model based on the standard MobileFaceNet. It specifically includes the following steps:
[0045] In step 201, the acquired data set containing real face and fake face photos is divided into a training set, a test set, and a validation set according to a certain ratio, and the faces of the photos are annotated to obtain cropped face images.
[0046] a. Divide the dataset into training set, validation set, and test set in a ratio of 7:1:2;
[0047] b. Identify the face images in the dataset and return the coordinate information of the face frame in the image;
[0048] This step mainly involves identifying the content of the photo, using a common face detection algorithm to frame the image, and determining the coordinate position information of the face in the image.
[0049] c. Double the coordinates of the face frame in all directions according to the coordinate information, and crop and extract the face according to the face frame to obtain a cropped face image.
[0050] This step normalizes the images to ensure that all faces are included in the captured image for subsequent iterative training and validation. The images in this solution are normalized to a size of 96*96 to facilitate subsequent model training.
[0051] Step 202: construct a backbone feature extraction network for a face liveness detection model and perform iterative training using a data set.
[0052] This solution improves the backbone network based on MobileFaceNet. The improved backbone feature extraction network includes three common convolution blocks CW, one deep convolution block DW, nine bottleneck layer BK modules and two point-by-point convolution blocks PW; its structure is as follows Figure 3As shown, it includes the first common convolution block CW, which is located before the nine BK modules. Its convolution kernel size is 3*3, the stride s=1, and the number of output channels is 2. After the 96*96 size photo passes through the convolution block, the output size is 48*48. The first common convolution block CW is followed by the depth convolution block DW. The convolution kernel of this depth convolution block is 3*3, the stride s=1, and the number of channels is also 2. The output size of the 48*48 size photo remains unchanged after passing through the convolution network. The depth convolution block is cascaded with nine BK modules. These nine BK modules are divided into three groups. The number of channels in the first two groups is all 4. A shortcut structure is set between the BK modules in each group. The stride s=2 for the first BK module in each group, and the stride s=1 for the second and third BK modules. The third group has eight channels, with shortcuts between the BK modules within the group. The step size of the first BK module is s = 2, and the step sizes of the second and third BK modules are s = 1. This cascaded structure of three BK modules replaces the original three complex BK modules, reducing the number of channels to nine and significantly reducing the number of channels per group.
[0053] After the BK module, the first point-by-point convolution block, the second normal convolution block, the third normal convolution block, and the third point-by-point convolution block are sequentially connected. The convolution kernel of the first point-by-point convolution block is 1*1, the stride s=1, and the number of output channels is 8. The convolution kernel of the second normal convolution block is 3*3, the stride s=2, and the number of output channels is 8. The convolution kernel of the third normal convolution block is 3*3, the stride s=1, and the number of output channels is 8. The convolution kernel of the point-by-point convolution block is 1*1, the stride s=2, and the number of output channels is 128. For detailed data of the network components of this backbone feature extraction network structure, please refer to Table 1.
[0054] Table 1 Detailed data of the network components of the feature extraction network structure
[0055] enter operate Number of channels n / number of superpositions S / step length 96×96 CW 2 1 2 48×48 DW 2 1 1 48×48 BK 4 3 2 24×24 BK 4 3 2 12×12 BK 8 3 2 6×6 PW 8 1 1 6×6 CW 8 1 2 3×3 CW 8 1 1 1×1 PW 128 1 1
[0056] As can be seen in Table 1, the 96*96 input size is reduced to 48*48 after CW and DW operations. After the first set of three BK modules, the size is reduced to 24*24. After the second set of BK modules, the size is reduced to 12*12, further reducing the size and increasing the number of channels. With the third set of BK modules, the output size is reduced to 6*6. Regarding the 6*6 global depthwise convolution cascaded after the BK module in the original model, this solution replaces it with two 3*3 depthwise convolutions. Because the two 3*3 depthwise convolutions have fewer parameters and computational complexity than the 6*6 global depthwise convolution, and have stronger feature extraction capabilities, the memory usage of this part of the model is also reduced compared to the previous one. Finally, the second point-by-point convolution block outputs the extracted 128-dimensional feature vector.
[0057] The classifier network, located after the backbone feature extraction network, is primarily used to identify and classify the extracted 128-dimensional feature vectors, determining whether an object is alive or not, that is, distinguishing between real and fake faces. This section is divided into two cascaded fourth and fifth ordinary convolution blocks. The fourth ordinary convolution block has a 3*3 convolution kernel, a stride s=1, and 32 output channels; the fifth ordinary convolution block has a 3*3 convolution kernel, a stride s=1, and 2 output channels, with the two output channels representing the probability of a live object and the probability of a non-live object, respectively. It is important to emphasize that the convolution block includes a convolutional network, a batch normalization layer, and a Prelude Unit (PReLU) activation function.
[0058] Step 203: construct a classifier network of a face liveness detection model, and perform iterative training using the data set and the trained backbone feature extraction network.
[0059] The backbone feature extraction network and the classifier network require separate iterative training, so the loss functions must also be constructed separately: the first loss function and the second loss function. Since attack methods are often difficult to defend against, we can consider the attack method as an open set, while the real face is a closed set. We view the liveness detection problem from the perspective of anomaly detection. For the backbone network, we use Triplet Focal Loss, which combines Triplet Loss and Focal Loss, as shown below:
[0060]
[0061] This loss function consists of two parts: the first part is SoftMax Loss, and the second part is Triplet Loss. Triplet Loss is a regularization constraint on the first part, which can maintain the distinguishability between features rather than directly using classification loss to guide the gradient. Where D(a,n) represents the distance between the sample to be predicted and the negative sample, and D(a,p) represents the distance between the sample to be predicted and the positive sample. The value of m is 0.2, and the value of σ is 0.3.
[0062] For the second loss function of the classifier, this scheme uses the cross entropy loss function for training. The cross entropy loss function is as follows:
[0063]
[0064] Among them, y i is the label value, a one-hot encoding consisting of 0 and 1, y i ' is the predicted value, y i '∈(0,1).
[0065] The face liveness detection model constructed based on the above content is iteratively trained. Before training, the dataset needs to be labeled, that is, the label of the live face image is set to 1, and the label of the fake face image is set to 0. Then, the label and face image are sent to the face liveness detection model.
[0066] First, the backbone feature extraction network was trained with a training epoch of 90 and a batch size of 128. The SGD optimizer had an initial learning rate of 0.01. When the input dataset reached 30 epochs, the learning rate was multiplied by 0.1, bringing it to 0.001. When the epoch reached 60, the learning rate was multiplied by 0.1 again, bringing it to 0.0001. During the iterative training process, the model batch size, learning rate, and gradient parameters were adjusted based on the live and non-live probabilities.
[0067] Similarly, after training the backbone feature extraction network, we used it to extract feature vectors and fed them into the classifier network for model training. The classifier network was trained for 90 epochs, with a batch size of 128, and an SGD optimizer with an initial learning rate of 0.01. When the number of epochs reached 30, the learning rate was increased to 0.001, and when the number of epochs reached 60, the learning rate was increased to 0.0001.
[0068] In this solution, the liveness probability + the non-liveness probability = 1. For a trained face liveness detection model, if the liveness probability minus the non-liveness probability of the input face photo is greater than 0.2, the predicted result is liveness; otherwise, it is non-liveness. After training, the model size of this solution is 51KB, significantly smaller than traditional network models, making it a lightweight model.
[0069] Step 204: export the model parameters of the face liveness detection model and deploy them to the embedded device. Use C language to execute the network structure and convolution operations between network layers, perform the forward reasoning process, and obtain the face liveness detection model.
[0070] The process specifically includes: a. Saving the model's weight parameters in the form of a TXT file.
[0071] b. Use C language to read the saved weight parameter file, import it into the embedded end, and use C language to build the network structure of the face liveness detection model and the convolution operation between the network layers, perform the forward reasoning process, and obtain the face liveness detection model.
[0072] Deploying an embedded device requires exporting the model's weight parameters and saving them as a TXT file. This file is then read using C language to access the model parameters on the embedded device. The network structure and convolution operations between layers are implemented in C language for the complete model's forward inference process. The network ultimately outputs a 1*1*2 tensor, representing the probabilities of live and non-live objects. The difference between these two probabilities is then used to output the final recognition result.
[0073] In summary, this application improves the backbone network based on the original MobileFaceNet, on the one hand reducing the number of Bottlenecks, and on the other hand reducing the number of channels of each convolutional block network and the number of Bottleneck channels. For the original 6*6 global depth convolution block, this solution splits it into two 3*3 depth convolutions, reducing the amount of computation and parameters, while also improving feature extraction capabilities. The improved face liveness detection model has greatly reduced memory, and when deployed on embedded devices, it uses TXT text export and C language format forwarding, without the assistance of additional hardware acceleration, it can complete face liveness detection under low latency requirements while maintaining high accuracy.
[0074] The above describes the preferred embodiments of the present invention; it should be understood that the present invention is not limited to the above-mentioned specific embodiments, and the devices and structures not described in detail should be understood to be implemented in a common manner in the art; any technician familiar with the art can make many possible changes and modifications without departing from the technical solution of the present invention, or modify them into equivalent embodiments with equivalent changes, which does not affect the essential content of the present invention; therefore, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention that do not depart from the content of the technical solution of the present invention are still within the scope of protection of the technical solution of the present invention.
Claims
1. A face liveness detection method based on MobileFaceNet, characterized in that: The method comprises: The dataset containing real and fake face photos is divided into training, test, and validation sets according to the proportions, and the faces of the photos are annotated to obtain cropped face images. Construct a backbone feature extraction network for a face liveness detection model and perform iterative training using a processed data set; the backbone feature extraction network is used to extract image features and includes three common convolution blocks CW, one depthwise convolution block DW, nine bottleneck layer BK modules, and two pointwise convolution blocks PW; the first common convolution block and the depthwise convolution block are connected before the nine bottleneck layer modules, and the first pointwise convolution block, the second common convolution block, the third common convolution block, and the second pointwise convolution block are sequentially cascaded after the nine bottleneck layer modules; Constructing a classifier network of the face liveness detection model and performing iterative training using the dataset and the trained backbone feature extraction network; the classifier network is used to make judgments based on the extracted image features to obtain liveness probability and non-liveness probability; The model parameters of the face liveness detection model are exported and deployed to an embedded device. The network structure and convolution operations between network layers are executed using C language, and a forward reasoning process is performed to obtain the face liveness detection model.
2. The method according to claim 1, characterized in that The data set containing real face and fake face photos is divided into a training set, a test set, and a validation set according to the proportion, and the faces of the photos are annotated to obtain cropped face images, including: The dataset is divided into a training set, a validation set, and a test set in a ratio of 7:1:2; Detecting the face images in the dataset and returning the coordinate information of the face frame in the image; The coordinates of the face frame are doubled in all directions according to the coordinate information, and the face is cropped and extracted according to the face frame to obtain a cropped face image.
3. The method according to claim 2, characterized in that The backbone feature extraction network inputs a face image of size 96*96; the output is a 128-dimensional feature vector extracted from the face image; The convolution kernel size of the first ordinary convolution block is 3*3, the step size s=3, and the number of output channels is 2; the convolution kernel size of the depthwise convolution block is 3*3, the step size s=1, and the number of output channels is 2; The nine BK modules are divided into three groups. The number of channels in the first two groups is 4. A direct connection shortcut structure is set between the BK modules in each group. The step size of the first BK module in each group is s=2, and the step size of the second and third BK modules is s=1. The number of channels in the third group is 8. A shortcut structure is set between the BK modules in the group. The step size of the first BK module is s=2, and the step size of the second and third BK modules is s=1. The convolution kernel of the first point-by-point convolution block is 1*1, the step size s=1, and the number of output channels is 8; the convolution kernel of the second ordinary convolution block is 3*3, the step size s=2, and the number of output channels is 8; the convolution kernel of the third ordinary convolution block is 3*3, the step size s=1, and the number of output channels is 8; the convolution kernel of the point-by-point convolution block is 1*1, the step size s=2, and the number of output channels is 128.
4. The method according to claim 3, characterized in that The training cycle of the backbone feature extraction network is 90 epochs, the batch size is set to 128, and the initial learning rate of the SGD optimizer is 0.01; When epochs reaches 30, the learning rate is adjusted to 0.001, and when epochs reaches 60, the learning rate is adjusted to 0.0001.
5. The method according to claim 2, characterized in that The input of the classifier network is the 128-dimensional feature vector output by the backbone feature extraction network, including a fourth ordinary convolution block and a fifth ordinary convolution block; the convolution kernel of the fourth ordinary convolution block is 3*3, the step size s=1, and the number of output channels is 32; the convolution kernel of the fifth ordinary convolution block is 3*3, the step size s=1, and the number of output channels is 2, and the two output channels represent the probability of liveness and the probability of non-liveness, respectively.
6. The method according to claim 5, characterized in that The training cycle of the classifier network is 90 epochs, the batch size is set to 128, and the initial learning rate of the SGD optimizer is 0.01; When epochs reaches 30, the learning rate is adjusted to 0.001, and when epochs reaches 60, the learning rate is adjusted to 0.0001.
7. The method according to any one of claims 4 or 6, characterized in that: Set the label of the live face image to 1 and the label of the fake face image to 0; Sending the label and the face image into the face liveness detection model to respectively train the backbone feature extraction network and the classifier network, and adjusting the model batch size, learning rate and gradient parameters according to the liveness probability and the non-liveness probability; When the probability of being alive minus the probability of being non-alive output by the classifier network is greater than 0.2, the output prediction result is alive, otherwise it is non-alive.
8. The method according to claim 7, characterized in that The first loss function L1 of the backbone feature extraction network is expressed as: Among them, the front part of L1 is SoftMax Loss, the back part is Triplet Loss, D(a,n) represents the distance between the sample to be predicted and the negative sample, D(a,p) represents the distance between the sample to be predicted and the positive sample, m is 0.2, and σ is 0.3; The second loss function L2 of the classifier network is expressed as: Among them, y i is the label value, a one-hot encoding consisting of 0 and 1, y i ' is the predicted value, y i '∈(0,1).
9. The method according to claim 4, characterized in that The model parameters of the face liveness detection model are exported and deployed to the embedded device, the network structure and the convolution operations between the network layers are executed by C language, and the forward reasoning process is performed to obtain the face liveness detection model; include: Save the model's weight parameters in the form of a TXT file; The saved weight parameter file is read using C language, imported into the embedded end, and the network structure of the face liveness detection model and the convolution operation between the network layers are constructed using C language, and the forward reasoning process is performed to obtain the face liveness detection model.
Citation Information
Patent Citations
Human face living body detection method and device, computer equipment and storage medium
CN111191521A
Face living body detection method and system, electronic equipment and storage medium
CN115376183A