Face living body detection method, electronic equipment and storage medium
By acquiring and analyzing pseudo-depth information of visible light and near-red images, differential thermal images are generated to judge the living state of the face, which solves the problem that users need to cooperate in performing actions in the prior art, and improves detection accuracy and user experience.
Patent Information
- Application Number
- CN202510639161.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-05-19
Smart Images

Figure CN120164264A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of face liveness detection, and particularly to a face liveness detection method, an electronic device, and a storage medium. Background Art
[0002] With the development of electronic technology, face recognition has become a potential biometric authentication method. While greatly improving people's living convenience, the security issue of face recognition technology has gradually attracted people's attention. To prevent people from cracking face recognition through photos (i.e., aligning a face photo with the camera of the face recognition system for face recognition), face liveness detection technology has also become a research hotspot. The purpose of face liveness detection is to distinguish whether an image is obtained by photographing a real person, and it is a basic component of the face recognition system.
[0003] In related face liveness detection technologies, the detection device will issue random action instructions, such as blinking, shaking the head, or even reading out a string of random numbers, etc., and then detect whether the user's response meets the expectation through the camera. Obviously, this requires the user to cooperate to perform corresponding actions to complete, with a long detection time and poor user experience.
[0004] In view of this, this application is specifically proposed. Summary of the Invention
[0005] This application aims to provide a face liveness detection method, an electronic device, and a storage medium, which can accurately identify whether the currently recognized face is a live face without the user's cooperation to perform corresponding actions.
[0006] In a first aspect, an embodiment of this application provides a face liveness detection method, including: Obtain a visible light image and a near-infrared image of the same face object; Determine a visible light pseudo-depth image corresponding to the visible light image, and a near-infrared pseudo-depth image corresponding to the near-infrared image; Generate a difference heat map according to the difference between the visible light pseudo-depth image and the near-infrared pseudo-depth image; Determine whether the face object is a live body according to the difference heat map.
[0007] According to the technical solution provided by the embodiment of this application, optionally, the generating a difference heat map according to the difference between the visible light pseudo-depth image and the near-infrared pseudo-depth image includes: Determine a visible light feature map of the visible light pseudo-depth image and a near-infrared feature map of the near-infrared pseudo-depth image based on a trained comparison network; Determine the difference between the visible light feature map and the near-infrared feature map; Generate a differential heat map based on the difference between the visible light feature map and the near-infrared feature map.
[0008] According to the technical solution provided by the embodiment of the present application, optionally, the trained contrast network includes a plurality of cascaded feature extraction layers in sequence. Determining the visible light feature map of the visible light pseudo-depth image and the near-infrared feature map of the near-infrared pseudo-depth image based on the trained contrast network includes: Input the visible light pseudo-depth image into the trained contrast network to obtain first feature sub-maps respectively output by each of the feature extraction layers. Through a Feature Pyramid Network (FPN), fuse the first feature sub-maps respectively output by each of the feature extraction layers to obtain the visible light feature map of the visible light pseudo-depth image; Input the near-infrared pseudo-depth image into the trained contrast network to obtain second feature sub-maps respectively output by each of the feature extraction layers; through a Feature Pyramid Network (FPN), fuse the second feature sub-maps respectively output by each of the feature extraction layers to obtain the near-infrared feature map of the near-infrared pseudo-depth image.
[0009] According to the technical solution provided by the embodiment of the present application, optionally, the trained contrast network is trained in the following manner: Input the visible light pseudo-depth image and the near-infrared pseudo-depth image in the normal sample group into the contrast network to be trained in sequence. Through the last-level feature extraction layer of the contrast network, output the visible light feature vector corresponding to the visible light pseudo-depth image and the near-infrared feature vector corresponding to the near-infrared pseudo-depth image; Input the visible light feature vector and the near-infrared feature vector into two different branch networks of an asymmetric contrast siamese network, and respectively obtain a first predicted value corresponding to the visible light pseudo-depth image and a second predicted value corresponding to the near-infrared pseudo-depth image through the two different branch networks; Calculate the loss value between the first predicted value and the second predicted value, and use the zero loss value as the training target to adjust the parameters of each feature extraction layer in the contrast network to be trained; Among them, the feature extraction layer includes a convolutional network module and a spatial transformation network module; the normal sample group refers to a combination of a visible light pseudo-depth image and a near-infrared pseudo-depth image of a living face.
[0010] According to the technical solution provided by the embodiment of the present application, optionally, determining the difference between the visible light feature map and the near-infrared feature map includes: Extract the first feature vector corresponding to each coordinate position in the visible light feature map, and extract the second feature vector corresponding to each coordinate position in the near-infrared feature map; Form a first eigenvector matrix by the first eigenvectors corresponding to all coordinate positions in the visible light feature map; form a second eigenvector matrix by the second eigenvectors corresponding to all coordinate positions in the near-infrared feature map, and calculate the covariance matrix between the first eigenvector matrix and the second eigenvector matrix; Calculate the Mahalanobis distance between the visible light feature map and the near-infrared feature map at each coordinate position according to the inverse matrix of the covariance matrix, the first eigenvector and the second eigenvector corresponding to each coordinate position; Determine the Mahalanobis distance at each coordinate position as the difference between the visible light feature map and the near-infrared feature map.
[0011] According to the technical solution provided by the embodiment of the present application, optionally, the generating a difference heat map according to the difference between the visible light pseudo-depth image and the near-infrared pseudo-depth image includes: Map the Mahalanobis distance to the color value of the difference heat map at the coordinate position.
[0012] According to the technical solution provided by the embodiment of the present application, optionally, the determining whether the face object is a live body according to the difference heat map includes: Convert the difference heat map into a grayscale image; Calculate the average value of the grayscale values of each pixel in the grayscale image; If the average value is less than the threshold, determine that the face object is a live body.
[0013] According to the technical solution provided by the embodiment of the present application, optionally, the determining the visible light pseudo-depth image corresponding to the visible light image and the near-infrared pseudo-depth image corresponding to the near-infrared image includes: Input the visible light image into a trained first feature extraction network to obtain the visible light pseudo-depth image; Input the near-infrared image into a trained second feature extraction network to obtain the near-infrared pseudo-depth image; Wherein, the training sample label of the first feature extraction network is the pseudo-depth image corresponding to the visible light image, and the training sample label of the second feature extraction network is the pseudo-depth image corresponding to the near-infrared image.
[0014] In a second aspect, an embodiment of the present application further provides an electronic device, and the electronic device includes: A processor and a memory; The processor is configured to execute the steps of the face liveness detection method according to any one of the embodiments by calling a program or an instruction stored in the memory.
[0015] In a third aspect, an embodiment of the present application further provides a computer-readable storage medium, which stores a program or instructions, and the program or instructions cause a computer to execute the steps of the face liveness detection method as described in any one of the embodiments.
[0016] In summary, the present application proposes a face liveness detection method, which acquires visible light images and near-infrared images of the same face object; determines a visible light pseudo-depth image corresponding to the visible light image and a near-infrared pseudo-depth image corresponding to the near-infrared image; generates a difference heat map according to the difference between the visible light pseudo-depth image and the near-infrared pseudo-depth image; determines whether the face object is a live body according to the difference heat map, and can accurately identify whether the currently identified face is a live face without the user's cooperation to perform corresponding actions, improving the intelligence level and accuracy of the detection and enhancing the user experience. Specifically, the difference between a live face and a non-live face (such as a printed photo, an electronic screen, a 3D face model) is reflected in the texture of local details, and the pseudo-depth image can combine the local detail texture of the image to generate real depth information, and use information such as color, texture, and the geometric shape of an object to estimate the relative depth of the object through calculation or reasoning, so as to create a pseudo-depth image with a sense of depth. At the same time, there are differences in the imaging principles between visible light images and near-infrared images. The difference in imaging principles leads to different imaging results for non-live faces, and there are obvious differences between the imaging results. In particular, due to its unique imaging principle, the near-infrared camera cannot image the face displayed on an electronic device and has natural defenses. For a printed face, there are texture differences between its visible light image and near-infrared image. Based on this, in the solution of the present application, a difference heat map is generated according to the difference between the visible light pseudo-depth image and the near-infrared pseudo-depth image, and whether the face object is a live body is determined according to the difference heat map, which can greatly improve the final detection accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is a flowchart of a face liveness detection method provided by an embodiment of the present application; Figure 2 is a schematic diagram of a depthwise separable convolution structure provided by an embodiment of the present application; Figure 3 is a schematic diagram of the structure of a trained contrast network provided by an embodiment of the present application; Figure 4 is a schematic diagram when training the contrast network provided by an embodiment of the present application; Figure 5 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] The present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the relevant invention and are not intended to limit the invention. Additionally, it should be noted that for ease of description, only the parts related to the invention are shown in the drawings.
[0019] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the drawings and embodiments.
[0020] The Driver Monitor System (DMS) is an important component in the intelligent cockpit and is of great significance for commercial vehicles and passenger cars. The driver monitoring system can be roughly divided into functions such as camera occlusion, face occlusion, dangerous behavior, fatigue monitoring, distraction monitoring, and face identification FACE ID. Face liveness detection is a key link in face recognition of the FACE ID function and can effectively resist attacks by fraud means such as photos, videos, and high-fidelity masks.
[0021] At present, the mainstream liveness detection technologies include interactive cooperative liveness detection technology, silent liveness detection technology, and 3D liveness detection technology. Among them, the interactive cooperative liveness detection technology requires users to perform specific actions (such as blinking, opening the mouth, shaking the head, etc.) according to the system prompts, which requires a high degree of user participation and cooperation and does not have a good user experience. The 3D liveness detection technology uses devices such as three-dimensional depth cameras or structured light cameras to obtain three-dimensional information of the face, and the device cost is relatively high. The silent liveness detection technology does not require users to cooperate to perform specific actions. It monitors by capturing the natural behaviors or physiological characteristics of users, has a good user experience and a low hardware cost. The current mainstream method is to use a single network model, collect real face images of live bodies and fraudulent face images of non-live bodies as the training data set. However, the acquisition of fraudulent face images of non-live bodies often requires a large amount of manpower and material resources, and the trained network model has poor generalization in various scenarios. When facing a new application scenario, it is necessary to collect new images in the new application scenario again to increment the training data set and iterate the previously trained network model to make it adapt to the new scenario.
[0022] In view of the above-mentioned defects existing in the prior art, the present application proposes a liveness detection method based on a multi-modal unsupervised contrast network. This method can be applied in the intelligent cockpit scenario and can also be applied in other face detection scenarios. This method belongs to the category of silent liveness detection and does not require users to cooperate to complete specific actions. Therefore, it has a good user experience and does not need to rely on special hardware devices. Therefore, the hardware cost is relatively low, and it has high detection accuracy and low implementation difficulty.
[0023] Figure 1The figure is a flowchart of a face liveness detection method provided by an embodiment of the present application. Refer to Figure 1 , the face liveness detection method specifically includes the following steps: S110. Obtain a visible light image and a near-infrared image of the same face object.
[0024] Among them, the visible light image and the near-infrared image are different types of images of the same face object from the same perspective. The visible light image is also called an RGB image, which is obtained by imaging with red, green, and blue color components. The near-infrared image is obtained by reflecting near-infrared light. Therefore, there are differences in the imaging principles between the visible light image and the near-infrared image. The differences in the imaging principles lead to different imaging results for non-live faces, and there are obvious differences between the imaging results.
[0025] In particular, due to its unique imaging principle, the near-infrared camera cannot image the face displayed on the electronic device and has natural defenses. For a printed face, there are texture differences between its visible light image and near-infrared image. Based on this, the solution of the present application uses multi-modal input data (specifically referring to visible light images and near-infrared images) to implement face liveness detection, improving the detection accuracy and reducing the detection difficulty.
[0026] S120. Determine the visible light pseudo-depth image corresponding to the visible light image and the near-infrared pseudo-depth image corresponding to the near-infrared image.
[0027] Among them, the pseudo-depth image is an image that simulates real depth information generated through a specific algorithm or technology. By using information such as color, texture, and the geometric shape of an object, the relative depth of the object is estimated through calculation or reasoning, so as to create a pseudo-depth image with a sense of depth.
[0028] Exemplarily, input the visible light image into a trained first feature extraction network to obtain the visible light pseudo-depth image; input the near-infrared image into a trained second feature extraction network to obtain the near-infrared pseudo-depth image; among them, the training sample label of the first feature extraction network is the pseudo-depth image corresponding to the visible light image, and the training sample label of the second feature extraction network is the pseudo-depth image corresponding to the near-infrared image. Exemplarily, both the first feature extraction network and the second feature extraction network can be convolutional neural networks.
[0029] Using the pseudo-depth map as the label for model training enables the model to extract the texture difference features between the visible light image and the near-infrared image, thereby training a multi-modal feature extraction model. There are differences between the visible light pseudo-depth image and the near-infrared pseudo-depth image inferred by the model.
[0030] The difference between living and non-living faces (such as printed photos, electronic screens, and 3D face models) lies in the texture of local details. By using pseudo-depth images rather than conventional classification labels (living and non-living) as real training labels, the local texture features of the image are greatly utilized to distinguish the detailed differences between living and non-living faces, thereby improving the final detection accuracy.
[0031] In the smart cockpit application scenario, during the model training phase, in order to adapt to the scenario where the computing power of the vehicle is limited, the backbone feature network adopts a deep separable convolutional network structure, such as the MobileNet series. Figure 2 As shown, it is a depth-wise separable convolution structure, which exemplarily includes 3 input channels and 4 convolution kernels. After the image of each input channel is detected by the convolution kernel, the corresponding feature map is obtained.
[0032] By separating the convolution operations using DepthWise and PointWise, the model parameters can be reduced to adapt to scenarios where the computing power of the vehicle is limited.
[0033] Among them, the loss Loss between the model output result and the label can be determined by the following calculation formula: Loss = LMSE + LCDL Among them, LMSE and LCDL are both mean square error losses, LMSE is the category loss (category refers to living or non-living) and LCDL is the depth loss of the pseudo depth map.
[0034] Optionally, the visible light pseudo depth image and near infrared pseudo depth image in the training data set can be automatically generated by the existing PRNET network, rather than being obtained by expensive 3D scanning equipment, which reduces the difficulty and cost of obtaining training samples. The PRNET network can directly predict the dense 3D face geometry (including feature information such as depth, normal, vertex coordinates, etc.) from a single 2D face image, and map the visible light image and near infrared image to a pseudo depth image of a unified form through a feature extraction network. The form of expression here refers to the visible light image as an image of three color channels of red, green, and blue, while the near infrared image is a single-channel grayscale image. Therefore, the two have different forms of expression and are different types of input for the model. Therefore, different types of input data need to be mapped to the same type in order to reduce the disturbance caused by the morphological differences of the input data to the subsequent model accuracy.
[0035] The multimodal visible light images and near-infrared images are mapped into pseudo depth maps, which are used as the input of the subsequent comparison network, rather than directly using the multimodal visible light images and near-infrared images as the input of the comparison network. This achieves the purpose of unifying the feature representation of multimodal data and reduces the accuracy disturbance of the comparison network caused by data morphology differences.
[0036] S130. Generate a differential heat map based on the difference between the visible light pseudo-depth image and the near-infrared pseudo-depth image.
[0037] Exemplarily, determine the visible light feature map of the visible light pseudo-depth image and the near-infrared feature map of the near-infrared pseudo-depth image based on the trained contrast network; determine the difference between the visible light feature map and the near-infrared feature map; generate a differential heat map according to the difference between the visible light feature map and the near-infrared feature map. That is, further extract features from the visible light pseudo-depth image and the near-infrared pseudo-depth image respectively, and compare the extracted features. If the quantization value of the feature difference between the visible light pseudo-depth image and the near-infrared pseudo-depth image exceeds a preset value, the detection result is considered non-living, otherwise the detection result is considered living.
[0038] As Figure 3 shown, the trained contrast network 300 includes a plurality of cascaded feature extraction layers 310 in sequence. Determining the visible light feature map of the visible light pseudo-depth image and the near-infrared feature map of the near-infrared pseudo-depth image based on the trained contrast network includes: Input the visible light pseudo-depth image into the trained contrast network to obtain first feature sub-maps respectively output by each of the feature extraction layers, and fuse the first feature sub-maps respectively output by each of the feature extraction layers through a Feature Pyramid Network (FPN) to obtain the visible light feature map of the visible light pseudo-depth image; input the near-infrared pseudo-depth image into the trained contrast network to obtain second feature sub-maps respectively output by each of the feature extraction layers; fuse the second feature sub-maps respectively output by each of the feature extraction layers through a Feature Pyramid Network (FPN) to obtain the near-infrared feature map of the near-infrared pseudo-depth image.
[0039] Among them, the feature extraction layer includes a convolutional network module (i.e., Figure 3 the "C" in Figure 3 ) and a Spatial Transformer Network (STN) module (i.e., the "S" in
[0040] ). It combines the powerful feature extraction ability of convolution and the flexible spatial invariance of the spatial transformer network module, improving the generalization of the feature extraction layer. STN is a spatial transformer network, a plug-and-play module in deep learning networks, which can endow the network with spatial invariance and has strong robustness to transformations such as translation, rotation, and scaling. In the first step, feature sub - graphs are collected through a bottom - up path, which is the forward propagation process of a traditional convolutional neural network. In this process, as the number of network layers deepens, the resolution of the feature sub - graphs gradually decreases, but the semantic information of the features gradually increases. The feature sub - graphs output by convolutional layers at different levels are used as the input of the FPN. In the second step, feature sub - graphs are continuously collected through a top - down path. Starting from the highest layer (the feature sub - graph output by the highest layer has the richest semantic information but the lowest resolution), an up - sampling operation is performed on it to increase its resolution to the same resolution as the feature sub - graph of the next layer. In the third step, the up - sampled feature sub - graph in the second step is fused with the corresponding - level feature sub - graph in the first step at the same resolution. Specifically, the feature sub - graphs can be added element - by - element. The purpose of doing this is to make the fused feature sub - graph include both rich semantic information and a high resolution. In the fourth step, a convolutional operation is performed on the fused feature sub - graph to adjust the number of channels of the feature sub - graph.
[0041] The differences between the visible - light pseudo - depth map and the near - infrared pseudo - depth map are compared through the idea of image comparison to obtain the difference heat map of the visible - light pseudo - depth map and the near - infrared pseudo - depth map. Based on this difference heat map, the live state of the face is determined. Among them, the contrast network is an unsupervised network. The unsupervised idea mainly lies in using a general image contrast network for zero - sample transfer to any other application scenario (such as the cockpit scenario) for face liveness detection. An unsupervised network means that it does not require an iterative training of the network using a face image data set in a specific application scenario (such as the cockpit scenario). The contrast network can be trained only with normal samples. It calculates the difference in a comparative way combined with the Mahalanobis distance, rather than simply using the main - feature extraction network to map the difference. Therefore, the main - feature extraction network of the face pseudo - depth map in a specific application scenario (such as the cockpit scenario) does not need to be specifically trained in the cockpit scenario. It can use a mainstream feature extraction network for zero - sample transfer to the cockpit scenario, and then fuse the sub - feature maps at multiple scales, calculate the Mahalanobis distance of the feature maps corresponding to the visible - light pseudo - depth map and the near - infrared pseudo - depth map respectively, and determine the live state of the face by calculating the difference between the two types of pseudo - depth maps. The contrast network can be trained with normal samples in any scenario, can adapt to any other scenario including the cockpit, has excellent scene - transfer generalization, does not need to deliberately obtain face data and non - live data in the cockpit environment, avoids the human and material input for obtaining the training data set of the contrast network, and reduces the difficulty of the implementation scheme of this application.
[0042] The concept of "normal" in normal samples comes from the two states of normal and abnormal samples in anomaly detection. In the embodiment of this application, normal samples refer to the pseudo - depth images of live human faces. Therefore, the construction of the training data set only requires normal samples and does not require abnormal samples, avoiding the time cost and economic cost required to obtain or collect abnormal samples.
[0043] Reference Figure 4 As shown, the contrast network is trained as follows: The visible light pseudo-depth image and the near-infrared pseudo-depth image in the normal sample group are sequentially input into the contrast network 410 to be trained. The visible light feature vector corresponding to the visible light pseudo-depth image and the near-infrared feature vector corresponding to the near-infrared pseudo-depth image are output through the last-level feature extraction layer of the contrast network; The visible light feature vector and the near-infrared feature vector are input into two different branch networks of the asymmetric contrast siamese network 420. The first prediction value Pa corresponding to the visible light pseudo-depth image and the second prediction value Za corresponding to the near-infrared pseudo-depth image are obtained through the two different branch networks respectively; calculate the loss value between the first prediction value Pa and the second prediction value Za, and use the loss value being zero as the training objective to adjust the parameters of each feature extraction layer in the contrast network to be trained; the normal sample group refers to a combination of a visible light pseudo-depth image and a near-infrared pseudo-depth image including a live human face. The asymmetric contrast siamese network can enhance data adaptability, extract complementary features, support multimodal data fusion, etc. Specifically, the asymmetric contrast siamese network can adopt different data augmentation strategies, different network structures or parameter settings for the input siamese samples, so as to be able to extract features from different angles for the input data and learn different levels and types of features. Among them, P and Z represent different projection modules, that is, different operations are performed on the input siamese samples (in this embodiment, referring to the visible light pseudo-depth image and the near-infrared pseudo-depth image), reflecting the "asymmetric" characteristic. E represents the encoding module shared by the siamese samples, usually composed of a convolutional neural network, a recurrent neural network, etc. Loss represents the loss calculation module. After the encoding module E transforms the feature vector output by the last-level feature extraction layer into a high-dimensional space, and then through the mapping of the mapping module P or Z, the first prediction value Pa and the second prediction value Za are obtained, and then the loss value between the first prediction value Pa and the second prediction value Za is calculated. Using the loss value being zero as the training objective, the parameters of each feature extraction layer in the contrast network to be trained are adjusted. Among them, a gradient truncation Stop Gradient operation can be introduced to prevent the gradient from backpropagating along a specific path, thereby controlling the parameter update path, preventing the parameters of certain modules from being updated during a specific training process, helping the model to learn more stably, and avoiding the problems of gradient explosion or gradient disappearance.
[0044] In some embodiments, determining the difference between the visible light feature map and the near-infrared feature map includes: Extract the first feature vector corresponding to each coordinate position in the visible light feature map, and extract the second feature vector corresponding to each coordinate position in the near-infrared feature map; form the first feature vector matrix by the first feature vectors corresponding to all coordinate positions in the visible light feature map; form the second feature vector matrix by the second feature vectors corresponding to all coordinate positions in the near-infrared feature map, and calculate the covariance matrix between the first feature vector matrix and the second feature vector matrix; calculate the Mahalanobis distance between the visible light feature map and the near-infrared feature map at the one coordinate position according to the inverse matrix of the covariance matrix, the first feature vector and the second feature vector corresponding to the one coordinate position; determine the Mahalanobis distance as the difference between the visible light feature map and the near-infrared feature map at the one coordinate position. Specifically, form the first feature vector matrix A = [a ij by the first feature vectors corresponding to all coordinate positions in the visible light feature map, where a ij represents the first feature vector corresponding to the coordinate position (i, j) in the visible light feature map. Form the second feature vector matrix B = [b ij by the second feature vectors corresponding to all coordinate positions in the near-infrared feature map, where b ij represents the second feature vector corresponding to the coordinate position (i, j) in the near-infrared feature map. Then calculate the covariance matrix of the first feature vector matrix and the second feature vector matrix.
[0045] Generating a difference heat map according to the difference between the visible light pseudo-depth image and the near-infrared pseudo-depth image includes: Map the Mahalanobis distance to the color value of the difference heat map at the one coordinate position.
[0046] S140. Determine whether the face object is a live body according to the difference heat map.
[0047] Specifically, the difference heat map can be converted into a grayscale image; calculate the average value of the gray values of each pixel in the grayscale image; if the average value is less than a threshold (the threshold can be 0.5, for example), then determine that the face object is a live body.
[0048] The face liveness detection method provided by the embodiments of the present application uses the pseudo-depth map as a label to train a multi-modal feature extraction network. Compared with the traditional face liveness detection method that uses a relatively simple binary classification model to map the entire image to a single classification probability and has difficulty in distinguishing the local texture information of the face, the solution of the present application can better utilize the local detail texture features of the image by using the pseudo-depth map as a label, more accurately represent the local texture information of the face, and better distinguish the image features of live bodies and non-live bodies.
[0049] By adopting the PNET network, a pseudo-depth map for training labels can be generated from optical images and near-infrared images, avoiding the need to scan a large number of real human faces with a 3D scanning device to generate depth maps, which greatly saves the economic cost of producing training set labels.
[0050] The idea of unsupervised image contrast is introduced into the field of live detection. The multi-modal optical images and near-infrared images are mapped into pseudo-depth maps, and the pseudo-depth maps are used as the input of the contrast network instead of directly using the multi-modal optical images and near-infrared images as the input of the contrast network, achieving the purpose of unifying the feature representation forms of multi-modal data and reducing the accuracy perturbation caused by data form differences to the contrast network.
[0051] By introducing the image contrast idea, the contrast network only needs normal samples for training, without the need to collect abnormal samples, avoiding a series of device investments and labor inputs such as the hardware device construction of the abnormal sample data set, image acquisition, selection, annotation, correction, and balanced labeling from the network design level. After the performance forms of the multi-modal input data are aligned, the contrast network can be applied to any scenario. Face live detection in the cockpit scenario is just one of the scenarios, truly achieving the effect of zero-sample adaptation to any scenario.
[0052] The Mahalanobis distance is used instead of the conventional Euclidean distance calculated by eigenvalues as the distance calculation method between the feature maps of the visible light pseudo-depth map and the near-infrared pseudo-depth map. Considering the relationship between the single-point difference and the overall distribution in the feature map and excluding the correlation interference between points, the difference calculation performance is better.
[0053] Figure 5 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 5 shown, the electronic device 500 includes one or more processors 501 and a memory 502.
[0054] The processor 501 can be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and can control other components in the electronic device 500 to perform desired functions.
[0055] The memory 502 may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage media, and the processor 501 may run the program instructions to implement the face liveness detection method of any embodiment of the present application described above and / or other desired functions. Various contents such as initial external parameters, thresholds, etc. may also be stored in the computer-readable storage media.
[0056] In one example, the electronic device 500 may further include: an input device 503 and an output device 504, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown). The input device 503 may include, for example, a keyboard, a mouse, etc. The output device 504 may output various information to the outside, including warning prompt information, braking force, etc. The output device 504 may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.
[0057] Of course, for simplicity, Figure 5 only some of the components related to the present application in the electronic device 500 are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, according to specific application scenarios, the electronic device 500 may further include any other appropriate components.
[0058] In addition to the above methods and devices, an embodiment of the present application may also be a computer program product, which includes computer program instructions, and when the computer program instructions are run by a processor, the processor is caused to execute the steps of the face liveness detection method provided by any embodiment of the present application.
[0059] The computer program product may be written in any combination of one or more programming languages to write program code for performing the operations of the embodiments of the present application. The programming languages include object-oriented programming languages, such as Java, C++, etc., and also include conventional procedural programming languages, such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as an independent software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0060] In addition, an embodiment of the present application may also be a computer-readable storage medium storing computer program instructions, which, when run by a processor, cause the processor to execute the steps of the face liveness detection method provided by any embodiment of the present application.
[0061] The computer-readable storage medium may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, include but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0062] It should be noted that the terms used in the present application are only for describing specific embodiments and do not limit the scope of the present application. As shown in the specification and claims of the present application, unless the context clearly indicates otherwise, words such as "a", "an", "one", and / or "the" are not specifically singular and may also include the plural. The term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, or device including the element.
[0063] It should also be noted that the orientation or positional relationship indicated by terms such as "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present application. Unless otherwise clearly specified and limited, terms such as "installed", "connected", "coupled" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the internal communication of two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances.
[0064] In this text, specific examples are used to elaborate on the principles and implementation manners of this application. The descriptions of the above embodiments are only used to help understand the method and its core idea of this application. The above is only the preferred implementation manner of this application. It should be noted that due to the limited nature of literal expression and objectively infinite specific structures, for those of ordinary skill in the art, without departing from the principles of this invention, several improvements, refinements or changes can be made, or the above technical features can be combined in an appropriate manner; these improvements, refinements, changes or combinations, or directly applying the concept and technical solution of the invention to other occasions without improvement, shall all be regarded as the protection scope of this application.
Claims
1. A method for detecting liveness of a face, characterized in that: include: Acquire a visible light image and a near infrared image of the same face object; Determine a visible light pseudo depth image corresponding to the visible light image, and a near infrared pseudo depth image corresponding to the near infrared image; generating a difference thermal image according to a difference between the visible light pseudo depth image and the near infrared pseudo depth image; Determine whether the human face object is a living body according to the difference thermal image.
2. The method according to claim 1, characterized in that The step of generating a difference thermal image according to a difference between the visible light pseudo depth image and the near infrared pseudo depth image comprises: Determine a visible light feature map of the visible light pseudo depth image and a near infrared feature map of the near infrared pseudo depth image based on the trained contrast network; determining a difference between the visible light characteristic pattern and the near infrared characteristic pattern; A difference thermal image is generated according to the difference between the visible light characteristic map and the near infrared characteristic map.
3. The method according to claim 2, characterized in that The trained contrast network includes a plurality of sequentially cascaded feature extraction layers, and determining the visible light feature map of the visible light pseudo depth image and the near infrared feature map of the near infrared pseudo depth image based on the trained contrast network includes: Inputting the visible light pseudo depth image into the trained contrast network to obtain the first feature sub-graphs respectively output by the feature extraction layers, and fusing the first feature sub-graphs respectively output by the feature extraction layers through a feature pyramid FPN to obtain a visible light feature graph of the visible light pseudo depth image; The near-infrared pseudo depth image is input into the trained contrast network to obtain the second feature sub-graphs output by each feature extraction layer; the second feature sub-graphs output by each feature extraction layer are fused through the feature pyramid FPN to obtain the near-infrared feature graph of the near-infrared pseudo depth image.
4. The method according to claim 3, characterized in that: The contrast network is trained as follows: The visible light pseudo depth image and the near infrared pseudo depth image in the normal sample group are sequentially input into the contrast network to be trained, and the visible light feature vector corresponding to the visible light pseudo depth image and the near infrared feature vector corresponding to the near infrared pseudo depth image are output through the last level feature extraction layer of the contrast network; Inputting the visible light feature vector and the near infrared feature vector into two different branch networks of the asymmetric contrast twin network, and obtaining a first prediction value corresponding to the visible light pseudo depth image and a second prediction value corresponding to the near infrared pseudo depth image through the two different branch networks respectively; Calculating a loss value between the first prediction value and the second prediction value, taking the loss value as zero as a training target, and adjusting parameters of each feature extraction layer in the comparison network to be trained; Among them, the feature extraction layer includes a convolutional network module and a spatial transformation network module; the normal sample group refers to a combination of visible light pseudo-depth images and near-infrared pseudo-depth images of living human faces.
5. The method according to claim 2, characterized in that: The determining the difference between the visible light characteristic graph and the near infrared characteristic graph comprises: Extracting a first eigenvector corresponding to each coordinate position in the visible light feature map, and extracting a second eigenvector corresponding to each coordinate position in the near infrared feature map; The first eigenvectors corresponding to all coordinate positions in the visible light feature map are combined into a first eigenvector matrix; the second eigenvectors corresponding to all coordinate positions in the near infrared feature map are combined into a second eigenvector matrix, and the covariance matrix between the first eigenvector matrix and the second eigenvector matrix is calculated; Calculate the Mahalanobis distance between the visible light feature map and the near infrared feature map at each coordinate position according to the inverse matrix of the covariance matrix and the first eigenvector and the second eigenvector corresponding to each coordinate position; The Mahalanobis distance at each coordinate position is determined as the difference between the visible light characteristic map and the near infrared characteristic map.
6. The method according to claim 5, characterized in that The step of generating a difference thermal image according to a difference between the visible light pseudo depth image and the near infrared pseudo depth image comprises: The Mahalanobis distance is mapped to a color value of the difference thermal image at the coordinate position.
7. The method according to claim 1, characterized in that The determining whether the face object is alive according to the difference thermal image comprises: Converting the difference thermal image into a grayscale image; Calculate the average value of the grayscale value of each pixel in the grayscale image; If the average value is less than a threshold value, it is determined that the face object is alive.
8. The method according to claim 1, characterized in that The determining of the visible light pseudo depth image corresponding to the visible light image and the near infrared pseudo depth image corresponding to the near infrared image includes: Inputting the visible light image into a trained first feature extraction network to obtain the visible light pseudo depth image; Inputting the near infrared image into a trained second feature extraction network to obtain the near infrared pseudo depth image; Among them, the training sample labels of the first feature extraction network are pseudo depth images corresponding to visible light images, and the training sample labels of the second feature extraction network are pseudo depth images corresponding to near infrared images.
9. An electronic device, characterized in that: The electronic device comprises: Processor and memory; The processor is used to execute the steps of the face liveness detection method according to any one of claims 1 to 8 by calling the program or instruction stored in the memory.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a program or instruction, and the program or instruction enables a computer to execute the steps of the face liveness detection method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Living body detection method and system
CN111259814A
Living body detection method and device, and electronic equipment
CN113792581A
Cross-modal living body fusion detection method and device, storage medium and computer equipment
CN117373138A
Cited By
Visible light and infrared image fused living body face detection method and system
CN121305696A