A method, device and electronic device for live detection

The generation of pseudo-depth maps through the depth prediction model solves the problem that the face recognition system is difficult to distinguish between real people and paper photos, realizes live detection without user cooperation, and improves system security and user experience.

CN113792581BActive Publication Date: 2025-06-17SHENZHEN YIXIN VISION TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110881598.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-02
Publication Date
2025-06-17
Estimated Expiration
2041-08-02

AI Technical Summary

Technical Problem

When performing facial recognition, it is difficult to distinguish between real people and photos on paper when existing facial recognition systems perform facial recognition, resulting in damage to security and poor user experience.

Method used

Using the depth prediction model, a pseudo-depth map is generated by obtaining the face image of the object to be detected. The pseudo-depth map reflects the depth relative relationship between multiple positions in the face, and then determines whether the object to be detected is a living body.

Benefits of technology

No need for users to cooperate with the specified actions to filter out paper attacks, improve user experience, and effectively distinguish between real people and photos, and enhance the security of the face recognition system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113792581B_ABST
    Figure CN113792581B_ABST
Patent Text Reader

Abstract

The live detection method provided by the embodiments of the present application pre - generates a depth prediction model by jointly using self - supervised learning and supervised learning. During detection, an image including the face of the object to be detected is obtained, and the image including the face of the object to be detected is input into the depth prediction model to obtain a pseudo - depth map corresponding to the image including the face of the object to be detected. According to the pseudo - depth map corresponding to the image including the face of the object to be detected, the detection result of the object to be detected is determined. In the embodiments of the present application, live detection is performed through the pseudo - depth map, which can effectively distinguish the differences between real people and photos, filter out paper attacks without the need for users to cooperate to complete specified actions, and can improve the user experience. In addition, jointly training the depth prediction model using self - supervised learning and supervised learning can improve the accuracy of the depth prediction model when predicting the pseudo - depth map, and further improve the accuracy of live judgment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of live detection, and particularly relates to a live detection method, apparatus, and electronic device. Background Art

[0002] Face recognition technology has currently been widely applied in various scenarios such as mobile payment, device unlocking, and access control management. Face recognition is based on the facial feature information of a person for identity recognition. Lawbreakers may use abnormal means to obtain the videos and photos of users, play the videos or photos using an electronic device, or print the photos on paper to impersonate the user for identity recognition and successfully pass the verification. Therefore, when performing face recognition, it is necessary to determine whether the recognized face is a live body or a video or photo to ensure the security of the face recognition system.

[0003] Playing videos or photos using an electronic device belongs to screen-type attacks, and printing photos on paper belongs to paper attacks. For screen-type attacks, such attacks can be filtered by setting the camera device of the face recognition system as a near-infrared camera.

[0004] For paper attacks, such attacks can be filtered by the user's cooperation to complete specified actions such as blinking, opening the mouth, shaking the head, etc. Since user cooperation is required, the user experience is not good. Summary of the Invention

[0005] In view of the above technical problems, embodiments of this application provide a live detection method, apparatus, and electronic device, which can filter out paper attacks without the need for the user to cooperate to complete specified actions when performing face recognition, and can improve the user experience.

[0006] In a first aspect, embodiments of this application provide a live detection method, which includes:

[0007] Obtain an image including the face of an object to be detected;

[0008] Input the image including the face of the object to be detected into a depth prediction model to obtain a pseudo-depth map corresponding to the image including the face of the object to be detected. The pseudo-depth map reflects the relative relationship in depth between multiple positions in the face. The depth prediction model is generated according to a first training set and a second training set. The first training set consists of multiple pairs of images including faces, and each pair of images including faces is obtained for the same object. The second training set consists of multiple images including faces and multiple pre-generated pseudo-depth maps;

[0009] Determine the detection result of the object to be detected according to the pseudo-depth map corresponding to the image including the face of the object to be detected.

[0010] In combination with the first aspect, in some implementations of the first aspect, the first training set consists of multiple pairs of images including faces, and each pair of images including faces is obtained for the same object at the same moment.

[0011] In combination with the first aspect and the above implementation, in some implementations of the first aspect, according to the pseudo-depth map corresponding to the image including the face of the object to be detected, the detection result of the object to be detected is determined, including: according to the pseudo-depth map corresponding to the image including the face of the object to be detected and the image including the face of the object to be detected, the detection result of the object to be detected is determined.

[0012] In combination with the first aspect and the above implementation, in some implementations of the first aspect, according to the pseudo-depth map corresponding to the image including the face of the object to be detected and the image including the face of the object to be detected, the detection result of the object to be detected is determined, including:

[0013] The pseudo-depth map corresponding to the image including the face of the object to be detected and the feature information related to the face in the image including the face of the object to be detected are respectively obtained, and the detection result of the object to be detected is determined according to the feature information related to the face.

[0014] In combination with the first aspect and the above implementation, in some implementations of the first aspect, the feature information related to the face includes at least one of global feature information and local feature information, the global feature information includes the feature information of the entire face area, and the local feature information includes the feature information of the local area in the face.

[0015] In combination with the first aspect and the above implementation, in some implementations of the first aspect, the image including the face of the object to be detected includes a first image and a second image, the first image and the second image are obtained based on a binocular camera, the pseudo-depth map corresponding to the image including the face of the object to be detected includes a first pseudo-depth map and a second pseudo-depth map, the first image corresponds to the first pseudo-depth map, and the second image corresponds to the second pseudo-depth map.

[0016] In a second aspect, an embodiment of the present application provides a method for training a model, and the method includes:

[0017] Obtain a first training set and a second training set, the first training set consists of multiple pairs of images including faces, and each pair of images including faces is obtained for the same object, and the second training set consists of multiple images including faces and multiple pre-generated pseudo-depth maps;

[0018] Train a depth prediction model according to the first training set and the second training set.

[0019] In combination with the second aspect, in some implementations of the second aspect, the first training set consists of multiple pairs of images including human faces, and each pair of images including human faces is obtained for the same object at the same moment.

[0020] In a third aspect, an embodiment of the present application provides a live detection device, which includes:

[0021] An acquisition module, configured to acquire an image including the face of an object to be detected;

[0022] A processing module, configured to input the image including the face of the object to be detected into a depth prediction model to obtain a pseudo-depth map corresponding to the image including the face of the object to be detected. The pseudo-depth map reflects the relative relationship in depth between multiple positions in the face. The depth prediction model is generated according to the first training set and the second training set. The first training set consists of multiple pairs of images including human faces, and each pair of images including human faces is obtained for the same object. The second training set consists of multiple images including human faces and multiple pre-generated pseudo-depth maps; determine the detection result of the object to be detected according to the pseudo-depth map corresponding to the image including the face of the object to be detected.

[0023] In a fourth aspect, an embodiment of the present application further provides an electronic device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the live detection method as described in the first aspect or the method for training a model as described in the second aspect.

[0024] In a fifth aspect, an embodiment of the present application further provides a computer-readable storage medium. The computer-readable storage medium stores computer instructions. When the computer instructions run on a computer, the computer is caused to execute the live detection method as described in the first aspect or the method for training a model as described in the second aspect.

[0025] In a sixth aspect, an embodiment of the present application further provides a computer program product. The computer program product includes a computer program. When the computer program product runs on a computer, it implements the live detection method as described in the first aspect or the method for training a model as described in the second aspect.

[0026] The live detection method provided by the embodiment of the present application pre-generates a depth prediction model according to the first training set and the second training set. During detection, an image including the face of the object to be detected is acquired, the image including the face of the object to be detected is input into the depth prediction model to obtain a pseudo-depth map corresponding to the image including the face of the object to be detected, and the detection result of the object to be detected is determined according to the pseudo-depth map corresponding to the image including the face of the object to be detected.

[0027] The relative depths between various organs in a real human face are different, while it is difficult for the relative depths between various organs in a paper face to reach a level close to that of a real human face. The pseudo-depth map can reflect the relative relationship in depth between multiple positions in the face. Therefore, by using the pseudo-depth map for liveness detection, the differences between real people and photos can be effectively distinguished, and paper attacks can be filtered out without the need for users to cooperate to complete specified actions, which can improve the user experience.

[0028] In addition, generating a depth prediction model based on the first training set belongs to self-supervised learning, and generating a depth prediction model based on the second training set belongs to supervised learning. In the embodiments of the present application, when training the depth prediction model, self-supervised learning and supervised learning are combined. When training the depth prediction model based on self-supervised learning, it is not necessary to pre-generate pseudo-depth maps. Based on the depth prediction model trained by self-supervised learning, supervised learning can be carried out, which can reduce the number of pseudo-depth maps that need to be pre-generated, avoid the shortage of the number of pseudo-depth maps that need to be pre-generated, and thus improve the generalization of the depth prediction model. In addition, pseudo-depth maps can be pre-generated in a targeted manner to make up for the deficiencies of self-supervised learning under certain imaging conditions. In short, compared with self-supervised learning or supervised learning, the method provided by the embodiments of the present application can improve the accuracy of the depth prediction model when predicting pseudo-depth maps, and further improve the accuracy of liveness judgment. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0030] Figure 1 is a schematic flowchart of a method for training a model provided by an embodiment of the present application;

[0031] Figure 2 is a schematic diagram of a pseudo-depth map corresponding to liveness provided by an embodiment of the present application;

[0032] Figure 3 is a schematic diagram of a pseudo-depth map corresponding to non-liveness provided by an embodiment of the present application;

[0033] Figure 4 is a schematic flowchart of a liveness detection method provided by an embodiment of the present application;

[0034] Figure 5 is a schematic flowchart of a method for generating a pseudo-depth map provided by an embodiment of the present application;

[0035] Figure 6It is a schematic flowchart of a process for stitching images provided by an embodiment of the present application;

[0036] Figure 7 It is a schematic flowchart of a process for making a judgment based on global feature information provided by an embodiment of the present application;

[0037] Figure 8 It is a schematic flowchart of a process for making a judgment based on local feature information provided by an embodiment of the present application;

[0038] Figure 9 It is a schematic structural diagram of a device for training a model provided by an embodiment of the present application;

[0039] Figure 10 It is a schematic structural diagram of a device for live detection provided by an embodiment of the present application;

[0040] Figure 11 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0041] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0042] Hereinafter, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present application, unless otherwise stated, the meaning of "a plurality" is two or more, and "at least one", "one or more" means one, two or more.

[0043] The reference to "one embodiment" or "some embodiments" etc. described in this specification means that a specific feature, structure or characteristic described in conjunction with the embodiment is included in one or more embodiments of the present application. Thus, the phrases "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments" etc. that appear in different places in this specification are not necessarily all referring to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "include", "comprise", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0044] Face recognition technology has been widely used in various scenarios such as mobile payment, device unlocking, access control management, etc. When a face recognition system conducts face recognition, the object to be recognized may be a real person, a video or photo played by an electronic device, or a photo printed on paper. Playing a video or photo using an electronic device belongs to screen-type attacks, and printing a photo on paper belongs to paper attacks.

[0045] Based on the principle of infrared imaging, when the face recognition system is under screen-type attacks, it cannot form an image. Therefore, screen-type attacks can be filtered out by setting the camera in the face recognition system as a near-infrared camera. For paper attacks, such attacks can be filtered out by the user's cooperation to complete specified actions such as blinking, opening the mouth, shaking the head, etc. Since user cooperation is required, the user experience is not good.

[0046] Currently, the methods of silent liveness detection without user cooperation can generally be divided into two categories. The first category is to extract the differences between the texture, image quality, etc. of the attack image and the liveness image, as well as information such as the borders leaked by various attacks for judgment. This type of method mainly directly extracts the features of binocular infrared images, such as traditional local binary pattern (LBP) features, histogram features, etc., or trains a neural network to extract features, and based on the extracted features, it distinguishes whether the object to be recognized is a live body or a non-live body. This type of method is prone to misjudging some attack samples with high similarity as live bodies.

[0047] The second category is to make a judgment based on the differences in parallax or depth between live bodies and non-live bodies. This type of method can be further divided into two types. The first method is to predict the depth of some key points and judge whether the object to be recognized is a live body based on the depth of some key points. This method first calibrates the binocular infrared camera to obtain the internal parameters of the binocular infrared camera, then detects multiple face key points in the image captured by the binocular infrared camera, and assumes that these face key points are matched. Furthermore, it estimates the external parameters of the binocular camera to obtain the depth values of these face key points. Finally, it extracts the depth value features of these key points to judge whether the object to be recognized is a live body. The premise of this method is that the key points of the binocular images are matched. However, there are errors in the estimation of the key points themselves, which will lead to the accumulation of errors, affecting the accuracy of the predicted depth and ultimately affecting the accuracy of liveness judgment.

[0048] The second method is to predict the depth map of the face area based on the method of stereo matching, extract the features of the depth map, and combine the features of the depth map and the features of the binocular infrared images captured by the camera to judge whether the object to be recognized is a live body. Among them, stereo matching can be further divided into traditional methods and deep learning-based methods. Traditional methods usually rely on some artificial empirical values and generally have poor robustness.

[0049] There are various deep learning-based methods. For example, based on the Spatial Transformer Networks (STN), a binocular depth map is predicted, and based on the depth map and the global features of the binocular infrared images, it is determined whether the recognized object is a live body. This method belongs to the self-supervised learning method. Under strong light conditions or some other imaging conditions, some areas of the face may become textureless areas. At this time, there is no obvious distinction between the depth maps generated only by the self-supervised learning method for live bodies and non-live bodies, which affects the accuracy of live body judgment. In addition, this method only uses global features to determine whether the recognized object is a live body, and the accuracy is relatively low.

[0050] For example, based on the Structure from Motion (SfM) algorithm or a trained neural network model, the binocular infrared images are input to generate corresponding pseudo-depth maps, and the pseudo-depth maps are compared with the three-dimensional information template maps to determine whether the recognized object is a live body. During the process of training the neural network model, pre-generated pseudo-depth maps are used in this method, which belongs to supervised learning. There may be a large error between the pseudo-depth maps predicted and generated only by supervised learning and the pre-generated pseudo-depth maps. For example, there is still a relative depth relationship between multiple positions of the face in the pseudo-depth map of a flat paper generated based on this method, which affects the accuracy of subsequent live body judgment. Moreover, the supervised learning method requires a large number of pre-generated pseudo-depth maps. When the number of pre-generated pseudo-depth maps is insufficient, the generalization of the neural network model obtained based on the supervised learning method is poor, which affects the accuracy of live body judgment. In addition, the three-dimensional information template map may not cover all the three-dimensional information situations of the human face, which may cause errors in the comparison and also affect the accuracy of subsequent live body judgment.

[0051] In view of the defects of the existing methods, this application proposes a method for live body detection based on stereo matching and deep learning. This method jointly pre-trains a depth prediction model using self-supervised learning and supervised learning. During live body detection, an image including the face of the detected object is obtained, and the image including the face of the detected object is input into the depth prediction model to obtain the corresponding pseudo-depth map, and a live body judgment is made based on the corresponding pseudo-depth map.

[0052] The relative depths between various organs in a real human face are different, while it is difficult for the relative depths between various organs in the face on a paper to reach a level close to that of a real human face. The pseudo-depth map can reflect the relative depth relationship between multiple positions in the face. Therefore, by using the pseudo-depth map for live body detection, the difference between a real person and a photo can be effectively distinguished, and it is possible to filter out paper attacks without the user's cooperation to complete a specified action, which can improve the user experience.

[0053] In addition, when training the depth prediction model in the embodiments of the present application, self-supervised learning and supervised learning are combined. When training the depth prediction model based on self-supervised learning, it is not necessary to pre-generate pseudo-depth maps in advance. Based on the depth prediction model trained by self-supervised learning, supervised learning can be performed, which can reduce the number of pseudo-depth maps that need to be pre-generated in advance and avoid the shortage of the number of pseudo-depth maps that need to be pre-generated in advance, thereby improving the generalization of the depth prediction model. In addition, pseudo-depth maps can be pre-generated specifically to make up for the deficiencies of self-supervised learning under certain imaging conditions. In short, compared with self-supervised learning or supervised learning, the method provided in the embodiments of the present application can improve the accuracy of the depth prediction model when predicting pseudo-depth maps, and further improve the accuracy of live body judgment.

[0054] The following will be combined with Figure 1 to illustrate the method 100 for training a model provided by the embodiments of the present application. As Figure 1 shown, the method 100 includes:

[0055] S101: Obtain a first training set and a second training set. The first training set consists of multiple pairs of images including human faces, and each pair of images including human faces is obtained for the same object. The second training set consists of multiple images including human faces and multiple pre-generated pseudo-depth maps.

[0056] S102: Train a depth prediction model according to the first training set and the second training set.

[0057] In the embodiments of the present application, the training of the depth prediction model consists of two parts: self-supervised learning training and joint training of self-supervised learning and supervised learning, which will be described in detail below.

[0058] (1) Self-supervised learning

[0059] First, the principle of self-supervised learning will be described. Self-supervised learning is learned according to the first training set, which consists of multiple pairs of images including human faces, and each pair of images including human faces is obtained for the same object.

[0060] Each pair of images including human faces is divided into a left image and a right image. The perspectives for obtaining the left image and the right image are different, that is, the positions of the left image and the right image in the imaging plane are different. For example, imagining the human eyes as a binocular camera and holding up a finger in front as the target object, and closing the left or right eye separately to observe the target object. At this time, the position of the target object in the imaging plane has moved. Since there are position differences between the left image and the right image, the three-dimensional geometric information of the target object can be restored by calculating the position deviation between the corresponding points of the two images, and then the depth map of the target object can be obtained.

[0061] Each pair of images including a human face can be acquired at different times or at the same time. Since the object may be in a moving state, when one image in each pair of images including a human face is acquired first and then the other image is acquired, there is a temporal difference between the two images, and the poses of the same object in the two images may not be the same, which is likely to cause errors when calculating the position deviation between corresponding points of the two images. When acquiring each pair of images including a human face at the same time, the temporal difference between the two images can be avoided, improving the accuracy of the calculation.

[0062] The depth map estimated by self-supervised learning is usually a pseudo-depth map, that is, the depth value is not the depth value in the real physical world. Nevertheless, the pseudo-depth can still reflect the relative relationship in depth between multiple positions on the human face. For example, the relative depth value from the nose to the mouth and the relative depth value from the eyes to the mouth are different. The difference between real people and photos can be effectively characterized through the pseudo-depth map.

[0063] The learning process of self-supervised learning is described below.

[0064] a. Obtain the first training set.

[0065] The first training set can be obtained by using a binocular camera for the same object, or by using a monocular camera to move parallel and photograph the same object. The object includes real people and photos printed on paper. Real people are living bodies, and photos printed on paper are non-living bodies.

[0066] Select one image in each pair of images including a human face as the source image, and the other image as the target image for supervision. The source image and the target image can be interchanged.

[0067] b. Let the source image be I s , and the target image be I d . Input the source image I s into the depth prediction neural network model to predict the depth map D s corresponding to the source image I s .

[0068] The depth prediction neural network model is an encoder-decoder network structure. The output depth map D s has the same size as the input source image I s . There are many choices for the encoder-decoder network structure, such as segnet, unet, unet++, RefineNet, etc.

[0069] c. Input the source image I s into the pose prediction model to predict the source image I s relative to the target image Id The pose transformation information T d→s , where the pose transformation information T d→s includes the rotation angle and translation information.

[0070] d. Combine the predicted depth map D s and the pose transformation information T d→s , estimate the pixel coordinates of the target image I d in the source image I s , and obtain the estimated target image I s→d through bilinear interpolation.

[0071] I s→d = I s <proj(D s , T d→s , K)>[[]]

[0072] where K is the camera internal parameter, and proj means converting the depth map D s of the source image I s into a 3D point cloud, and then through the predicted pose transformation information T d→s , converting the 3D point cloud of the source image I s to the 3D point cloud of the target image I d , and combining the camera internal parameter K to obtain the 2D coordinate points corresponding to the target image I d on the source image I s . <> indicates that the estimated target image I s→d is obtained by bilinear interpolation of the 2D coordinate points. Since the estimated here is a pseudo-depth map, therefore, there is no need to calibrate the internal parameters of the camera, and a set of camera internal parameters can be defined.

[0073] e. Calculate the photometric consistency error loss between the estimated target image I s→d and the target image I d as the loss function of the self-supervised learning training module:

[0074] L p = pe(I d , I s→d )

[0075] where pe(I a , I b ) = 0.5α(1 - SSIM(I a , I b )) + (1 - α)||I a - I b ||1. Usually, α = 0.85 is set, SSIM is the structural similarity loss function, and ‖‖1 is the L1 loss function.

[0076] (2) Supervised learning

[0077] First, the principle of supervised learning is described. Supervised learning is learned according to the second training set. The second training set consists of multiple images including human faces and multiple pre-generated pseudo-depth maps. The images including human faces correspond one-to-one with the pre-generated pseudo-depth Figure 1 maps. Supervised learning usually first obtains images including human faces, pre-generates pseudo-depth maps corresponding to the images including human faces, and then uses the images including human faces as inputs and the pre-generated pseudo-depth maps corresponding to the images including human faces as the supervised targets for learning.

[0078] The following describes the learning process of supervised learning.

[0079] a. Obtain images including human faces. The objects obtained include real people and photos printed on paper. The obtaining method can be using a binocular camera or using a monocular camera to take parallel moving shots.

[0080] b. Pre-generate a pseudo-depth map for each image including a human face. As Figure 2 shown, based on the 3D face reconstruction algorithm (prnet), pre-generate the pseudo-depth map 22 of the real person 21 (live body). The pseudo-depth map 22 is three-dimensional. As Figure 3 shown, for the photo 31 (non-live body) printed on paper, through an artificial assistance method, pre-generate the pseudo-depth map 32 of the photo 31 printed on paper, and set all the values of the pixels of the pseudo-depth map 32 to 0, represented in black.

[0081] c. Define the L1 loss function (L1 loss). The input of the L1 loss is the pseudo-depth map predicted by the depth prediction model, and the target target is the pseudo-depth map pre-generated in step b. Use the L1 loss as the loss function of the supervised learning training module:

[0082] L1 = ‖pred - target‖1

[0083] where pred is the pseudo-depth map predicted by the depth prediction model. If the input sample is a live body, then target is the pseudo-depth map obtained based on prnet. If the input sample is a non-live body, then target is the pseudo-depth map with all values set to 0.

[0084] (3) Joint training of self-supervised learning and supervised learning

[0085] On the basis of the loss function L p of the aforementioned self-supervised learning, add the L1 loss to complete the training of the entire depth prediction model. Define the loss function of the entire training as L = λ·L p+(1 - λ)·L1, where λ controls the weight of the loss functions of self-supervised learning and supervised learning. From the formula of this loss function, it can be seen that if λ is taken as 0, the entire loss function degenerates into the loss function of supervised learning; if λ is taken as 1, the entire loss function degenerates into the loss function of self-supervised learning.

[0086] The process of jointly training self-supervised learning and supervised learning is as follows.

[0087] Train for 20 epochs based on the Adaw optimizer. One epoch is the process of training all training samples once. When the number of samples in one epoch (that is, all training samples) may be too large, the samples can be divided into multiple batches for training. The size of each batch of samples can be set to 128.

[0088] Set λ to 1 and the learning rate to 1e-4 for the first 10 epochs, that is, only train the self-supervised learning module in the first 10 epochs. After 10 epochs of training, the parameters W of the trained depth prediction model provide a good initial value for the entire optimization objective. At this time, the pseudo-depth map estimated by the depth prediction model satisfies the constraint conditions of the self-supervised learning module and already has the characteristics of discriminating whether the recognized object is a live body.

[0089] Set λ to 0.5 and the learning rate to 1e-5 for the last 10 epochs. Use the first training set and the second training set as the overall training set, and jointly train self-supervised learning and supervised learning. The goal of supervised learning training is to make the estimated pseudo-depth map approach the predefined target pseudo-depth map. The discrimination of the predefined target pseudo-depth map between live bodies and non-live bodies is significant. Therefore, in order to further improve the discrimination of features, the supervised learning loss function is added in the last 10 epochs.

[0090] However, the depth map estimated by the supervised learning module may not be optimal in the self-supervised learning module. That is, when the model parameters W make L1 decrease, it may cause L p to increase. The joint optimization of the two can achieve an overall balance. The advantage of joint training is that under the condition of satisfying the constraint conditions of self-supervised learning, the estimated pseudo-depth map is made to approach the supervised learning goal of the supervised learning module as much as possible, so that the estimated pseudo-depth map is more discriminative between live bodies and non-live bodies.

[0091] After the depth prediction model is trained, the depth prediction model can be used for live detection during face recognition. The following combines Figure 4 to illustrate the live detection method 400 provided in the embodiments of the present application. As Figure 4 shown, this method 400 includes:

[0092] S401: Obtain an image including the face of the object to be detected.

[0093] S402: Input the image including the face of the object to be detected into the depth prediction model to obtain a pseudo-depth map corresponding to the image including the face of the object to be detected. The pseudo-depth map reflects the relative relationship in depth between multiple positions in the face. The depth prediction model is generated based on a first training set and a second training set. The first training set consists of multiple pairs of images including faces, and each pair of images including faces is obtained for the same object. The second training set consists of multiple images including faces and multiple pre-generated pseudo-depth maps.

[0094] S403: Determine the detection result of the object to be detected according to the pseudo-depth map corresponding to the image including the face of the object to be detected.

[0095] In the embodiments of the present application, based on the imaging device of the face recognition system, an image including the face of the object to be detected is obtained. The object to be detected may be a real person or a person in a photo.

[0096] The imaging device may be a monocular camera or a binocular camera. The binocular camera may be a near-infrared camera and a visible light camera, or two near-infrared cameras. When the binocular camera is a near-infrared camera and a visible light camera, the image obtained by the visible light camera needs to be converted into a grayscale image before detection. For the sake of simplicity, hereinafter, the "image including the face of the object to be detected" will be simply referred to as the "infrared image".

[0097] When the imaging device is a monocular camera, the infrared image is a single image. When the imaging device is a binocular camera, the infrared image includes two images, namely the left image and the right image. The left image corresponds to the first image, and the right image corresponds to the second image.

[0098] Taking the binocular camera as an example for illustration, input the left image and the right image into the depth prediction model, as Figure 5 shown, to obtain a left pseudo-depth map corresponding to the left image and a right pseudo-depth map corresponding to the right image. The left pseudo-depth map corresponds to the first pseudo-depth map, and the right pseudo-depth map corresponds to the second pseudo-depth map. In the embodiments of the present application, the estimated pseudo-depth map corresponds to the entire face area, not limited to the estimation of the depth at certain key point positions, without artificial empirical values, and has good robustness.

[0099] After obtaining the left pseudo-depth map and the right pseudo-depth map, the detection result of the object to be detected can be determined according to the feature information in the left pseudo-depth map and the right pseudo-depth map.

[0100] In the first implementation manner, obtain the feature information related to the face in the left pseudo-depth map and the right pseudo-depth map, and determine whether the object to be detected is a live body or a non-live body according to the feature information related to the face.

[0101] In the second implementation manner, feature information related to the human face in the left pseudo-depth map and the right pseudo-depth map is obtained, and feature information related to the human face in the left image and the right image is obtained, and it is determined whether the object to be detected is a live body or a non-live body according to the feature information related to the human face.

[0102] Among them, the feature information related to the human face includes at least one of global feature information and local feature information, that is, only global feature information can be obtained, or only local feature information can be obtained, or both global feature information and local feature information can be obtained. The global feature information includes the feature information of the entire human face region, and the entire human face region includes the contour of the human face and the region within the contour of the human face, and may also include a preset range of regions outside the contour of the human face. The local feature information includes the feature information of local regions in the human face, such as the feature information of key positions such as the nose and eyes.

[0103] The following will be described in detail in combination with the second implementation manner, and the feature information related to the human face includes global feature information and local feature information.

[0104] (1) As Figure 6 shown, the i-channel image and the i-channel pseudo-depth map are spliced into a two-channel image to obtain the two-channel spliced image of the i-channel. i represents {left, right}.

[0105] (2) Based on the multi-task convolutional neural network (Joint Face Detection and Alignment using Multi-task Cascaded Convolutional Networks, MTCNN), the human face region of the i-channel infrared image is located. The coordinates of the human face region in the infrared image and the pseudo-depth map are the same. After the human face region of the infrared image is located, the human face region of the pseudo-depth map can be located based on the same coordinates, optimizing the process of locating the human face region.

[0106] The human face region of the i-channel infrared image and the human face region of the i-channel pseudo-depth map are extracted to generate the spliced image of the human face region of the i-channel. The spliced image of the human face region of the i-channel is input into the global classification network, as Figure 7 shown, to obtain the live body judgment result 1 of the i-channel.

[0107] (3) Based on MTCNN, 5 human face key points (left eye, right eye, nose, left mouth corner, right mouth corner) of the i-channel infrared image are located, and a local region is extracted for each key point to obtain 5 local region images of the i-channel infrared image and 5 local region images of the i-channel pseudo-depth map.

[0108] The specific steps are as follows: Normalize the original image to a resolution of N*N. For the j-th key point, extract the local area with the coordinates of the key point as the center point, and extract an area of (ratio*N, ratio*N). If it exceeds the boundary, fill it by copying the pixel values on the boundary, where j represents one of the aforementioned 5 key points, and the value range of ratio is (0-1), for example, 0.25 can be taken.

[0109] Stitch the local area image in the i-th infrared image with the corresponding local area image in the i-th pseudo-depth map to generate 5 stitched images of the i-th local areas. The stitched images of the left eye, right eye, nose, left mouth corner, and right mouth corner correspond to local area 1, local area 2, local area 3, local area 4, and local area 5 respectively.

[0110] (4) Define 5 convolutional neural networks, namely patch net1, patch net2, patch net3, patchnet4, and patch net5. These 5 convolutional neural networks can have exactly the same network structure, or completely different network structures, or partially the same network structure, as long as the input (1*2*N*N) and output (1*2) of each convolutional neural network are the same, that is, the output is the judgment result of live or non-live.

[0111] (5) As Figure 8 shown, patch net1, patch net2, patch net3, patch net4, and patch net5 correspond to local area 1, local area 2, local area 3, local area 4, and local area 5 respectively. Input the 5 stitched images of the i-th local areas into the 5 convolutional neural networks respectively, and output the discriminant features 1 - discriminant feature 5 of 1*2. Calculate the average value of the 5 judgment features, and finally pass through the softmax function to obtain the i-th live judgment result 2.

[0112] (6) After steps (1) to (5), the left path and the right path will both obtain the live judgment result 1 and the live judgment result 2. Therefore, there are a total of 4 live judgment results in the end. When the proportion of the judgment results being live is greater than or equal to the preset threshold, it is determined that the object to be detected is live. For example, if the proportion is set to 100%, when all 4 live judgment results are live, it is determined that the object to be detected is live; otherwise, the object to be detected is non-live.

[0113] In the embodiments of the present application, the functions of the global classification network and the local classification network are to extract the feature information in the image. There are many choices for the network structure, such as various classic networks and their variants: Inception, ResNet, ShuffleNet, MobileNet, etc.

[0114] The embodiments of the present application combine the global feature information and local feature information of infrared images and pseudo-depth maps, further improving the accuracy of live body judgment.

[0115] The methods for training a model and for live body detection provided by the embodiments of the present application are described above. The devices and electronic devices provided by the embodiments of the present application are described below.

[0116] Figure 9 A device for training a model provided by an embodiment of the present application, the device 900 includes an acquisition module 901 and a processing module 902.

[0117] The acquisition module 901 is configured to acquire a first training set and a second training set. The first training set consists of multiple pairs of images including faces, and each pair of images including faces is acquired for the same object. The second training set consists of multiple images including faces and multiple pre-generated pseudo-depth maps.

[0118] The processing module 902 is configured to train a depth prediction model according to the first training set and the second training set.

[0119] Specifically, the first training set consists of multiple pairs of images including faces, and each pair of images including faces is acquired for the same object at the same moment.

[0120] It should be understood that the device 900 of the embodiments of the present application can be implemented by an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The above PLD can be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof. It can also be implemented by software Figure 1 The method for training a model shown is implemented by software Figure 1 When the method for training a model shown is implemented, the device 900 and its respective modules can also be software modules.

[0121] Figure 10 A device for live body detection provided by an embodiment of the present application, the device 100 includes an acquisition module 1001 and a processing module 1002.

[0122] The acquisition module 1001 is configured to acquire an image including the face of an object to be detected;

[0123] The processing module 1002 is configured to input an image of a face including an object to be detected into a depth prediction model to obtain a pseudo-depth map corresponding to the image of the face including the object to be detected. The pseudo-depth map reflects the relative relationship in depth between multiple positions in the face. The depth prediction model is generated based on a first training set and a second training set. The first training set consists of multiple pairs of images including faces, and each pair of images including faces is obtained for the same object. The second training set consists of multiple images including faces and multiple pre-generated pseudo-depth maps. Determine the detection result of the object to be detected according to the pseudo-depth map corresponding to the image of the face including the object to be detected.

[0124] In particular, the first training set consists of multiple pairs of images including faces, and each pair of images including faces is obtained for the same object at the same time.

[0125] In particular, the processing module 1002 is further configured to determine the detection result of the object to be detected according to the pseudo-depth map corresponding to the image of the face including the object to be detected and the image of the face including the object to be detected.

[0126] In particular, the processing module 1002 is further configured to respectively obtain the pseudo-depth map corresponding to the image of the face including the object to be detected and the feature information related to the face in the image of the face including the object to be detected, and determine the detection result of the object to be detected according to the feature information related to the face.

[0127] In particular, the feature information related to the face includes at least one of global feature information and local feature information. The global feature information includes the feature information of the entire face region, and the local feature information includes the feature information of the local region in the face.

[0128] In particular, the image of the face including the object to be detected includes a first image and a second image. The first image and the second image are obtained based on a binocular camera. The pseudo-depth map corresponding to the image of the face including the object to be detected includes a first pseudo-depth map and a second pseudo-depth map. The first image corresponds to the first pseudo-depth map, and the second image corresponds to the second pseudo-depth map.

[0129] It should be understood that the device 100 in the embodiments of the present application can be implemented by an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The above PLD can be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof. It can also be implemented by software. Figure 4 The living body detection method shown is implemented by software. Figure 4 When implementing the living body detection method shown, the device 100 and its various modules can also be software modules.

[0130] Figure 11 FIG. is a schematic structural diagram of an electronic device 110 provided by an embodiment of the present application. As Figure 11 shown, the device 110 includes a processor 1101, a memory 1102, a communication interface 1103, and a bus 1104. Among them, the processor 1101, the memory 1102, and the communication interface 1103 communicate through the bus 1104, and can also communicate through other means such as wireless transmission. The memory 1102 is used to store instructions, and the processor 1101 is used to execute the instructions stored in the memory 1102. The memory 1102 stores program code 1021, and the processor 1101 can call the program code 1021 stored in the memory 1102 to execute Figure 1 the method of training a model shown or Figure 4 the living body detection method shown.

[0131] It should be understood that in the embodiments of the present application, the processor 1101 can be a CPU, and the processor 1101 can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.

[0132] The memory 1102 may include a read-only memory and a random access memory, and provide instructions and data to the processor 1101. The memory 1102 may also include a non-volatile random access memory. The memory 1102 may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0133] In addition to including a data bus, the bus 1104 may also include a power bus, a control bus, a status signal bus, etc. However, for the sake of clarity, in Figure 11 all kinds of buses are labeled as the bus 1104.

[0134] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more collections of available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid state drive (SSD).

[0135] The above-described embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application and should all be included in the protection scope of the present application.

Claims

1. A method for live detection, characterized in that, The method includes: Obtaining an image including the face of an object to be detected; Inputting the image including the face of the object to be detected into a depth prediction model to obtain a pseudo-depth map corresponding to the image including the face of the object to be detected, where the pseudo-depth map reflects the relative relationship in depth between multiple positions in the face. The depth prediction model is generated based on a first training set and a second training set. The first training set consists of multiple pairs of images including faces, and each pair of images including faces is obtained for the same object. The second training set consists of multiple images including faces and multiple pre-generated pseudo-depth maps; Determining a detection result of the object to be detected according to the pseudo-depth map corresponding to the image including the face of the object to be detected; The learning process of supervised learning includes: a. Obtaining images including faces, and the objects obtained include real people and photos printed on paper; b. Pre-generating a pseudo-depth map for each image including a face. The pseudo-depth map of a real person is pre-generated based on a 3D face reconstruction algorithm, and the pseudo-depth map is three-dimensional; for a photo printed on paper, with the assistance of artificial means, a pseudo-depth map of the photo printed on paper is pre-generated, and all values of the pixels of the pseudo-depth map are set to 0 and represented in black; c. Defining L1 loss, the input of L1 loss is the pseudo-depth map predicted by the depth prediction model, and the target is the pseudo-depth map pre-generated in step b. L1 loss is used as the loss function of the supervised learning training module: L1 = ‖pred - target‖1 where pred is the pseudo-depth map predicted by the depth prediction model. If the input sample is a live body, then target is the pseudo-depth map obtained based on the 3D face reconstruction algorithm. If the input sample is a non-live body, then target is the pseudo-depth map with all values set to 0; among them, the depth prediction model is pre-trained by combining self-supervised learning and supervised learning; The process of jointly training the self-supervised learning and the supervised learning is as follows: Training for 20 epochs based on the Adaw optimizer. One epoch refers to the process of training all training samples once; Setting λ to 1 and the learning rate to 1e-4 for the first 10 epochs. Only the self-supervised learning module is trained in the first 10 epochs. After the 10 epochs of training are completed, the parameters of the trained depth prediction model provide initial values for the entire optimization objective; Setting λ to 0.5 and the learning rate to 1e-5 for the last 10 epochs. Using the first training set and the second training set as an overall training set, jointly training the self-supervised learning and the supervised learning.

2. The method according to claim 1, characterized in that, The determining the detection result of the object to be detected according to the pseudo-depth map corresponding to the image including the face of the object to be detected includes: Determining the detection result of the object to be detected according to the pseudo-depth map corresponding to the image including the face of the object to be detected and the image including the face of the object to be detected.

3. The method according to claim 2, characterized in that, Determining a detection result of the object to be detected according to the pseudo-depth map corresponding to the image of the face including the object to be detected and the image of the face including the object to be detected includes: Obtaining the pseudo-depth map corresponding to the image of the face including the object to be detected and the feature information related to the face in the image of the face including the object to be detected respectively; Determining the detection result of the object to be detected according to the feature information related to the face.

4. The method according to claim 3, characterized in that, The feature information related to the face includes at least one of global feature information and local feature information, the global feature information includes the feature information of the entire face area, and the local feature information includes the feature information of the local area in the face.

5. The method according to any one of claims 1 to 4, characterized in that, The image of the face including the object to be detected includes a first image and a second image, the first image and the second image are obtained based on a binocular camera, the pseudo-depth map corresponding to the image of the face including the object to be detected includes a first pseudo-depth map and a second pseudo-depth map, the first image corresponds to the first pseudo-depth map, and the second image corresponds to the second pseudo-depth map.

6. A method for training a model, characterized in that, The method includes: Obtaining a first training set and a second training set, the first training set consists of multiple pairs of images including faces, each pair of images including faces is obtained for the same object, and the second training set consists of multiple images including faces and multiple pre-generated pseudo-depth maps; Training a depth prediction model according to the first training set and the second training set; The learning process of supervised learning includes: a. Obtaining an image including a face, and the objects obtained include real people and photos printed on paper; b. Generating a pseudo-depth map for each image including a face in advance. The pseudo-depth map of a real person is pre-generated based on a 3D face reconstruction algorithm, and the pseudo-depth map is three-dimensional; for a photo printed on paper, with the assistance of artificial means, the pseudo-depth map of the photo printed on paper is pre-generated, and all values of the pixels of the pseudo-depth map are set to 0 and represented in black; c. Defining an L1 loss, the input of the L1 loss is the pseudo-depth map predicted by the depth prediction model, and the target is the pseudo-depth map pre-generated in step b. The L1 loss is used as the loss function of the supervised learning training module: L1 = ‖pred - target‖1 where pred is the pseudo-depth map predicted by the depth prediction model. If the input sample is a live body, then target is the pseudo-depth map obtained based on the 3D face reconstruction algorithm. If the input sample is a non-live body, then target is the pseudo-depth map with all values set to 0; among them, the depth prediction model is pre-trained by combining self-supervised learning and supervised learning; The process of jointly training the self-supervised learning and the supervised learning is as follows: Training for 20 epochs based on the Adaw optimizer, and one epoch refers to the process of training all training samples once; Set λ for the first 10 epochs to 1 and the learning rate to 1e-4. Only train the self-supervised learning module in the first 10 epochs. After the 10 epochs of training are completed, the parameters of the trained depth prediction model provide initial values for the entire optimization objective; Set λ for the last 10 epochs to 0.5 and the learning rate to 1e-5. Use the first training set and the second training set as an overall training set, and jointly train using the self-supervised learning and the supervised learning.

7. The method according to claim 6, wherein The first training set consists of multiple pairs of images including faces. Each pair of images including a face is obtained for the same object at the same moment.

8. A living body detection device, characterized in that The device includes: An acquisition module for acquiring an image including the face of an object to be detected; A processing module for inputting the image including the face of the object to be detected into a depth prediction model to obtain a pseudo-depth map corresponding to the image including the face of the object to be detected. The pseudo-depth map reflects the relative relationship in depth between multiple positions in the face. The depth prediction model is generated based on a first training set and a second training set. The first training set consists of multiple pairs of images including faces, and each pair of images including a face is obtained for the same object. The second training set consists of multiple images including faces and multiple pre-generated pseudo-depth maps; determining the detection result of the object to be detected according to the pseudo-depth map corresponding to the image including the face of the object to be detected; The learning process of the supervised learning includes: a. Acquire images including faces, and the acquired objects include real people and photos printed on paper; b. Generate a pseudo-depth map for each image including a face in advance. Generate a pseudo-depth map of a real person in advance based on a 3D face reconstruction algorithm. The pseudo-depth map is three-dimensional; for a photo printed on paper, generate a pseudo-depth map of the photo printed on paper through manual assistance, and set all values of the pixels of the pseudo-depth map to 0, represented in black; c. Define L1 loss. The input of L1 loss is the pseudo-depth map predicted by the depth prediction model, and the target target is the pseudo-depth map pre-generated in step b. Use L1 loss as the loss function of the supervised learning training module: L1 = ‖pred - target‖1 where pred is the pseudo-depth map predicted by the depth prediction model. If the input sample is a live body, target is the pseudo-depth map obtained based on the 3D face reconstruction algorithm. If the input sample is a non-live body, target is the pseudo-depth map with all values set to 0; among them, the depth prediction model is pre-trained jointly using self-supervised learning and supervised learning; The process of jointly training the self-supervised learning and the supervised learning is as follows: Train for 20 epochs based on the Adaw optimizer. One epoch refers to the process of training all training samples once; Set λ for the first 10 epochs to 1 and the learning rate to 1e-4. Only train the self-supervised learning module in the first 10 epochs. After the training of 10 epochs is completed, the parameters of the trained depth prediction model provide initial values for the entire optimization objective; Set λ for the subsequent 10 epochs to 0.5 and the learning rate to 1e-5. Use the first training set and the second training set as an overall training set, and jointly train the self-supervised learning and the supervised learning.

9. An electronic device, characterized in that Including: A memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, it implements the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and when the computer instructions run on an electronic device, the electronic device is caused to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Living body detection method and device

    CN105930710A

  • Training method and device of label proportion learning model based on self-supervised learning

    CN113139651A