A computer-implemented method for object liveness detection, a system for object liveness detection, a computer device and a computer-readable storage medium
A method for liveness detection on devices with auto-focus cameras uses multiple images and 3D model comparisons to identify fraudulent facial images, addressing the limitations of complex sensor requirements and improving security on mid-range smartphones.
Patent Information
- Application Number
- PCT/IB2023/063272
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-27
- Publication Date
- 2025-07-03
AI Technical Summary
Existing liveness detection methods require complex sensor systems like depth or infrared sensors, making them unsuitable for mid-range smartphones with limited capabilities, and lack efficiency in identifying fraudulent attempts using facial images.
A method utilizing a terminal device with an auto-focus camera to capture multiple images at different distances, extract keypoints, reconstruct depth maps, and compare 3D models using machine learning algorithms to assess similarity, enabling liveness detection without intricate sensors.
Enables accurate fraud detection on devices with limited capabilities by reducing data requirements and streamlining the process, enhancing security and speed in systems like banking and gambling.
Smart Images

Figure 00000030_0000 
Figure 00000031_0000 
Figure 00000032_0000
Abstract
Description
[0001] A computer-implemented method for object liveness detection, a system for object liveness detection, a computer device and a computer-readable storage medium
[0002] The present invention relates to a computer-implemented method for object liveness detection, a system for object liveness detection, a computer device and a computer-readable storage medium. The objects of the present invention are applicable in detection of object liveness, especially a human face, based on an image taken by a terminal device equipped with a camera with an auto-focus function. These inventions ensure detection of fraud attempts during user authorization using a photo or image presented to the terminal's camera and can be used in banking, insurance or gambling systems, wherever the user's face, especially biometric features of the user face, are used to confirm the authenticity of authorized user.
[0003] An European patent application EP3719694A1 discloses a method, an apparatus and an electronic device for face liveness detection based on a neural network model. The method includes the steps of: obtaining a target visible light image and a target infrared image of a target object to be detected; extracting a first face image from the target visible light image, and extracting a second face image from the target infrared image; generating a target image array of the target object based on multiple monochromatic components of the first face image and a monochromatic component of the second face image; and feeding the target image array into a pre-trained neural network model for detection, to obtain a face liveness detection result of the target object. U.S. patent application publication no. US2017345146A1 discloses a liveness detection method and a liveness detection system. The liveness detection method includes: obtaining first and second face image data of an object to be detected, and at least one of the first and the second face image data being a depth image; determining a first face region and a second face region, determining whether the first and the second face regions correspond to each other, and extracting, when it is determined that the first and the second face region corresponds to each other, a first and a second face image from the first and the second face region respectively; determining a first classification result for the extracted first face image and a second classification result for the extracted second face image; and determining, based on the first classification result and the second classification result, a detection result for the object to be detected.
[0004] U.S. patent publication no. US10546183B2 discloses a liveness detection system comprising a controller, a video input, a feature recognition module, and a liveness detection module. The controller is configured to control an output device to provide randomized outputs to an entity over an interval of time. The video input is configured to receive a moving image of the entity captured by a camera over the interval of time. The feature recognition module is configured to process the moving image to detect at least one human feature of the entity. The liveness detection module is configured to compare with the randomized outputs a behaviour exhibited by the detected human feature over the interval of time to determine whether the behaviour is an expected reaction to the randomized outputs, thereby determining whether the entity is a living being.
[0005] A publication of the international patent application no. W02019056310A1 discloses a method performed by an electronic device, wherein the method includes receiving an image depicting a face. The method also includes detecting at least one facial landmark of the face in the image. The method further includes receiving a depth image of the face. The method additionally includes determining at least one landmark depth by mapping the at least one facial landmark to the depth image. The method also includes determining a plurality of scales of depth image pixels based on at least one landmark depth. The method further includes determining a smoothness measure based on the scales of the depth image pixels. The method additionally includes determining facial liveness based on the smoothness measure.
[0006] U.S. patent publication no. US10127639B2 discloses an image processing device and the like that can generate a composite image in a desired focusing condition. In a smartphone, an edge detecting section detects an edge as a feature from a plurality of input images taken with different focusing distances, and detects the intensity of the edge as a feature value. A depth estimating section then estimates the depth of a target pixel, which is information representing which of the plurality of input images is in focus at the target pixel, by using the edge intensity detected by the edge detecting section. A depth map generating section then generates a depth map based on the estimation results by the depth estimating section.
[0007] The technical problem facing the present invention is to provide such a liveness detection method capable of identifying fraudulent attempts during user authorization by capturing a facial image. This method aims to prevent fraud where someone tries to use a displayed or printed image of a face on the terminal's camera for authorization. The goal is to provide a method and system that can detect an object's liveness while being feasible for user devices with limited capabilities, specifically those equipped only with an auto-focus camera and lacking complex sensor systems. Additionally, the aim is to provide a simpler implementation method, with fewer stages, yet capable of accurately determining an object's liveness based on a reduced amount of input data, particularly a small set of images captured by the camera. Moreover, there's a desire to provide a method feasible for implementation on mid-range smartphones that have front cameras with restricted capabilities. In one aspect, the present invention provides a computer-implemented method for object liveness detection, wherein the method comprises: at a first distance xi of the camera from the object for detection obtaining a first distance xi base image comprising the object for detection using a terminal device equipped with a camera with an auto-focus system, extracting at least two keypoints of the object from the base image using a machine learning keypoints detection algorithm to create a map of object keypoints, obtaining at least a first image of the object for detection with the auto-focus of the camera set on the first keypoint of the object, and a second image of the object for detection with the auto-focus of the camera set on the second keypoint of the object, changing the distance between the camera and the object for detection from xi to X2, wherein the distance xi is different than the distance X2, at a second distance X2 of the camera from the object for detection obtaining a second distance X2 base image comprising the object for detection using a terminal device equipped with a camera with an auto-focus system, extracting at least two keypoints of the object from the base image using a machine learning keypoints detection algorithm to create a map of object keypoints, wherein the at least two keypoints correspond to at least two keypoints extracted at the first distance xi, obtaining at least a first image of the object for detection with the auto-focus of the camera set on the first keypoint of the object, and a second image of the object for detection with the auto-focus of the camera set on the second keypoint of the object, using a machine learning depth map reconstruction algorithm reconstructing a first distance xi depth map from the images obtained at the first distance xi of the camera from the object for detection, using a machine learning depth map reconstruction algorithm reconstructing a second distance X2 depth map from the images obtained at the second distance X2 of the camera from the object for detection, using a machine learning similarity comparison algorithm comparing at least a first distance xi 3D model and a second distance X2 3D model, wherein the step of comparing comprises: extracting of area corresponding to the object for detection from the first distance xi base image, performing an object pose estimation on the extracted part of the first distance xi base image, performing an inverse transformation of the first distance xi depth map, extracting of area corresponding to the object for detection from the second distance X2 base image, performing an object pose estimation on the extracted part of the second distance X2 base image, performing an inverse transformation of the second distance X2 depth map, scaling the inversely transformed depth maps for distance xi and X2, forming the first distance xi 3D model and the second distance X2 3D model, performing similarity assessment between the first distance xi 3D model and the second distance X2 3D model.
[0008] Preferably, the first distance xi base image and / or the second distance X2 base image is taken with the auto-focus of the camera set on a central region of the image. Preferably, changing the distance between the camera and the object for detection from xi to X2 includes moving the camera closer or farther from the object, moving the object closer or farther from the camera.
[0009] Preferably, at least two keypoints of the object are selected as the keypoints having a depth distance greater than the average depth distance between any two keypoints of the object.
[0010] Preferably, the keypoints of the object are selected from the group comprising: left_eye_center, right_eye_center, left_eye_inner_corner, left_eye_outer_corner, right_eye_inner_corner, right_eye_outer_corner, left_eyebrow_inner_end, left_eyebrow_outer_end, right_eyebrow_inner_end, right_eyebrow_outer_end, nose_tip, mouth_left_corner, mouth_right_corner, mouth_center_top_lip, mouth_center_bottom_lip.
[0011] Preferably, the terminal device is an electronic device selected from the group comprising: a smartphone, a tablet, a laptop, a PDA.
[0012] Preferably, the camera is a front camera of the electronic device.
[0013] In another aspect, the present invention provides a system for object liveness detection comprising a terminal device equipped with a camera with an autofocus system and a liveness detection server, wherein the terminal device is configured to: at a first distance xi of the camera from the object for detection obtain a first distance xi base image comprising the object for detection, extract at least two keypoints of the object from the base image using a machine learning keypoints detection algorithm to create a map of object keypoints, obtain at least a first image of the object for detection with the auto-focus of the camera set on the first keypoint of the object, and a second image of the object for detection with the auto-focus of the camera set on the second keypoint of the object, at a second distance X2 of the camera from the object for detection obtain a second distance X2 base image comprising the object for detection, extract at least two keypoints of the object from the base image using a machine learning keypoints detection algorithm to create a map of object keypoints, wherein the at least two keypoints correspond to at least two keypoints extracted at the first distance xi, obtain at least a first image of the object for detection with the auto-focus of the camera set on the first keypoint of the object, and a second image of the object for detection with the auto-focus of the camera set on the second keypoint of the object, communicate with the liveness detection server for sending and receiving data including images taken by the camera, the liveness detection server is configured to: communicate with the terminal device for sending and receiving data including images taken by the camera, using a machine learning depth map reconstruction algorithm reconstruct a first distance xi depth map from the images obtained at the first distance xi of the camera from the object for detection, using a machine learning depth map reconstruction algorithm reconstruct a second distance X2 depth map from the images obtained at the second distance X2 of the camera from the object for detection, using a machine learning similarity comparison algorithm compare a first distance xi 3D model and a second distance X2 3D model, wherein the comparison comprises: extracting of area corresponding to the object for detection from the first distance xi base image, performing an object pose estimation on the extracted part of the first distance xi base image performing an inverse transformation of the first distance xi depth map, extracting of area corresponding to the object for detection from the second distance X2 base image, performing an object pose estimation on the extracted part of the second distance X2 base image, performing an inverse transformation of the second distance X2 depth map, scaling the inversely transformed depth maps for distance xi and X2, forming the first distance xi 3D model and the second distance X2 3D model, performing similarity assessment between the first distance xi 3D model and the second distance X2 3D model.
[0014] Preferably, the first distance xi base image and / or the second distance X2 base image is taken with the auto-focus of the camera set on a central region of the image.
[0015] Preferably, at least two keypoints of the object are selected as the keypoints having a depth distance greater than the average depth distance between any two keypoints of the object.
[0016] Preferably, the keypoints of the object are selected from the group comprising: left_eye_center, right_eye_center, left_eye_inner_corner, left_eye_outer_corner, right_eye_inner_corner, right_eye_outer_corner, left_eyebrow_inner_end, left_eyebrow_outer_end, right_eyebrow_inner_end, right_eyebrow_outer_end, nose_tip, mouth_left_corner, mouth_right_corner, mouth_center_top_lip, mouth_center_bottom_lip.
[0017] Preferably, the terminal device is an electronic device selected from the group comprising: a smartphone, a tablet, a laptop, a PDA. Preferably, the camera is a front camera of the electronic device.
[0018] In yet another aspect, the present invention provides a computer device, comprising: at least one processor and a memory communicatively connected to the at least one processor, wherein the memory stores a computer-readable instructions, and the computer-readable instructions when executed by the one or more processors configure the one or more processors to implement the method according to the first aspect of the present invention.
[0019] In further another aspect, the present invention provides a computer-readable storage medium, storing computer-readable instructions, and the computer- readable instructions when executed by the one or more processors configure the one or more processors to implement the method according to the first aspect of the present invention.
[0020] The method for object liveness detection enables the identification of fraudulent attempts during user authorization by capturing the user's facial photo. This method operates solely on a terminal equipped with an auto-focus camera, eliminating the need for intricate sensor systems like depth or infrared sensors (e.g., LIDAR) to identify fraud attempts involving presenting images of the authorized user's face in front ofthe terminal's camera. Implementing the liveness detection method of the present invention only requires a user device, such as a mid-range smartphone, with a front camera of limited capabilities and basic autofocus functionality. This method needs only a few photos to detect fraud attempts, reducing the data needed for identifying fraud, streamlining the method itself and data management procedures. Additionally, the limited input data enhances the speed of fraud detection, thereby significantly enhancing the security of systems utilizing this method.
[0021] Examples of the invention are presented in the drawing, where Fig. 1 shows the flowchart illustrating one embodiment of a method for object liveness detection according to the present invention, Fig. 2 the flowchart illustrating in details the sub-sets of preforming comparison step of the method for object liveness detection according to the present invention, Fig. 3 shows the block diagram illustrating one embodiment of a computer device according to the present invention.
[0022] Example 1
[0023] An embodiment of the computer-implemented method for object liveness detection according to the present invention is shown in the flowchart diagram in Fig. 1. The method for object liveness detection is implemented in a system for object liveness detection comprising a terminal device equipped with a camera with an auto-focus system and a liveness detection server.
[0024] In this embodiment, the terminal device is selected in the form of a smartphone equipped with a front RGB camera with auto-focus functionality. However, it should be emphasized that the type of the terminal device used is not limited to the smartphone presented in this example, and in alternative embodiments other terminal devices equipped with a camera with auto-focus functionality, such as a tablet, laptop or PDA, can be used.
[0025] Additionally, in the system for object liveness detection, the terminal device is connected to the liveness detection server via appropriate communication means, wireless or wired. The terminal device and the liveness detection server provide two-way communication allowing the transfer of images between units, as well as other data used in the method for object liveness detection according to the present invention. The liveness detection server stores the algorithms used at various steps of the method for object liveness detection and performs appropriate logical operations and simulations to obtain the desired results.
[0026] The method for object liveness detection according to the present invention starts with the step of obtaining 101 a first distance xi base image comprising the object for detection. The xi is the distance between the camera of the terminal device and the object being detected, usually the user's face. Typically, this step is carried out by the front camera of a smartphone device, and the image taken is a classic selfie photo, where most of the frame is filled with the user's face. As already mentioned, the camera of the terminal device is equipped with an auto-focus system. Obtaining 101 the first distance xi base image is performed with autofocus set to the central region of the image. As used herein, the term central image region includes a two-dimensional image region covering 20% of the image area with its center located at the center of symmetry of the image. Therefore, when obtaining 101 the first distance xi base image, the camera's auto-focus is set in the mentioned area of 20% of the image area with the center located at the center of image symmetry. Selecting this area ensures the highest probability of capturing the user's face for selecting the first auto-focus in a location that allows facial features to be identified.
[0027] After obtaining 101 the first distance xi base image, the method for object liveness detection proceeds to the next step, in which two keypoints (KPs) of the object from the first distance xi base image are extracted 102 in the terminal device. The extraction 102 is carried out using a machine learning (ML) keypoints (KPs) detection algorithm and, as a result, allows the creation of a map of object keypoints. In this embodiment, MediaPipe (from Google) is used as the ML keypoints detection algorithm, using BlazeFace, which is an implementation of: SSD: Single Shot MultiBox Detector, Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian Szegedy, Scott Reed, Cheng-Yang Fu, Alexander C. Berg, ECCV2016. In alternative embodiments, it is possible to use different ML keypoints detection algorithms, provided that the keypoints characteristic of the face are effectively extracted and allow the method according to the invention to be further carried out. An alternative ML keypoints detection algorithm can be, for example, DLib.
[0028] In the present embodiment of the method for object liveness detection, two KPs of the object are selected, namely left_eye_inner_corner as KPi and nose_tip as KP2. The selected two KPs have a depth distance greater than the average depth distance between any two keypoints of the object. This means that there is a significant distance between these two KPs relative to the camera, which increases the sensitivity of the method, especially in the next steps of reconstructing a depth map.
[0029] After extracting 102 two KPs of the object (face) from the first distance xi base image, the method for object liveness detection according to the present invention proceeds to the step in which a first image of the object for detection is obtained 103 with the auto-focus of the camera of the terminal device set on the first KP of the object, and a second image of the object for detection with the autofocus of the camera set on the second KP of the object. Therefore in this step 103 two images are taken in which in the first image the camera's AF is set to left_eye_inner_corner as KPi and in the second image the camera's AF is set to nose_tip as KP2.
[0030] The images obtained at this step 103, along with information about the KPs to which the camera's AF was set, are sent from the terminal device to the liveness detection server, and the method for object liveness detection moves to the next step, in which the distance between the camera and the object for detection is changed 104 from xi to X2. In this embodiment of the invention, changing 104 the distance of the camera from the object is performed by moving the camera towards the user's face, therefore the distance X2 is smaller than the distance xi. In alternative embodiments, changing 104 the distance between the camera and the object for detection from xi to X2 is not limited to moving the terminal device towards the object and can also be implemented by moving the camera farther from the object, moving the object closer or farther from the camera.
[0031] Then, in the method for object liveness detection, in which the camera is already at a distance X2 from the object, i.e. the user's face, the step of obtaining 105 a second distance X2 base image comprising the object for detection occurs via a terminal device equipped with a camera. As in step 101, in step 105 when obtaining the second distance X2 base image, AF defaults to the central region of the image.
[0032] In the next step of the method, using a ML keypoints detection algorithm from step 102, two KPs are extracted to create a map of object keypoints, wherein two keypoints correspond to two keypoints extracted at the first distance xi, i.e. in this example left_eye_inner_corner as KPi and in nose_tip as KP2.
[0033] After creating the map of object keypoints, the next step of the method is to obtain 107 the first image of the object for detection with the auto-focus of the camera set on the first KP of the object, and the second image of the object for detection with the auto-focus of the camera set on the second KP of the object. Therefore in this step 107 two images are taken in which in the first image the camera AF is set to left_eye_inner_corner as KPi and in the second image the camera AF is set to nose_tip as KP2.
[0034] The images obtained at this step 107, along with information about the KPs to which the camera's AF was set, are sent from the terminal device to the liveness detection server, and the method for object liveness detection moves to the next step, in which the liveness detection server uses a ML depth map reconstruction algorithm for reconstructing 108 a first distance xi depth map from the images obtained at the first distance xi of the camera from the object for detection. The present embodiment uses the ML depth map reconstruction algorithm disclosed in Suwajanakorn, S., Hernandez, C. and Seitz, S.M., 2015. Depth from focus with your mobile phone. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 3497-3506). The type of ML algorithm used to reconstruct 108 the depth map is not limited to that mentioned above, and in alternative embodiments, other algorithms may be used, including without limitation the algorithm disclosed in: Hazirbas, C., Soyer, S.G., Staab, M.C., Leal- Taixe, L. and Cremers, D., 2019. Deep depth from focus. In Computer Vision-ACCV 2018: 14th Asian Conference on Computer Vision, Perth, Australia, December 2- 6, 2018, Revised Selected Papers, Part III 14 (pp. 525-541); Bhat, S.F., Birkl, R., Wofk, D., Wonka, P. and Muller, M., 2023. Zoedepth: Zero-shot transfer by combining relative and metric depth. arXiv preprint arXiv:2302.12288; or Yang, X., Fu, Q., Elhoseiny, M. and Heidrich, W., 2023. Aberration-aware depth-from-focus. IEEE Transactions on Pattern Analysis and Machine Intelligence.
[0035] Similarly, on the liveness detection server, using ML depth map reconstruction algorithm, a second distance X2 depth map from the images obtained at the second distance X2 of the camera from the object for detection is reconstructed 109. As a result, after performing steps 108 and 109, two depth maps are obtained: the first distance xi depth map and the second distance X2 depth map. The obtained depth maps essentially create 3D models of the object for detection, which additionally include the depth component of the images.
[0036] Having these two depth maps, the method moves to the step of performing 110 a comparison of the first distance xi 3D model and the second distance X2 3D model for assessing the similarity between 3D models of the object for detection, using a machine learning similarity comparison algorithm.
[0037] The step of preforming 110 the comparison is presented in the flowchart diagram in Fig. 2 and includes a substep in which areas corresponding to the object for detection are extracted 1101 from the first distance xi base image (in the present embodiment, the face is extracted). The extraction 1101 of areas corresponding to the object for detection from the image is carried out using any suitable ML algorithm, such as MediaPipe (from Google), DLib, etc.
[0038] Then, in the next substep object pose estimation 1102 is performed on the extracted part of the image to obtain parameters defining at least the rotation of the object relative to the axis. In the present embodiment, for an object to be detected in the form of a user's face, head pose estimation is performed in this substep 1102, which returns the pitch, yaw and roll parameters for the analyzed object. The algorithm used to perform head pose estimation is disclosed in Y. Zhou, J. Gregson, Real-time fine-grained estimation for wide range head pose, arXiv:2005.10353.
[0039] In an alternative embodiment, another appropriate object pose estimation may be performed in this substep, such as, without limitation, 6D object pose estimation, which returns the parameters pitch, yaw, roll and additionally x, y, z coordinates. A suitable algorithm to perform 6D object pose estimation is Thorsten Hempel, Ahmed A. Abdelrahman, Ayoub Al-Hamadi, 6D Rotation Representation For Unconstrained Head Pose Estimation, 2022 IEEE International Conference on Image Processing (ICIP), 2022. Alternatively, any of the algorithms disclosed in Asperti, A., Filippini, D. Deep Learning for Head Pose Estimation: A Survey. SN COMPUT. SCI. 4, 349 (2023) may be used to perform substep 1102.
[0040] In the next substep of the comparison 110 step, the inverse transformation 1103 of the first distance xi depth map (forming the 3D object model) is performed, using the parameters (pitch, yaw, roll) calculated in the substep 1102, to the position corresponding to the pitch, yaw, roll parameters equal to 0, 0, 0, respectively. The inverse transformation 1103 used in this substep is implemented via the algorithm providing base linear algebra operations, such as rotation. This type of algorithm will be known for the skilled in the art.
[0041] Similarly, for the second distance X2 from the second distance X2 base image, areas corresponding to the object for detection are extracted 1104, and then the object pose estimation 1105 is performed to obtain parameters defining at least the rotation of the object relative to the axis, i.e. parameters pitch, yaw, roll. Using the calculated parameters pitch, yaw and roll for the second distance X2 base image, an inverse transformation 1106 of the second distance X2 depth map (forming the 3D model) of the object is performed to the position corresponding to the parameters pitch, yaw, roll equal to 0, 0, 0, respectively.
[0042] Then the inversely transformed depth maps for distance xi and X2 are scaled 1107 to obtain the same model sizes. Scaling 1107 includes any of: scaling the model for distance xi to the model for distance X2, scaling the model for distance X2 to the model for distance xi, scaling the model for distance xi and X2 to a model size different from the initial sizes of the models for distance xi and X2. Scaling 1107 substep is implemented via a common scaling algorithm known from the computer graphics field.
[0043] The depth maps scaled in substep 1107 constitute the first distance xi 3D model and the second distance X2 3D model, respectively.
[0044] In the last substep of the comparison 110 step, a similarity assessment 1108 is performed between the first distance xi 3D model and the second distance X2 3D model. The similarity assessment 1108 is performed using a ML similarity comparison algorithm as disclosed in for example Dai, G., Xie, J. and Fang, Y., Siamese CNN-BiLSTM architecture for 3D shape representation learning, I JCAI'18 : Proceedings of the 27th International Joint Conference on Artificial Intelligence, July 2018; Patel, A. and Smith, W.A., 2009, June. 3d morphable face models revisited. In 2009 IEEE conference on computer vision and pattern recognition (pp. 1327-1334). IEEE or Zhang, D., Wu, Z., Wang, X., Lv, C. and Liu, N., 2021. 3D skull and face similarity measurements based on a harmonic wave kernel signature. The Visual Computer, 37, pp.749-764.
[0045] In the present embodiment, a computer device is provided. The computer device may be a terminal device or a liveness detection server, and an internal structural diagram thereof may be as shown in Fig. 3. When the computer device is the terminal device, the computer device may include a camera with an auto-focus system. The computer device includes a processor, a memory, and a communication means. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer-readable instructions. The internal memory provides an environment for running the operating system and the computer-readable instructions in the non-volatile storage medium. The communication means of the computer device is configured to connect and communicate with another computer device via a wireless or wired connection. The computer-readable instructions when executed by the processor implements the foregoing method for object liveness detection.
[0046] The structure as shown in Fig. 3 is block diagram of some structures related to the present invention, and does not constitute a limitation on the computer devices to which the present invention is applicable.
[0047] In the present embodiment, a computer-readable storage medium is provided. The computer-readable storage medium stores computer-readable instructions, and the computer-readable instructions when executed by a processor implement the steps in the foregoing method for object liveness detection.
[0048] Example 2
[0049] The second embodiment of the computer-implemented method for object liveness detection according to the present invention is similar to the computer- implemented method for object liveness detection according to the present invention presented in the first embodiment, and therefore similar steps will not be described again for the clarity of this disclosure.
[0050] Unlike in the first embodiment, the second embodiment of the computer- implemented method for object liveness detection, after the step of obtaining 101 the first distance xi base image comprising the object for detection, comprises the step of extracting 102 four keypoints of the object from the base image using a ML keypoints detection algorithm to create a map of object keypoints. In the present embodiment of the method for object liveness detection, four KPs of the object are selected, namely right_eye_outer_corner as KPi, left_eye_center as KP2, mouth_center_top_lip as KP3, and right_eyebrow_inner_end as KP4. As a consequence in the next step of the method four images are obtained 103, that is a first image of the object with the auto-focus of the camera of the terminal device set on the first KPi of the object, a second image of the object with the auto-focus of the camera set on the second KP2 of the object, a third image of the object with the auto-focus of the camera set on the third KP3 of the object and the fourth image of the object with the auto-focus of the camera set on the fourth KP4 of the object. Therefore in this step 103 four images are obtained in which in the first image the camera's AF is set to right_eye_outer_corner as KPi, in the second image the camera's AF is set to left_eye_center as KP2, in the third image the camera's AF is set to mouth_center_top_lip as KP3 and in the fourth image the camera's AF is set to right_eyebrow_inner_end as KP4. Increasing the number of KPs in which the camera's AF is set to obtain an image of the object ensures obtaining a more sensitive the method for object liveness detection.
[0051] Similarly, after changing the distance of the object from the camera from xi to X2 and obtaining 105 a second distance X2 base image comprising the object for detection, in the next step of the method the extraction 106 of four keypoints of the object from the base image using a ML keypoints detection algorithm is performed to create a map of object keypoints, wherein the four KPs of the object are right_eye_outer_corner as KPi, left_eye_center as KP2, mouth_center_top_lip as KP3, and right_eyebrow_inner_end as KP4. Then four images are obtained 107 in which in the first image the camera's AF is set to right_eye_outer_corner as KPi, in the second image the camera's AF is set to left_eye_center as KP2, in the third image the camera's AF is set to mouth_center_top_lip as KP3 and in the fourth image the camera's AF is set to right_eyebrow_inner_end as KP4.
[0052] Unlike the method for object liveness detection as set in the first embodiment, in the present embodiment the images are obtained for another distances X3, X4 and xs, that are all different than distances xi and X2, to increase the sensitivity of the method. Therefore, in the next step of the method, the distance of the camera of the terminal device in relation to the object is changed from X2 to X3 and the step of obtaining a third distance X3 base image is performed, followed by the extraction of four KPs, i.e. KPi, KP2, KP3 and KP4. Further, four images are obtained, in which the camera's AF is set to KPi, KP2, KP3 and KP4, respectively. After this step, the distance is changed again from X3 to X4 and a fourth distance X4 base image is obtained again, from which four KPs are extracted, corresponding to the previous KPs, which are used to set the AF of the camera of the terminal device to obtain four images respectively. Finally, the last distance change from X4 to xs is performed and a fifth distance xs base image is obtained, after which four KPs are extracted as mentioned above, which are used to set the AF of the camera of the terminal device to obtain four images respectively. As a result, four images for a distance of xi, four images for a distance of X2, four images for a distance of X3, four images for a distance of X4 and four images for a distance of xs were obtained, i.e. a total of 20 images that are sent to the liveness detection server.
[0053] Similarly to the first embodiment, depth maps corresponding to the appropriate distances are reconstructed from the obtained images, i.e. a first distance xi depth map is reconstructed 108, a second distance X2 depth map is reconstructed 109 and additionally for this embodiment a third distance X3 depth map, a fourth distance X4 depth map and a fifth distance xs depth map are reconstructed.
[0054] Having these five depth maps, the method moves to the step of performing 110 a comparison of the first distance xi 3D model, the second distance X23D model, the third distance X3 3D model, the fourth distance X4 3D model and the fifth distance xs 3D model for assessing the similarity between 3D models of the object for detection, using a machine learning similarity comparison algorithm.
[0055] Similarly to the first embodiment, the step 110 comprises: extraction 1101 of the area corresponding to the object for detection from the first distance xi base image, extraction 1104 of the area corresponding to the object from the second distance X2 base image, and additionally extraction of the areas corresponding to the object from the third distance X3 base image, from the fourth distance X4 base image and from the fifth distance xs base image.
[0056] Then, from the extracted areas of the images the object pose estimations 1102, 1105 are performed to obtain parameters defining at least the rotation of the object relative to the axis for each base image. In the next substep the inverse transformation 1103, 1106 of the first distance xi depth map, the second distance X2 depth map, the third distance X3 depth map, the fourth distance X4 depth map and the fifth distance xs depth map is performed, using the parameters (pitch, yaw, roll) calculated in the previous substep, respectively. Each depth map for distances xi - xs is transformed to the position corresponding to the pitch, yaw, roll parameters equal to 0, 0, 0, respectively.
[0057] Then the inversely transformed depth maps for distance xi- xs are scaled 1107 to obtain the same model sizes. The depth maps scaled in substep 1107 constitute the first distance xi 3D model, the second distance X2 3D model, the third distance X3 3D model, the fourth distance X4 3D model, and the fifth distance xs 3D model, respectively.
[0058] In the last substep of the comparison 110 step, a similarity assessment 1108 is performed between the first distance xi 3D model , the second distance X2 3D model, the third distance X3 3D model, the fourth distance X4 3D model, and the fifth distance xs 3D model.
[0059] List of references:
[0060] 101 - obtaining a first distance xi base image comprising the object for detection
[0061] 102 - extracting at least two keypoints of the object from the base image using a machine learning keypoints detection algorithm to create a map of object keypoints
[0062] 103 - obtaining at least a first image of the object for detection with the autofocus of the camera set on the first keypoint of the object, and a second image of the object for detection with the auto-focus of the camera set on the second keypoint of the object
[0063] 104 - changing the distance between the camera and the object for detection from xi to X2,
[0064] 105 - obtaining a second distance X2 base image comprising the object for detection
[0065] 106 - extracting at least two keypoints of the object from the base image using a machine learning keypoints detection algorithm to create a map of object keypoints
[0066] 107 - obtaining at least a first image of the object for detection with the autofocus of the camera set on the first keypoint of the object, and a second image of the object for detection with the auto-focus of the camera set on the second keypoint of the object
[0067] 108 - reconstructing a first distance xi depth map from the images obtained at the first distance xi of the camera from the object for detection
[0068] 109 - reconstructing a second distance X2 depth map from the images obtained at the second distance X2 of the camera from the object for detection
[0069] 110 - comparing the first distance xi 3D model and the second distance X2 3D model 1101 - extracting of area corresponding to the object for detection from the first distance xi base image
[0070] 1102 - performing an object pose estimation on the extracted part of the first distance xi base image 1103 - performing an inverse transformation of the first distance xl depth map
[0071] 1104 - extracting of area corresponding to the object for detection from the second distance x2 base image
[0072] 1105 - performing an object pose estimation on the extracted part of the second distance x2 base image 1106- performing an inverse transformation of the second distance x2 depth map
[0073] 1107 - scaling the inversely transformed depth maps for distance xi and X2
[0074] 1108 - performing similarity assessment between the first distance xi 3D model and the second distance X2 3D model
Claims
Claims1. A computer-implemented method for object liveness detection, characterized in that, the method comprises: at a first distance xi of the camera from the object for detection obtaining (101) a first distance xi base image comprising the object for detection using a terminal device equipped with a camera with an auto-focus system, extracting (102) at least two keypoints of the object from the base image using a machine learning keypoints detection algorithm to create a map of object keypoints, obtaining (103) at least a first image of the object for detection with the auto-focus of the camera set on the first keypoint of the object, and a second image of the object for detection with the auto-focus of the camera set on the second keypoint of the object, changing (104) the distance between the camera and the object for detection from xi to X2, wherein the distance xi is different than the distance X2, at a second distance X2 of the camera from the object for detection obtaining (105) a second distance X2 base image comprising the object for detection using a terminal device equipped with a camera with an autofocus system, extracting (106) at least two keypoints of the object from the base image using a machine learning keypoints detection algorithm to create a map of object keypoints, wherein the at least two keypoints correspond to at least two keypoints extracted at the first distance xi, obtaining (107) at least a first image of the object for detection with the auto-focus of the camera set on the first keypoint of the object, and asecond image of the object for detection with the auto-focus of the camera set on the second keypoint of the object, using a machine learning depth map reconstruction algorithm reconstructing (108) a first distance xi depth map from the images obtained at the first distance xi of the camera from the object for detection, using a machine learning depth map reconstruction algorithm reconstructing (109) a second distance X2 depth map from the images obtained at the second distance X2 of the camera from the object for detection, using a machine learning similarity comparison algorithm comparing (110) at least a first distance xi 3D model and a second distance X2 3D model, wherein the step of comparing (110) comprises: extracting (1101) of area corresponding to the object for detection from the first distance xi base image, performing (1102) an object pose estimation on the extracted part of the first distance xi base image, performing (1103) an inverse transformation of the first distance xi depth map, extracting (1104) of area corresponding to the object for detection from the second distance X2 base image, performing (1105) an object pose estimation on the extracted part of the second distance X2 base image, performing (1106) an inverse transformation of the second distance X2 depth map,scaling (1107) the inversely transformed depth maps for distance xi and X2, forming the first distance xi 3D model and the second distance X2 3D model, performing (1108) similarity assessment between the first distance xi 3D model and the second distance X2 3D model.
2. The computer-implemented method for object liveness detection according to claim 1, characterized in that, the first distance xi base image and / or the second distance X2 base image is taken with the auto-focus of the camera set on a central region of the image.
3. The computer-implemented method for object liveness detection according to claim 1 or 2, characterized in that, changing the distance between the camera and the object for detection from xi to X2 includes moving the camera closer or farther from the object, moving the object closer or farther from the camera.
4. The computer-implemented method for object liveness detection according to any one of claims 1 to 3, characterized in that, at least two keypoints of the object are selected as the keypoints having a depth distance greater than the average depth distance between any two keypoints of the object.
5. The computer-implemented method for object liveness detection according to any one of claims 1 to 4, characterized in that, the keypoints of the object are selected from the group comprising: left_eye_center, right_eye_center, left_eye_inner_corner, left_eye_outer_corner, right_eye_inner_corner, right_eye_outer_corner, left_eye browj n ne r_e nd, left_eye brow_oute r_e nd, right_eyebrow_inner_end, right_eyebrow_outer_end, nose_tip, mouth_left_corner, mouth_right_corner, mouth_center_top_lip, mouth_center_bottom_lip.
6. The computer-implemented method for object liveness detection according to any one of claims 1 to 5, characterized in that, the terminal device is an electronic device selected from the group comprising: a smartphone, a tablet, a laptop, a PDA.
7. The computer-implemented method for object liveness detection according to claim 6, characterized in that, the camera is a front camera of the electronic device.
8. A system for object liveness detection comprising a terminal device equipped with a camera with an auto-focus system and a liveness detection server, characterized in that the terminal device is configured to: at a first distance xi of the camera from the object for detection obtain (101) a first distance xi base image comprising the object for detection, extract (102) at least two keypoints of the object from the base image using a machine learning keypoints detection algorithm to create a map of object keypoints, obtain (103) at least a first image of the object for detection with the auto-focus of the camera set on the first keypoint of the object, and a second image of the object for detection with the auto-focus of the camera set on the second keypoint of the object, at a second distance X2 of the camera from the object for detection obtain (105) a second distance X2 base image comprising the object for detection, extract (106) at least two keypoints of the object from the base image using a machine learning keypoints detection algorithm to create a map of object keypoints, wherein the at least twokeypoints correspond to at least two keypoints extracted at the first distance xi, obtain (107) at least a first image of the object for detection with the auto-focus of the camera set on the first keypoint of the object, and a second image of the object for detection with the auto-focus of the camera set on the second keypoint of the object, communicate with the liveness detection server for sending and receiving data including images taken by the camera, the liveness detection server is configured to: communicate with the terminal device for sending and receiving data including images taken by the camera, using a machine learning depth map reconstruction algorithm reconstruct (108) a first distance xi depth map from the images obtained at the first distance xi of the camera from the object for detection, using a machine learning depth map reconstruction algorithm reconstruct (109) a second distance X2 depth map from the images obtained at the second distance X2 of the camera from the object for detection, using a machine learning similarity comparison algorithm compare (110) a first distance xi 3D model and a second distance X2 3D model, wherein the comparison (110) comprises: extracting (1101) of area corresponding to the object for detection from the first distance xi base image,performing (1102) an object pose estimation on the extracted part of the first distance xi base image performing (1103) an inverse transformation of the first distance xi depth map, extracting (1104) of area corresponding to the object for detection from the second distance X2 base image, performing (1105) an object pose estimation on the extracted part of the second distance X2 base image, performing (1106) an inverse transformation of the second distance X2 depth map, scaling (1107) the inversely transformed depth maps for distance xi and X2, forming the first distance xi 3D model and the second distance X2 3D model, performing (1108) similarity assessment between the first distance xi 3D model and the second distance X2 3D model.
9. The system for object liveness detection according to claim 8, characterized in that, the first distance xi base image and / or the second distance X2 base image is taken with the auto-focus of the camera set on a central region of the image.
10. The system for object liveness detection according to claim 8 or 9, characterized in that, at least two keypoints of the object are selected as the keypoints having a depth distance greater than the average depth distance between any two keypoints of the object.
11. The system for object liveness detection according to any one of claims 8 to 10, characterized in that, the keypoints of the object are selected from the group comprising: left_eye_center, right_eye_center, left_eye_inner_corner, left_eye_outer_corner, right_eye_inner_corner,right_eye_outer_corner, left_eyebrow_inner_end, left_eyebrow_outer_end, right_eyebrow_inner_end, right_eyebrow_outer_end, nose_tip, mouth_left_corner, mouth_right_corner, mouth_center_top_lip, mouth_center_bottom_lip.
12. The system for object liveness detection according to any one of claims 8 to 11, characterized in that, the terminal device is an electronic device selected from the group comprising: a smartphone, a tablet, a laptop, a PDA.
13. The system for object liveness detection according to claim 12, characterized in that, a camera is a front camera of the electronic device.
14. A computer device, comprising: at least one processor and a memory communicatively connected to the at least one processor, wherein the memory stores a computer-readable instructions, and the computer- readable instructions when executed by the one or more processors configure the one or more processors to implement the method according to one of claims 1 to 7.
15. A computer-readable storage medium, storing computer-readable instructions, and the computer-readable instructions when executed by the one or more processors configure the one or more processors to implement the method according to one of claims 1 to 7.
Citation Information
Patent Citations
Neural network model-based human face living body detection
EP3719694A1
Image processing device having depth map generating unit, image processing method and non-transitory computer readable recording medium
US10127639B2
Liveness detection
US10546183B2
Liveness detection method and liveness detection system
US20170345146A1
Systems and methods for facial liveness detection
WO2019056310A1