A method, apparatus, storage medium, and electronic device for image liveness detection.

By using a depth model for depth estimation and inter-frame fusion in image liveness detection, a high-precision second depth image is generated, which solves the problems of detection accuracy and robustness in low-cost and low-device-requirement environments, and realizes effective image liveness detection in complex environments and low-performance hardware.

CN116129534BActive Publication Date: 2026-01-30ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211089114.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-07
Publication Date
2026-01-30
Estimated Expiration
2042-09-07

AI Technical Summary

Technical Problem

Existing image liveness detection technologies are difficult to effectively defend against attacks such as photos, face swaps, masks, and screen re-photographs in low-cost and low-equipment environments, and require complex operations by users, resulting in insufficient detection accuracy and robustness.

Method used

By acquiring at least two target color images, depth estimation is performed using a first depth model, and inter-frame depth fusion is performed using a second depth model to generate a second depth image. Liveness detection is then performed based on this image, reducing the requirements for image accuracy and quality and improving the detection effect.

Benefits of technology

It improves the accuracy and robustness of image liveness detection in complex environments and with low-performance hardware, effectively resists various attacks, and reduces equipment costs and requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116129534B_ABST
    Figure CN116129534B_ABST
Patent Text Reader

Abstract

This specification discloses an image liveness detection method, apparatus, storage medium, and electronic device. The method includes: acquiring a target color image of a target object; performing depth estimation processing based on a first depth model to obtain a first depth image corresponding to each frame of the target color image; performing inter-frame depth fusion processing based on a second depth model to obtain a second depth image of the target object; and then performing image liveness detection processing on the target object based on the second depth image and the target color image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to an image liveness detection method, apparatus, storage medium, and electronic device. Background Technology

[0002] In recent years, biometric technology has been widely applied to people's production and daily life. For example, technologies such as facial recognition payment, facial recognition access control, facial recognition attendance, and facial recognition station entry all rely on biometrics. In biometric scenarios such as facial recognition attendance, facial recognition station entry, and facial recognition payment, the demand for image liveness detection is becoming increasingly prominent. Image liveness detection needs to verify whether the user is a real, living person operating the device, and it needs to be able to effectively resist common attack methods such as photos, face swapping, masks, occlusion, and screen capture, in order to identify fraudulent behavior and protect user rights. Summary of the Invention

[0003] This specification provides an image liveness detection method, apparatus, storage medium, and electronic device, the technical solution of which is as follows:

[0004] Firstly, this specification provides an image liveness detection method, the method comprising:

[0005] Acquire at least two color images of the target object;

[0006] Based on the first depth model, depth estimation processing is performed on each of the target color images to obtain the first depth image corresponding to each frame of the target color image;

[0007] Based on the second depth model, inter-frame depth fusion processing is performed on each of the first depth images to obtain a second depth image for the target object;

[0008] Image liveness detection processing is performed on the target object based on the second depth image and the target color image.

[0009] Secondly, this specification provides an image liveness detection device, the device comprising:

[0010] The image acquisition module is used to acquire at least two color images of the target object.

[0011] The depth estimation module is used to perform depth estimation processing on each of the target color images based on the first depth model to obtain the first depth image corresponding to each frame of the target color image;

[0012] A deep fusion module is used to perform inter-frame deep fusion processing on each of the first depth images based on a second depth model to obtain a second depth image for the target object.

[0013] The liveness detection module is used to perform image liveness detection processing on the target object based on the second depth image and the target color image.

[0014] Thirdly, this specification provides a computer storage medium storing a plurality of instructions adapted for loading by a processor and executing the above-described method steps.

[0015] Fourthly, this specification provides an electronic device that may include: a processor and a memory; wherein the memory stores a computer program adapted to be loaded by the processor and to execute the above-described method steps.

[0016] The beneficial effects of the technical solutions provided in some embodiments of this specification include at least the following:

[0017] In one or more embodiments of this specification, an electronic device obtains a first depth image by estimating the depth of multiple target color images based on a first depth model, and performs depth fusion with the inter-frame depth relationships between multiple first depth images focusing on the same object by mining and focusing on a second depth model. This results in a second depth image corresponding to a higher-precision depth estimate, thereby reducing the detection requirements for image accuracy and image quality during target color image acquisition. It can resist detection interference in complex application environments and achieve a higher-precision second depth image based on a color two-dimensional image with lower image accuracy or lower image quality. Thus, image liveness detection can be performed based on the higher-precision second depth image and the target color image, improving the detection effect of image liveness detection in complex environments and low-performance hardware environments, and enhancing the liveness detection effect and robustness. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in this specification or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a scene diagram of an image liveness detection system provided in this manual;

[0020] Figure 2 This is a flowchart illustrating an image liveness detection method provided in this specification;

[0021] Figure 3 This is a flowchart illustrating another image-based liveness detection method provided in this manual;

[0022] Figure 4 This is a flowchart illustrating another image-based liveness detection method provided in this manual;

[0023] Figure 5 This is a schematic diagram of the structure of an image liveness detection device provided in this specification;

[0024] Figure 6 This is a schematic diagram of the structure of a depth estimation module provided in this specification;

[0025] Figure 7 This is a structural diagram of a deep fusion module provided in this specification;

[0026] Figure 8 This is a schematic diagram of another image liveness detection device provided in this specification;

[0027] Figure 9 This is a schematic diagram of the structure of an electronic device provided in this specification;

[0028] Figure 10 This is a schematic diagram of the operating system and user space provided in this manual;

[0029] Figure 11 yes Figure 10 Architecture diagram of the Android operating system in China;

[0030] Figure 12 yes Figure 10 Architecture diagram of the iOS operating system. Detailed Implementation

[0031] The technical solutions in this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0032] In the description of this application, it should be understood that the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. In the description of this application, it should be noted that, unless otherwise expressly specified and limited, "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances. Furthermore, in the description of this application, unless otherwise stated, "multiple" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist; for example, A and / or B can represent: A alone, A and B simultaneously, and B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship.

[0033] In related technologies, image liveness detection scenarios such as image liveness detection and interactive recognition detection often combine multimodal image data to achieve accurate image liveness detection. These methods add more modalities to the camera, such as adding NIR and 3D modalities to the RGB modalities, and even thermal imaging modalities. Adding multiple modalities significantly enhances the performance of the entire liveness detection system and improves its ability to defend against various types of attacks. However, these methods have drawbacks: the overall cost of image liveness detection increases significantly, and the equipment requirements also increase, making them unsuitable for low-cost scenarios or those with less demanding equipment requirements. Furthermore, these techniques may require highly cooperative user actions such as head shaking and blinking under prompts for accurate image liveness detection, which is often performed in less than ideal environments. Therefore, image liveness detection technologies in these areas have significant limitations.

[0034] The present application will now be described in detail with reference to specific embodiments.

[0035] Please see Figure 1 This is a scene diagram of an image-based liveness detection system provided in this specification. Figure 1 As shown, the image liveness detection system may include at least a client cluster and a service platform 100.

[0036] The client cluster may include at least one client, such as Figure 1 As shown, it specifically includes client 1 corresponding to user 1, client 2 corresponding to user 2, ..., client n corresponding to user n, where n is an integer greater than 0.

[0037] Each client in a client cluster can be an electronic device with communication capabilities, including but not limited to: wearable devices, handheld devices, personal computers, tablets, in-vehicle devices, smartphones, computing devices, or other processing devices connected to a wireless modem. Electronic devices may have different names in different networks, such as: user equipment, access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication equipment, user agent or user device, cellular phone, cordless phone, personal digital assistant (PDA), and electronic devices in 5G networks or future evolved networks.

[0038] The service platform 100 can be a standalone server device, such as a rack-mount, blade, tower, or cabinet-type server device, or a workstation, mainframe, or other hardware device with strong computing power; or it can be a server cluster composed of multiple servers. The servers in the service cluster can be composed in a symmetrical manner, wherein each server is functionally and hierarchically equivalent in the transaction chain, and each server can provide services independently. The independent provision of services can be understood as not requiring the assistance of other servers.

[0039] In one or more embodiments of this specification, the service platform 100 can establish a communication connection with at least one client in the client cluster. Based on this communication connection, data exchange is completed during the image liveness detection process, such as the exchange of at least two frames of target color image data of an online target object. For example, the client can collect at least two frames of target color images of the target object and send them to the service platform 100. The service platform 100 then executes the image liveness detection method described in this specification to perform image liveness detection, obtains the image liveness detection result, and feeds it back to the client. Alternatively, the service platform 100 can distribute relevant depth models for image liveness detection, such as a first depth model, a second depth model, and a third depth model, to several clients to instruct the clients to execute the image liveness detection method described in this specification to perform image liveness detection, and obtain the image liveness detection result. Furthermore, the service platform 100 can obtain training sample data, such as sample images, from the client for training relevant depth models.

[0040] It should be noted that the service platform 100 establishes a communication connection with at least one client in the client cluster via a network for interactive communication. This network can be a wireless network or a wired network. Wireless networks include, but are not limited to, cellular networks, wireless LANs, infrared networks, or Bluetooth networks. Wired networks include, but are not limited to, Ethernet, universal serial bus (USB), or controller area networks. In one or more embodiments of the specification, technologies and / or formats including Hyper Text Markup Language (HTML), Extensible Markup Language (XML), etc., are used to represent data exchanged over the network (such as target compressed packets). Furthermore, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), and Internet Protocol Security (IPsec) can be used to encrypt all or some links. In other embodiments, customized and / or dedicated data communication technologies can be used to replace or supplement the aforementioned data communication technologies.

[0041] The image recognition system embodiments provided in this specification and the image recognition methods described in one or more embodiments belong to the same concept. The execution entity corresponding to the image recognition method involved in one or more embodiments of this specification can be the aforementioned service platform 100; the execution entity corresponding to the image recognition method involved in one or more embodiments of this specification can also be a client, depending on the actual application environment. The implementation process of the image recognition system embodiments can be detailed in the following method embodiments, and will not be repeated here.

[0042] based on Figure 1 The following is a detailed description of the image liveness detection method provided by one or more embodiments of this specification, as illustrated in the scene diagram.

[0043] Please see Figure 2 This document provides a flowchart illustrating an image liveness detection method according to one or more embodiments. This method can be implemented using a computer program and can run on an image liveness detection device based on the von Neumann architecture. The computer program can be integrated into an application or run as a standalone utility application. The image liveness detection device can be an electronic device.

[0044] Specifically, the image liveness detection method includes:

[0045] S102: Acquire at least two frames of target color images for the target object;

[0046] While biometrics provides convenience, it also brings new risks and challenges. The most common threat to biometric systems is liveness detection, which involves attempting to bypass image biometric verification through means such as device screens or printed photos. To detect liveness detection, anti-liveness detection technology has become an essential component of biometric applications. Image liveness detection in one or more embodiments of this specification is also a crucial part of biometrics.

[0047] In related technologies, image liveness detection is a detection method used in identity verification scenarios to determine the true physiological characteristics of an object. In facial recognition applications, image liveness detection needs to verify whether the target object is a real, living person. Image liveness detection needs to be able to effectively resist common liveness attack methods such as photos, face swapping, masks, occlusion, and screen capture, thereby helping users identify fraudulent activities and protecting users' rights.

[0048] In one or more embodiments of this specification, the image liveness detection task achieves both low-cost image acquisition and accurate image liveness detection. In related technologies, image liveness detection scenarios such as image liveness detection and interaction recognition detection often combine multimodal image data to achieve accurate image liveness detection. These methods add more modalities to the camera, such as adding NIR and 3D modalities to the RGB modalities, or even adding thermal imaging modalities. Adding multiple modalities significantly enhances the performance of the entire liveness detection system and improves its defense against various types of attacks. However, the disadvantage of these methods is that the cost of image liveness detection and the equipment requirements increase significantly, presenting considerable limitations. By executing the image liveness detection method of this specification, a first depth image corresponding to each frame of the target color image can be predicted or estimated based on at least two frames of target color images. A second depth image is obtained by depth fusion of several first depth images, and then image liveness detection processing is performed based on the target color image and the second depth image.

[0049] The target color image is a two-dimensional color image of the target object obtained based on the image liveness detection task, such as an RGB image of the target object.

[0050] In a practical image application scenario, at least two frames of target color images can be acquired using devices such as RGB cameras or monocular cameras to detect the target object to be identified or detected, based on the corresponding image liveness detection task. Typically, the target color images are two-dimensional color images.

[0051] Optionally, the acquired at least two target color images can be consecutive images, or at least two target color images acquired consecutively within a preset time period (e.g., 2 seconds) with the target frame interval as a reference.

[0052] S104: Perform depth estimation processing on each of the target color images based on the first depth model to obtain the first depth image corresponding to each frame of the target color image;

[0053] Understandably, multiple frames of target color images of the target object are input into the first depth model, and the first depth model is used to perform depth estimation processing on the target color images to obtain the first depth image corresponding to each frame of target color image.

[0054] In one or more embodiments of this specification, an initial first depth model is pre-constructed and trained for depth estimation. After the training conditions for the model are met, a first depth model is obtained, which can be applied to actual image liveness detection tasks to perform depth estimation on multiple consecutive frames of target color images of the same target object to obtain a first depth image corresponding to each target color image.

[0055] S106: Perform inter-frame depth fusion processing on each of the first depth images based on the second depth model to obtain a second depth image for the target object;

[0056] In actual image liveness detection scenarios, the second depth model is used to perform inter-frame depth fusion on the first depth images corresponding to multiple frames of target color images of the same target object to obtain a second depth image for the target object. The second depth image is the depth image after inter-frame depth fusion of the same target object. In short, the input of the second depth model is "each first depth image corresponding to multiple frames of target color images of the same target object", and the output of the second depth model is "the second depth image for the target object".

[0057] In one feasible implementation, an initial first depth model for single-frame depth estimation and an initial second depth model for multi-frame inter-frame depth fusion can be pre-created; at least one set of image sample data is acquired, the image sample data including at least two consecutive sample images for the same sample object;

[0058] The electronic device performs depth estimation training on an initial first depth model and inter-frame depth fusion training on an initial second depth model based on image sample data. After the training conditions for the model are met, the trained first depth model and second depth model are obtained.

[0059] Optionally, the image sample data can be publicly available image data obtained from relevant databases, and several image sample data of the same sample object can be grouped to form several image sample data corresponding to the same sample object. Relevant databases include one or more of CIFAR-10, CIFAR-100, Tiny ImageNet, etc., or it can be user-defined multiple sets of image sample data for different sample objects collected in actual image detection tasks for transaction scenarios. A set of image sample data for the same sample object consists of multiple image data, usually multiple frames of continuous image sample data for the sample object, such as image datasets created by labeling image data collected from the Internet.

[0060] To illustrate, acquiring multiple sets of image sample data can be done as follows: Using an RGB camera, image sample data is collected during the user's face recognition process, with 1-3 seconds of image sample data collected for each user, approximately 25-30 frames per second; the users collected can cover various ages, genders, etc.; simultaneously, image sample data of various image attack types is collected, such as device screen type image sample data (sample objects displayed on the device screen), printed photo type image sample data (sample objects contained in printed photos), and object model type image sample data (object models are a certain sample object, such as figurines, etc.), also with 1-3 seconds of images collected, approximately 25-30 frames per second; during collection, multiple sets of image sample data can cover sample images of various image attack types;

[0061] Indicatively, the initial first deep model and the initial second deep model can be built based on machine learning models. Machine learning models can include one or more fitting implementations of machine learning models such as Convolutional Neural Network (CNN), Deep Neural Network (DNN), Recurrent Neural Networks (RNN), embedding models, Gradient Boosting Decision Tree (GBDT) models, and Logistic Regression (LR) models. The model training process for the initial first deep model and the initial second deep model can be referred to the explanations of other embodiments in this specification.

[0062] S108: Perform image liveness detection processing on the target object based on the second depth image and the target color image.

[0063] Understandably, after obtaining the second depth image based on multiple frames of target color images of the target object, image liveness detection processing can be performed based on the second depth image and the target color image to determine the liveness detection type of the target object.

[0064] In one feasible implementation, the electronic device performing the image liveness detection processing of the target object based on the second depth image and the target color image may be as follows:

[0065] The electronic device inputs the second depth image and the target color image into the liveness detection model, and outputs a first liveness probability value for the second depth image and a second liveness probability value for the target color image; then, based on the first liveness probability value and the second liveness probability value, it determines the liveness detection type of the target object.

[0066] Understandably, a pre-trained liveness detection model can be used with "the second depth image and the target color image" as input to output a first liveness probability value for the second depth image and a second liveness probability value for the target color image.

[0067] The liveness probability value can be understood as the liveness classification probability of the corresponding image by the liveness detection model.

[0068] After obtaining "the first liveness probability value corresponding to the second depth image and the second liveness probability value for the target color image", the electronic device can determine the target liveness probability based on the first liveness probability value and the second liveness probability value.

[0069] Optionally, the rule for determining the target liveness probability can be: selecting one of the first liveness probability value and the second liveness probability value as the target liveness probability, such as taking the maximum probability value between the first liveness probability value and the second liveness probability value as the target liveness probability;

[0070] Optionally, the rule for determining the target liveness probability can be: pre-setting a first weighting factor for the second depth image and a second weighting factor for the target color image, and using a weighted fusion method to obtain the final target liveness probability.

[0071] For illustration, assuming the first liveness probability is p1 and the second liveness probability is p2, and a first weighting factor is pre-set as a for the second depth image, and a second weighting factor is pre-set as b for the target color image, then the target liveness probability P can be calculated using the following formula:

[0072] P = p1*a + p2*b

[0073] Optionally, the first weighting factor and the second weighting factor are both 1, such as the first weighting factor being 0.5 and the second weighting factor being 0.5.

[0074] Optionally, there are usually multiple target color images. When performing image liveness detection processing, one of the multiple target color images can be included in the reference, that is, one of the multiple target color images is selected, and image liveness detection processing is performed on the target object based on the second depth image and the selected target color image.

[0075] Optionally, there are usually multiple target color images. When performing image liveness detection processing, multiple target color images can be included as a reference. The electronic device inputs the second depth image and all target color images into the liveness detection model, outputting the first liveness probability value corresponding to the second depth image and the second liveness probability value for each target color image. Then, the second liveness probability values ​​of each target color image are fitted to obtain the optimal second liveness probability value. For example, the optimal second liveness probability value can be selected by calculating the average, median, maximum, and minimum values ​​of all second liveness probability values.

[0076] Furthermore, after obtaining the target liveness probability, the target liveness probability can be compared with the target threshold to determine the liveness detection result.

[0077] Optionally, if the target liveness probability is greater than the target threshold, then the target object is determined to be a live object type;

[0078] Optionally, if the target liveness probability is less than or equal to the target threshold, then the target object is determined to be an attack target type.

[0079] In one or more embodiments of this specification, the liveness detection model can be created based on a machine learning model. The liveness detection model can be implemented by fitting one or more of the following models: Convolutional Neural Network (CNN), Deep Neural Network (DNN), Recurrent Neural Networks (RNN), Embedding model, Gradient Boosting Decision Tree (GBDT) model, Logistic Regression (LR) model, etc.

[0080] Understandably, an initial liveness detection model can be created in advance, and a large number of sample training images can be obtained as the model image training set for the liveness detection model. The model image training set includes depth sample images and color sample images corresponding to the same sample object. Then, the initial liveness detection model is trained based on the model image training set until the model training termination condition is met, and a trained liveness detection model is obtained.

[0081] In one or more embodiments of this specification, an electronic device obtains a first depth image by estimating the depth of multiple target color images based on a first depth model, and performs depth fusion with the inter-frame depth relationships between multiple first depth images focusing on the same object by mining and focusing on a second depth model. This results in a second depth image corresponding to a higher-precision depth estimate, thereby reducing the detection requirements for image accuracy and image quality during target color image acquisition. It can resist detection interference in complex application environments and achieve a higher-precision second depth image based on a color two-dimensional image with lower image accuracy or lower image quality. Thus, image liveness detection can be performed based on the higher-precision second depth image and the target color image, improving the detection effect of image liveness detection in complex environments and low-performance hardware environments, and enhancing the liveness detection effect and robustness.

[0082] Please see Figure 3 , Figure 3 This is a schematic flowchart illustrating another embodiment of an image liveness detection method proposed in one or more embodiments of this specification. Specifically:

[0083] S202: Create the initial first depth model and the initial second depth model;

[0084] S204: Acquire at least one set of image sample data, the image sample data including at least two consecutive sample images for the same sample object;

[0085] For details, please refer to the relevant explanations in other embodiments of this specification, which will not be repeated here.

[0086] S206: Input each of the sample images in the image sample data into the initial first depth model and output the sample depth estimation map corresponding to each of the sample images;

[0087] Indicatively, the initial first depth model can be created based on a machine learning model. In some embodiments, the model architecture corresponding to the initial first depth model can be the UNET model architecture. The initial first depth model based on the UNET model architecture can include at least two parts: an Encoder and a Decoder.

[0088] Furthermore, during model training, the input to the first depth model is a single-frame sample image, and the output of the first depth model is the sample depth estimation map corresponding to the sample image. The sample depth estimation map can be understood as a depth map containing pixel depth information.

[0089] In one feasible implementation, the training process of the initial first deep model can be as follows:

[0090] The electronic device predetermines the label depth map corresponding to each of the sample images;

[0091] In each round of training for the initial first depth model: each sample image is input into the initial first depth model for depth estimation, and the sample depth estimation map corresponding to each sample image is output; for example, in a certain round of training, b sample images (sample images are usually two-dimensional images) for the same sample object A are input into the initial first depth model for depth estimation. The initial first depth model predicts or estimates the pixel depth features of each pixel in the sample image, thereby generating the sample depth estimation map corresponding to the sample image.

[0092] Furthermore, during the training of the initial first depth model, a label depth map corresponding to each image sample data is pre-determined. The label depth map is used in the backpropagation training process of the initial first depth model. During the backpropagation training process, the model parameters of the initial first depth model are adjusted by backpropagation based on the label depth map and the sample depth estimation map output by the current initial first depth model.

[0093] The following is an illustrative explanation of the label depth map corresponding to the sample image:

[0094] In one feasible implementation, the electronic device can pre-set the label depth of the sample images after acquiring several sample images corresponding to the image sample data, so that the label depth map corresponding to each sample image can be determined in the subsequent training process.

[0095] Optionally, a label depth map can be determined for the images of all sample objects based on methods for generating image depth of sample objects in related technologies. For example, three-dimensional object (such as facial object) reconstruction (such as 3DMM model) and three-dimensional depth estimation techniques can be used to perform three-dimensional depth reconstruction on the sample images corresponding to all sample objects, and the resulting depth map is also the label depth map of the sample image.

[0096] Optionally, the electronic device can determine the label depth map corresponding to the sample image by combining the image sample type at the time of sample image acquisition and using different depth map settings after acquiring each sample image based on the image liveness detection task.

[0097] As an illustration, in the context of image liveness detection, the image sample type can include at least two types: attack image type and liveness image type.

[0098] Sample images of the attack image type can be sample images collected using common liveness attack methods such as photos, face swaps, masks, occlusion, and screen captures. These sample images are usually classified as attack images compared to sample images collected in the environment of a real live object.

[0099] Liveness image type can be understood as the type of sample image collected in the environment of a real living object.

[0100] Furthermore, when the electronic device executes the process of setting the label depth map corresponding to the image sample data based on the image sample type, it may do so in the following ways:

[0101] The electronic device obtains the first sample image corresponding to the attack image type from all the collected sample images, and sets the depth pixel value of the first label depth map corresponding to the first sample image as the target depth pixel value.

[0102] For example, for the first sample image corresponding to the attack image type, a first label depth map can be generated with all depth pixel values ​​equal to the target depth pixel values. For instance, an image with all depth pixel values ​​of 0 can be generated as a depth map, and it can be assumed that the depth of the first sample image corresponding to the attack image type is 0.

[0103] The electronic device obtains the second sample image corresponding to the liveness image type from all the collected sample images, and calls the target image depth service to determine the second label depth map corresponding to the second sample image.

[0104] Understandably, the label depth map corresponding to the sample image includes a first label depth map and a second label depth map;

[0105] Understandably, methods for generating image depth of sample objects based on related technologies determine the second label depth map for all second sample images. For example, three-dimensional object (such as facial object) reconstruction (such as 3DMM model) and three-dimensional depth prediction technology are used to perform three-dimensional depth reconstruction on the sample images corresponding to all sample objects. The resulting depth map is also the label depth map of the sample image.

[0106] In one or more embodiments of this specification, the label depth map of the sample image serves as a supervision signal during model training to optimize the model training effect.

[0107] S208: Based on the sample depth estimation map, perform depth estimation training on the initial first depth model until the initial first depth model completes training, and obtain the trained first depth model.

[0108] The sample depth estimation map is the output result of the initial first depth model performing depth estimation processing on the sample image.

[0109] During the training of the initial first depth model in each round, the label depth map corresponding to the sample image is determined. The sample image is then input into the initial first depth model for depth estimation processing, and the corresponding sample depth estimation map is output. The model parameters of the initial first depth model are then adjusted by backpropagation, which is combined with the output sample depth estimation map and the label depth map as a supervision signal.

[0110] Furthermore, during the training process of each initial first depth model, the electronic device calculates the pixel estimation loss focused on the pixel dimension based on the sample depth estimation map and the label depth map using the set loss calculation function, and adjusts the model parameters of the initial first depth model based on the pixel estimation loss through backpropagation, such as adjusting the connection weights and / or thresholds between neurons in each layer of the model based on the pixel estimation loss.

[0111] In a schematic manner, the electronic device can set a first loss calculation formula for the initial first depth model. The first loss calculation formula is the loss function of the initial first depth model. By inputting the sample depth estimation map and the label depth map into the first loss calculation formula during each round of model training, the pixel estimation loss is determined. The model parameters of the initial first depth model are adjusted based on the pixel estimation loss until the initial first depth model meets the model end training condition, and the trained first depth model is obtained.

[0112] Optionally, the first loss calculation formula satisfies the following formula:

[0113]

[0114] Among them, the Loss A Estimate the loss for each pixel. p is the estimated depth value of the i-th depth pixel in the sample depth estimation map. i γ is the label depth value of the i-th depth pixel in the label depth map, γ is the loss adaptive parameter, i is an integer, and I is the total number of pixels in the sample depth estimation map.

[0115] Indicatively, the magnitude of the estimated depth value of the i-th depth pixel in the sample depth estimation map can characterize the prediction probability value of the initial first depth model for the corresponding pixel. In some embodiments, the prediction probability value is in the range of 0-1.

[0116] Indicatively, γ is a loss adaptive parameter, which can be understood as an adjustment parameter for sample loss.

[0117] In one or more embodiments of this specification, the first loss calculation formula adopts the above form. The pixel estimation loss calculated by the first loss calculation formula focuses on the pixel dimension of the depth map estimation. Compared with the distance reconstruction loss in related technologies, the pixel estimation loss obtained by the above first loss calculation formula can focus on or even perceive the regions in the input sample image that are difficult to reconstruct in depth, and adaptively increase and adjust their weights, thereby improving the depth estimation effect during the model training process and obtaining a better output, namely the depth estimation map.

[0118] S210: Input the sample depth estimation map corresponding to each of the sample images into the initial second depth model to output the sample depth fusion map;

[0119] The sample depth estimation map is the output of the initial first depth model. When performing depth reconstruction on the initial second depth model, it is necessary to accumulate sample depth estimation maps corresponding to several frames of sample images of the same sample object based on the initial first depth model, and then input these sample depth estimation maps of the same sample object into the initial second depth model.

[0120] Understandably, during each round of training of the initial second depth model, the electronic device inputs the sample depth estimation map corresponding to each sample image of the same sample object into the initial second depth model for depth reconstruction processing, and obtains the depth map estimation feature corresponding to the sample depth estimation map and the sample depth reconstruction map corresponding to the sample depth estimation map. During each round of training, the sample depth reconstruction map is output, and the model parameters of the initial second depth model are adjusted by backpropagation in combination with the output sample depth reconstruction map and the depth map estimation feature.

[0121] To illustrate, the depth map estimation features are obtained by extracting depth features from the depth estimation map of each frame sample. By extracting depth features from the depth estimation map of each frame sample, the depth map estimation features of each frame sample depth estimation map can be obtained.

[0122] In one or more embodiments of this specification, the initial second depth model can be regarded as an inter-frame relation depth model or network. During the training process of the initial second depth model, the features of the corresponding sample depth estimation maps of multiple sample images are calculated, and then the depth relationship between the multiple frame depth estimation maps is mined. The inter-frame relation matrix in the intermediate depth reconstruction process of the model is used to fuse the multiple frame depth estimation maps.

[0123] S212: Based on the sample depth fusion map, perform inter-frame depth fusion training on the initial second depth model until the initial second depth model completes training, and obtain the trained second depth model.

[0124] Understandably, during the training of each round of the initial second depth model, inter-frame depth fusion will be performed based on the depth map estimation features corresponding to each sample depth estimation map to output a sample depth fusion map, and the model parameters of the initial second depth model will be adjusted based on the sample depth reconstruction map and the sample depth estimation map.

[0125] In one feasible implementation, the initial second deep model structure comprises at least a first deep encoding network, a second deep decoding network, and a third self-attention network;

[0126] Indicatively, the first deep coding network can be used as a deep feature encoder, such as a ResNet18 network, to extract deep features from the model input;

[0127] In illustrative terms, the second depth decoding network can be used for depth reconstruction. In some embodiments, the second depth decoding network can be a decoder, with the output of the first depth encoding network serving as the input to the second depth decoding network to obtain the reconstructed depth map, which is mainly used for subsequent network parameter adjustments.

[0128] Indicatively, the third self-attention network can be understood as a network module based on self-attention. In some embodiments, the third self-attention network can be a non-local self-attention network module. The input of the third self-attention network is the depth map estimation features of the multi-frame depth estimation map. Based on the input features, it performs inter-frame relationship prediction and outputs an inter-frame relationship matrix and a reference depth map.

[0129] The following explains the training process of the initial second-depth model:

[0130] In each round of training for the initial second depth model: the sample depth estimation map corresponding to each sample image of the same sample object is input into the first depth coding network, and the first depth coding network extracts features from each sample depth estimation map to obtain the depth map estimation features corresponding to each sample depth estimation map;

[0131] In a schematic way, the sample depth estimation maps corresponding to each sample image of the same sample object are input into the initial second depth model. The first depth encoding network of the initial second depth model extracts features from each sample depth estimation map, and the depth map estimation features corresponding to each sample depth estimation map are obtained in turn. It can be understood that the output of the first depth encoding network (such as ResNet18) is the first depth encoding network, and the input is the sample depth estimation map.

[0132] Furthermore, the electronic device controls the initial second depth model to input the depth map estimation features into the second depth decoding network to obtain the sample depth reconstruction map corresponding to the sample depth estimation map;

[0133] Indicatively, the initial second depth model of the electronic device control is reconstructed by the second depth decoding network (such as Decoder) based on several depth map estimation features, resulting in a sample depth reconstruction map corresponding to the reconstructed sample depth estimation map. Here, the sample depth reconstruction map is mainly used for subsequent network parameter adjustment, which can be understood as calculating the model loss based on the sample depth reconstruction map after depth reconstruction and then adjusting the subsequent network parameters.

[0134] Furthermore, the electronic device inputs the depth map estimation features corresponding to each sample depth estimation map into a third self-attention network to obtain a reference depth map and an inter-frame relationship matrix, wherein the reference depth map is one of the sample depth estimation maps;

[0135] Indicatively, by controlling the initial second depth model, a third self-attention network (such as a non-local self-attention network) estimates features based on the depth maps corresponding to the depth estimation maps of each sample object, and mines the inter-frame relationships between multiple sample images to obtain a reference depth map and an inter-frame relationship matrix.

[0136] In a schematic representation, the reference depth map is one of several input sample depth estimation maps. Using the selected sample depth estimation map as the reference depth image, the inter-frame relationships between other sample depth estimation maps and the reference depth image are mined and represented by an inter-frame relationship matrix, assumed to be denoted by w. n,j In the diagram, n represents the nth row of the inter-frame relation matrix, and w n,j The label j represents the j-th column of the inter-frame relation matrix. In the inter-frame relation matrix w, w... n,j The coefficient represents the relationship between the j-th pixel in the fused depth map and the n-th pixel in the selected reference depth map. The fused depth map is the depth map other than the reference depth map among all the depth estimation maps of the same sample object.

[0137] Additionally, n and j can be understood as the labels of the pixels in the corresponding depth map.

[0138] Furthermore, the electronic device performs inter-frame depth fusion based on the inter-frame relationship matrix and the reference depth map using the initial second depth model to output a sample depth fusion map.

[0139] Indicatively, during the training process of the initial second depth model, after determining the inter-frame relation matrix and the reference depth map through the initial second depth model, the inter-frame relation matrix feeds back the weights corresponding to the fused depth map based on the reference depth image. The weighted fusion is performed by combining the parameters of the inter-frame relation matrix with the reference depth image to obtain the weighted fused sample depth map.

[0140] In one feasible implementation, the step of performing inter-frame depth fusion based on the inter-frame relation matrix and the reference depth map to output a sample depth fusion map can be:

[0141] Based on the inter-frame relationship matrix and the reference depth map, a second inter-frame fusion calculation formula is used to perform weighted fusion to obtain the sample depth fusion map.

[0142] The second inter-frame fusion calculation formula is used to perform weighted fusion of the depth maps corresponding to several sample images based on the inter-frame relationship matrix and the reference depth map, to obtain a sample depth fusion map.

[0143] The second inter-frame fusion calculation formula satisfies the following formula:

[0144]

[0145] Wherein, the depth multi The sample depth fusion map is defined as follows: N is the total number of pixels in the reference depth map, n is an integer, and the depth is... n w is the depth pixel value of the nth pixel in the original depth map. n,j In the figure, n represents the nth row of the inter-frame relation matrix, and w n,j The label j represents the j-th column of the inter-frame relation matrix, where w... n,j The coefficient represents the relationship between the j-th pixel in the fused depth map and the n-th pixel in the reference depth map. The fused depth map is the depth map other than the reference depth map among all the depth estimation maps of the same sample object.

[0146] Schematic, based on the second inter-frame fusion calculation formula described above, using each pixel of the reference depth map as a reference, and combining it with the inter-frame relationship matrix, according to w in the inter-frame relationship matrix... n,jThe relationship coefficient between the j-th pixel in the fused depth map and the n-th pixel in the reference depth map can be used to obtain the weight of each pixel. Based on the above second inter-frame fusion calculation formula, the fused depth value of each fused depth pixel can be obtained. By determining all fused depth values, the sample depth fused map can be obtained.

[0147] Understandably, the above method can output sample depth images in each round of model training for the initial second depth model, and simultaneously adjust the model parameters of the initial second depth model based on the sample depth reconstruction map and the sample depth estimation map.

[0148] In one feasible implementation, adjusting the model parameters of the initial second depth model based on the sample depth reconstruction map and the sample depth estimation map can be:

[0149] The electronic device can input the depth estimation map and the depth reconstruction map of each sample for the same sample object into the third loss calculation formula to determine the depth reconstruction loss; then, based on the depth reconstruction loss, the model parameters of the initial second depth model are adjusted. The model parameters of the initial second depth model are adjusted through backpropagation, such as adjusting the connection weights and / or thresholds between neurons in each layer of the model based on the pixel estimation loss, until the initial second depth model meets the model end training condition, and the trained second depth model is obtained.

[0150] The third loss calculation formula satisfies the following formula:

[0151]

[0152] Where Loss B is the depth reconstruction loss, l is an integer, L is the total number of depth estimation maps corresponding to the same sample object, and I pred-l For the l-th sample depth reconstruction map, the I GT-l This is the depth estimation map for the l-th sample.

[0153] S214: Based on the image liveness detection task, acquire at least two frames of target color images of the target object;

[0154] S216: Perform depth estimation processing on each of the target color images based on the first depth model to obtain a first depth image corresponding to each frame of the target color image; perform inter-frame depth fusion processing on each of the first depth images based on the second depth model to obtain a second depth image for the target object; perform image liveness detection processing on the target object based on the second depth image and the target color image.

[0155] For details, please refer to the method steps in other embodiments of this specification, which will not be repeated here.

[0156] In one or more embodiments of this specification, an electronic device obtains a first depth image by estimating the depth of multiple target color images based on a first depth model, and performs depth fusion with the first depth images that are mined and focused on the same object using a second depth model to mine and focus on the inter-frame depth relationships. This yields a second depth image corresponding to a higher-precision depth estimate, thereby reducing the detection requirements for image accuracy and image quality during target color image acquisition. It can resist detection interference in complex application environments and achieve a higher-precision second depth image based on a color 2D image with lower image accuracy or lower image quality. Therefore, image liveness detection can be performed based on the higher-precision second depth image and the target color image. This technology improves the detection performance of liveness detection in complex environments and on low-performance hardware, enhancing both the effectiveness and robustness of liveness detection. Furthermore, it introduces an innovative first loss calculation formula for single-frame depth estimation, focusing on pixel estimation loss. This first loss formula better addresses regions in each image that are difficult to fit, resulting in better single-frame depth quality. Additionally, during inter-frame depth fusion, the initial second depth model is instructed to calculate features of multi-frame depth estimation images and mine the relationships between multiple frames to obtain an inter-frame relationship matrix. Based on this matrix, multi-frame depth maps are fused, resulting in better depth fusion performance and improved model output depth quality, leading to a second depth map with better depth estimation performance.

[0157] Please see Figure 4 , Figure 4 This is a schematic flowchart illustrating another embodiment of an image liveness detection method proposed in one or more embodiments of this specification. Specifically:

[0158] S302: Acquire at least two frames of target color images for the target object;

[0159] S304: Perform depth estimation processing on each of the target color images based on the first depth model to obtain a first depth image corresponding to each frame of the target color image; perform inter-frame depth fusion processing on each of the first depth images based on the second depth model to obtain a second depth image for the target object;

[0160] S306: Perform quality enhancement processing on the second depth image based on the third depth model to obtain a quality-enhanced third depth image;

[0161] Understandably, in real-world image liveness detection scenarios, after depth estimation and inter-frame depth fusion based on the first and second depth models, a fused second depth image is obtained. Considering objective factors such as the limitations of image quality in multiple color images of the same object and bottlenecks in model recognition processing, there is a certain probability that the fused second depth image will have discontinuous depth values ​​in local or small areas. Based on this, depth quality enhancement can be performed on the fused second depth image. A third depth model can be used to enhance depth quality and optimize the subsequent detection interference caused by the aforementioned objective factors. Further improving the data quality on the fused depth image can yield a depth estimation map with higher depth quality, i.e., the third depth image, thereby improving the accuracy and effectiveness of subsequent image liveness detection.

[0162] Understandably, the third depth model is used to perform quality enhancement on the depth image after fusing the depth estimation maps corresponding to multiple color images, in order to resist detection interference caused by objective factors and improve the quality of the fused depth image. The input of the third depth model is the second depth image output from the second depth model, and the output of the third depth model is the third depth image after quality enhancement processing.

[0163] In one feasible implementation, an initial third-depth model is pre-created, and at least one sample depth fusion map corresponding to the initial second-depth model is obtained. All or part of the sample depth fusion maps output during each round of training of the initial second-depth model are used as model training samples. The initial third-depth model is then subjected to quality enhancement training using these training samples until the initial third-depth model completes its training, resulting in a trained third-depth model. In other words, the steps involve performing quality enhancement training on the initial third-depth model based on each sample depth fusion map to obtain the trained third-depth model.

[0164] In one or more embodiments of this specification, the initial third deep model can be built based on a machine learning model. The machine learning model may include one or more fitting implementations of machine learning models such as Convolutional Neural Network (CNN), Deep Neural Network (DNN), Recurrent Neural Networks (RNN), embedding models, Gradient Boosting Decision Tree (GBDT) models, and Logistic Regression (LR) models. During the training process of the initial third deep model, an error backpropagation algorithm is introduced in combination with the model loss to optimize parameters and improve the processing effect of the machine learning model.

[0165] As an illustration, the initial third-depth model can adopt the UNET model structure built based on a machine learning model.

[0166] Optionally, the step of performing quality enhancement training on the initial third depth model based on the depth fusion maps of each of the samples to obtain the trained third depth model can be:

[0167] During each round of model training for the initial third depth model, the electronic device can acquire depth map enhancement labels corresponding to the sample depth fusion map.

[0168] In some embodiments, the depth image enhancement label can be a label depth map corresponding to a sample image of a sample object. The sample depth fusion map is generated based on multiple frames of sample images of the same sample object, and the label depth map can be one of these multiple frames of sample images. For example, the label depth map of the sample image corresponding to the reference depth map in the initial second depth model can be selected. It is understood that the reference depth map is usually one of the sample depth fusion maps corresponding to multiple frames of sample images, so the label depth map of the sample image corresponding to the reference depth map in the initial second depth model can be selected. For the initial third depth model, the depth image enhancement label serves as the intensity optimization target in the model processing stage of the initial third depth model.

[0169] The electronic device trains the initial third depth model by first perturbing the pixel depth fusion maps of each sample to obtain the perturbed sample depth fusion maps; the electronic device inputs each sample depth fusion map into the initial third depth model for quality enhancement processing, outputs the sample enhanced depth map corresponding to the sample depth fusion map, and adjusts the model parameters of the initial third depth model based on the sample enhanced depth map and the depth map enhancement label until the initial third depth model completes training, thus obtaining the trained third depth model;

[0170] Understandably, pixel perturbation is first applied to the sample depth estimation maps corresponding to each sample image of the same sample object to obtain pixel-perturbed sample depth estimation maps. Schematic, pixel perturbation of the sample depth estimation images can improve the depth reconstruction and depth fusion effect of the depth model during training, simulate attack interference in a real environment, and reconstruct the pixel-perturbed sample depth estimation maps during the model training phase so that the model has a better depth quality enhancement capability after training.

[0171] Optionally, pixel perturbation processing can be performed on the sample depth estimation map using pixel perturbation algorithms in related technologies. For example, differential evolution can be used to perturb the depth values ​​of a few pixels in the sample depth estimation map (such as perturbing only a few pixels out of 1024 pixels).

[0172] Schematic illustration: The adjustment of model parameters for the initial third depth model based on the enhanced depth map of the sample and the enhanced labels of the depth map can be:

[0173] During the training of the initial third-depth model in each round: after the initial third-depth network outputs sample-enhanced depth maps, the sample-enhanced depth maps and depth map enhancement labels are input into the fourth loss calculation formula to determine the quality enhancement loss; the model parameters of the initial third-depth model are adjusted based on the quality enhancement loss.

[0174] The fourth loss calculation formula satisfies the following formula:

[0175]

[0176] Wherein, Loss C is the mass enhancement loss, and I re For the sample enhancement depth map, the I GT Enhance the labels on the depth map.

[0177] Indicatively, in each round: after the initial third-depth network outputs a sample-enhanced depth map, the quality enhancement loss is calculated based on the sample-enhanced depth map and the depth map enhancement labels. Then, the model parameters of the initial third-depth model are adjusted based on the quality enhancement loss. The model parameters of the initial third-depth model are adjusted through backpropagation, such as adjusting the connection weights and / or thresholds between neurons in each layer of the model based on pixel estimation loss, until the initial third-depth model meets the model termination condition, thus obtaining the trained third-depth model.

[0178] Understandably, after training and generating the third depth model, in the practical application stage, after obtaining the second depth image for the target object, the second depth image can be input into the third depth model for quality enhancement processing to obtain the quality-enhanced third depth image.

[0179] S308: Using the third depth image as the second depth image, perform the step of image liveness detection processing on the target object based on the second depth image and the target color image.

[0180] Understandably, after obtaining the third depth image after quality enhancement processing, the electronic device can use the third depth image as the second depth image and perform image liveness detection processing on the target object based on the second depth image and the target color image. For details, please refer to the method steps of other embodiments of this specification, which will not be repeated here.

[0181] In one or more embodiments of this specification, an electronic device obtains a first depth image by estimating the depth of multiple target color images based on a first depth model, and performs depth fusion with the inter-frame depth relationships between multiple first depth images that focus on the same object by mining and focusing on them using a second depth model. This yields a second depth image corresponding to a higher-precision depth estimate, thereby reducing the detection requirements for image accuracy and image quality during target color image acquisition. It can resist detection interference in complex application environments and achieve a higher-precision second depth image based on color 2D images with lower image accuracy or lower image quality. Thus, image liveness detection can be performed based on the higher-precision second depth image and the target color image, improving the detection effect of image liveness detection in complex environments and low-performance hardware environments, and enhancing the liveness detection effect and robustness. Furthermore, after inter-frame depth fusion, a third depth model is introduced for frame rate enhancement, which can further improve the data quality of the fused depth estimation image, effectively resist environmental interference, and improve the stability and accuracy of liveness detection.

[0182] The following will combine Figure 5 This manual provides a detailed description of the image liveness detection device provided. It should be noted that... Figure 5 The image liveness detection device shown is used to perform the present application. Figures 1-4 The methods of the embodiments shown are illustrated only in connection with this specification for ease of explanation; for specific technical details not disclosed, please refer to this application. Figures 1-4 The example shown.

[0183] Please see Figure 5 This diagram illustrates the structure of the image liveness detection device described in this specification. The image liveness detection device 1 can be implemented as all or part of a user terminal through software, hardware, or a combination of both. According to some embodiments, the image liveness detection device 1 includes an image acquisition module 11, a depth estimation module 12, a depth fusion module 13, and a liveness detection module 14, specifically used for:

[0184] Image acquisition module 11 is used to acquire at least two frames of target color images of the target object based on the image liveness detection task;

[0185] Depth estimation module 12 is used to perform depth estimation processing on each of the target color images based on the first depth model to obtain the first depth image corresponding to each frame of the target color image;

[0186] The deep fusion module 13 is used to perform inter-frame deep fusion processing on each of the first depth images based on the second depth model to obtain a second depth image for the target object;

[0187] The liveness detection module 14 is used to perform image liveness detection processing on the target object based on the second depth image and the target color image.

[0188] Optional, such as Figure 8 As shown, the device 1 includes:

[0189] Model training module 15 is used to create the initial first depth model and the initial second depth model;

[0190] The model training module 15 is used to acquire at least one set of image sample data, the image sample data including at least two consecutive sample images for the same sample object;

[0191] The model training module 15 is used to instruct the depth estimation module 12 to perform depth estimation training on the initial first depth model based on the image sample data and to instruct the depth fusion module 13 to perform inter-frame depth fusion training on the initial second depth model, so as to obtain the trained first depth model and the second depth model.

[0192] Optionally, the depth estimation module 12 is used to input each of the sample images in the image sample data into the initial first depth model to output the sample depth estimation map corresponding to each of the sample images, and to perform depth estimation training on the initial first depth model based on the sample depth estimation map until the initial first depth model completes training to obtain the trained first depth model.

[0193] The deep fusion module 13 is used to input the sample depth estimation map corresponding to each sample image into the sample depth fusion map output by the initial second depth model, and perform inter-frame depth fusion training on the initial second depth model based on the sample depth fusion map until the initial second depth model completes training to obtain the trained second depth model.

[0194] Optional, such as Figure 6 As shown, the depth estimation module 12 includes:

[0195] The depth reconstruction unit 121 is used to determine the label depth map corresponding to each of the sample images; input each of the sample images into the initial first depth model for depth estimation processing, and output the sample depth estimation map corresponding to each of the sample images.

[0196] The parameter adjustment unit 122 is used to determine the pixel estimation loss based on the label depth map and adjust the model parameters of the initial first depth map based on the pixel estimation loss.

[0197] Optionally, the parameter adjustment unit 122 is used for:

[0198] The sample depth estimation map and the label depth map are input into the first loss calculation formula to determine the pixel estimation loss;

[0199] The first loss calculation formula satisfies the following formula:

[0200]

[0201] Among them, the Loss A Estimate the loss for each pixel. p is the estimated depth value of the i-th depth pixel in the sample depth estimation map. i γ is the label depth value of the i-th depth pixel in the label depth map, γ is the loss adaptive parameter, i is an integer, and I is the total number of pixels in the sample depth estimation map.

[0202] Optionally, the model training module 15 is used for:

[0203] Based on the image liveness detection task, each sample image is acquired, and the image sample type of the sample image is determined;

[0204] The label depth map corresponding to the sample image is set based on the image sample type.

[0205] Optionally, the image sample types include attack image types and liveness image types, and the model training module 15 is used for:

[0206] Obtain the first sample image corresponding to the attack image type, and set the depth pixel value of the first label depth map corresponding to the first sample image as the target depth pixel value;

[0207] Obtain the second sample image corresponding to the liveness image type, and call the target image depth service to determine the second label depth map corresponding to the second sample image.

[0208] Optional, such as Figure 7 As shown, the deep fusion module 13 includes:

[0209] The depth reconstruction unit 131 is used to input the sample depth estimation map corresponding to each sample image into the initial second depth model for depth reconstruction processing, and obtain the depth map estimation feature corresponding to the sample depth estimation map and the sample depth reconstruction map corresponding to the sample depth estimation map.

[0210] The parameter adjustment unit 132 is used to perform inter-frame depth fusion to output a sample depth fusion map based on the depth map estimation features corresponding to each sample depth estimation map, and to adjust the model parameters of the initial second depth model based on the sample depth reconstruction map and the sample depth estimation map.

[0211] Optionally, the initial second deep model includes at least a first deep encoding network, a second deep decoding network, and a third self-attention network.

[0212] The depth reconstruction unit 131 is configured to: input the sample depth estimation maps corresponding to each sample image of the same sample object into the first depth coding network to obtain depth map estimation features corresponding to each sample depth estimation map; and input the depth map estimation features into the second depth decoding network to obtain the sample depth reconstruction map corresponding to the sample depth estimation map.

[0213] The parameter adjustment unit 132 is used to input the depth map estimation features corresponding to each of the sample depth estimation maps into the third self-attention network to obtain a reference depth map and an inter-frame relationship matrix, wherein the reference depth map is one of the sample depth estimation maps.

[0214] Inter-frame depth fusion is performed based on the inter-frame relationship matrix and the reference depth map to output a sample depth fusion map.

[0215] Optionally, the parameter adjustment unit 132 is used to: perform weighted fusion based on the inter-frame relationship matrix and the reference depth map using a second inter-frame fusion calculation formula to obtain a sample depth fusion map;

[0216] The second inter-frame fusion calculation formula satisfies the following formula:

[0217]

[0218] Wherein, the depth multi The sample depth fusion map is defined as follows: N is the total number of pixels in the reference depth map, n is an integer, and the depth is... n w is the depth pixel value of the nth pixel in the original depth map. n,j In the figure, n represents the nth row of the inter-frame relation matrix, and w n,j The label j represents the j-th column of the inter-frame relation matrix, where w...n,j The coefficient represents the relationship between the j-th pixel in the fused depth map and the n-th pixel in the reference depth map. The fused depth map is the depth map other than the reference depth map among all the depth estimation maps of the same sample object.

[0219] Optionally, the parameter adjustment unit 132 is used to: input the sample depth estimation map and the sample depth reconstruction map for each sample object into the third loss calculation formula to determine the depth reconstruction loss;

[0220] The model parameters of the initial second depth model are adjusted based on the depth reconstruction loss.

[0221] The third loss calculation formula satisfies the following formula:

[0222]

[0223] Where Loss B is the depth reconstruction loss, l is an integer, L is the total number of depth estimation maps corresponding to the same sample object, and I pred-l For the l-th sample depth reconstruction map, the I GT -l represents the depth estimation map of the l-th sample.

[0224] Optionally, the device 1 is further configured to: perform quality enhancement processing on the second depth image based on the third depth model to obtain a quality-enhanced third depth image;

[0225] The liveness detection module 14 is also used for:

[0226] Using the third depth image as the second depth image, perform image liveness detection processing on the target object based on the second depth image and the target color image.

[0227] Optionally, the device 1 is further used for:

[0228] Create an initial third-depth model;

[0229] Obtain at least one sample depth fusion map corresponding to the initial second depth model of the second depth model;

[0230] The initial third depth model is trained with enhanced quality based on the depth fusion maps of the samples to obtain the trained third depth model.

[0231] Optionally, the device 1 is further configured to: acquire depth map enhancement labels corresponding to the sample depth fusion map;

[0232] Pixel perturbation processing is performed on each of the sample depth fusion maps to obtain the sample depth fusion map after pixel perturbation;

[0233] Each of the sample depth fusion maps is input into the initial third depth model for quality enhancement processing, and the sample enhanced depth map corresponding to the sample depth fusion map is output. The model parameters of the initial third depth model are adjusted based on the sample enhanced depth map and the enhancement label of the depth map until the initial third depth model completes training, and the trained third depth model is obtained.

[0234] Optionally, the device 1 is further configured to: input the sample enhancement depth map and the depth map enhancement label into the fourth loss calculation formula to determine the quality enhancement loss;

[0235] The model parameters of the initial third depth model are adjusted based on the aforementioned quality enhancement loss.

[0236] The fourth loss calculation formula satisfies the following formula:

[0237]

[0238] Wherein, Loss C is the mass enhancement loss, and I re For the sample enhancement depth map, the I GT Enhance the labels on the depth map.

[0239] Optionally, the liveness detection module 14 is used for:

[0240] The second depth image and the target color image are input into the liveness detection model, and the first liveness probability value corresponding to the second depth image and the second liveness probability value corresponding to the target color image are output.

[0241] The liveness detection type of the target object is determined based on the first liveness probability value and the second liveness probability value.

[0242] Optionally, the liveness detection module 14 is used for:

[0243] The target liveness probability is determined based on the first liveness probability value and the second liveness probability value;

[0244] If the target liveness probability is greater than the target threshold, then the target object is determined to be a live object type;

[0245] If the target liveness probability is less than or equal to the target threshold, then the target object is determined to be an attack target type.

[0246] It should be noted that the image liveness detection device provided in the above embodiments is only illustrated by the division of the above functional modules when performing the image liveness detection method. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the image liveness detection device and the image liveness detection method embodiments provided in the above embodiments belong to the same concept, and the implementation process is detailed in the method embodiments, which will not be repeated here.

[0247] The serial numbers in this specification are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0248] In one or more embodiments of this specification, an electronic device obtains a first depth image by estimating the depth of multiple target color images based on a first depth model, and performs depth fusion with the inter-frame depth relationships between multiple first depth images focusing on the same object by mining and focusing on a second depth model. This results in a second depth image corresponding to a higher-precision depth estimate, thereby reducing the detection requirements for image accuracy and image quality during target color image acquisition. It can resist detection interference in complex application environments and achieve a higher-precision second depth image based on a color two-dimensional image with lower image accuracy or lower image quality. Thus, image liveness detection can be performed based on the higher-precision second depth image and the target color image, improving the detection effect of image liveness detection in complex environments and low-performance hardware environments, and enhancing the liveness detection effect and robustness.

[0249] This specification also provides a computer storage medium capable of storing multiple instructions adapted to be loaded and executed by a processor as described above. Figures 1-4 The image liveness detection method described in the illustrated embodiment can be found in the following document for a detailed execution process. Figures 1-4 The specific details of the illustrated embodiments will not be elaborated here.

[0250] This application also provides a computer program product storing at least one instruction, which is loaded and executed by the processor as described above. Figures 1-4 The image liveness detection method described in the illustrated embodiment can be found in the following document for a detailed execution process. Figures 1-4 The specific details of the illustrated embodiments will not be elaborated here.

[0251] Please refer to Figure 9This diagram illustrates a structural block diagram of an electronic device provided in an exemplary embodiment of this application. The electronic device in this application may include one or more components such as a processor 110, a memory 120, an input device 130, an output device 140, and a bus 150. The processor 110, memory 120, input device 130, and output device 140 may be connected via the bus 150.

[0252] Processor 110 may include one or more processing cores. Processor 110 connects to various parts of the electronic device via various interfaces and lines, and performs various functions and processes data of electronic device 100 by running or executing instructions, programs, code sets, or instruction sets stored in memory 120, and by calling data stored in memory 120. Optionally, processor 110 may be implemented using at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). Processor 110 may integrate one or more of the following: central processing unit (CPU), graphics processing unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the displayed content; and the modem handles wireless communication. It is understood that the modem may also not be integrated into processor 110 and may be implemented separately through a communication chip.

[0253] The memory 120 may include random access memory (RAM) or read-only memory (ROM). Optionally, the memory 120 may include a non-transitory computer-readable storage medium. The memory 120 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 120 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as touch functionality, sound playback functionality, image playback functionality, etc.), instructions for implementing the various method embodiments described below, etc. The operating system may be the Android system, including systems deeply developed based on the Android system, the iOS system developed by Apple Inc., including systems deeply developed based on the iOS system, or other systems. The data storage area may also store data created by the electronic device during use, such as phonebook data, audio and video data, chat log data, etc.

[0254] See Figure 10 As shown, the memory 120 can be divided into operating system space and user space. The operating system runs in the operating system space, while native and third-party applications run in the user space. To ensure that different third-party applications can achieve good running performance, the operating system allocates corresponding system resources for each application. However, different application scenarios within the same third-party application have different requirements for system resources. For example, in local resource loading scenarios, third-party applications have high requirements for disk read speed; in animation rendering scenarios, third-party applications have high requirements for GPU performance. Since the operating system and third-party applications are independent of each other, the operating system often cannot promptly perceive the current application scenario of a third-party application, resulting in the operating system's inability to adapt system resources accordingly to the specific application scenario of the third-party application.

[0255] In order for the operating system to distinguish the specific application scenarios of third-party applications, it is necessary to establish data communication between the third-party applications and the operating system. This would allow the operating system to obtain the current scenario information of the third-party applications at any time, and then perform targeted system resource adaptation based on the current scenario.

[0256] Taking the Android operating system as an example, the programs and data stored in memory 120 are as follows: Figure 11As shown, the memory 120 can store the Linux kernel layer 320, the system runtime library layer 340, the application framework layer 360, and the application layer 380. The Linux kernel layer 320, system runtime library layer 340, and application framework layer 360 belong to the operating system space, while the application layer 380 belongs to the user space. The Linux kernel layer 320 provides low-level drivers for various hardware components of the electronic device, such as display drivers, audio drivers, camera drivers, Bluetooth drivers, Wi-Fi drivers, and power management. The system runtime library layer 340 provides support for key features of the Android system through several C / C++ libraries. For example, the SQLite library provides database support, the OpenGL / ES library provides 3D graphics support, and the Webkit library provides browser kernel support. The system runtime library layer 340 also provides the Android runtime library, which mainly provides core libraries that allow developers to write Android applications using the Java language. The Application Framework Layer 360 provides various APIs that may be used when building applications. Developers can also use these APIs to build their own applications, such as activity management, window management, view management, notification management, content provider, package management, call management, resource management, and location management. At least one application runs in the Application Layer 380. These applications can be native applications that come with the operating system, such as contacts, SMS, clock, and camera apps; or third-party applications developed by third-party developers, such as games, instant messaging, and photo editing apps.

[0257] Taking the operating system as an example (iOS), the programs and data stored in memory 120 are as follows: Figure 9As shown, the iOS system includes: Core OS layer 420, Core Services layer 440, Media layer 460, and Cocoa Touch layer 480. Core OS layer 420 includes the operating system kernel, drivers, and low-level program frameworks. These low-level program frameworks provide hardware-level functionality for use by the program frameworks located in Core Services layer 440. Core Services layer 440 provides system services and / or program frameworks required by applications, such as Foundation framework, account framework, advertising framework, data storage framework, network connectivity framework, geolocation framework, motion framework, etc. Media layer 460 provides applications with audiovisual interfaces, such as interfaces related to graphics and images, audio technology, video technology, and AirPlay (wireless playback of audio and video transmission technologies). Cocoa Touch layer 480 provides various commonly used interface-related frameworks for application development and is responsible for user touch interaction on electronic devices. Examples include local notification services, remote push services, advertising frameworks, game tool frameworks, message user interface (UI) frameworks, UIKit frameworks, map frameworks, and so on.

[0258] exist Figure 12 The framework shown includes, but is not limited to, the base framework in the core service layer 440 and the UIKit framework in the touchable layer 480. The base framework provides many basic object classes and data types, offering the most basic system services to all applications, and is independent of the UI. The UIKit framework, on the other hand, provides a basic UI class library for creating touch-based user interfaces. iOS applications can use the UIKit framework to provide their UI, thus providing the application's infrastructure for building user interfaces, drawing, handling user interaction events, responding to gestures, and so on.

[0259] The methods and principles for implementing data communication between third-party applications and the operating system in the iOS system can be referenced from the Android system, and will not be elaborated here.

[0260] The input device 130 is used to receive input instructions or data, and includes, but is not limited to, a keyboard, mouse, camera, microphone, or touch device. The output device 140 is used to output instructions or data, and includes, but is not limited to, a display device and a speaker. In one example, the input device 130 and the output device 140 can be combined into a touch screen, which is used to receive touch operations from the user using a finger, stylus, or any suitable object on or near it, and to display the user interface of various applications. The touch screen is usually located on the front panel of the electronic device. The touch screen can be designed as a full-screen, curved screen, or irregularly shaped screen. The touch screen can also be designed as a combination of a full-screen and a curved screen, or a combination of an irregularly shaped screen and a curved screen; this specification does not limit this.

[0261] In addition, those skilled in the art will understand that the structure of the electronic device shown in the above figures does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements. For example, the electronic device may also include radio frequency circuits, input units, sensors, audio circuits, wireless fidelity (WiFi) modules, power supplies, Bluetooth modules, etc., which will not be described in detail here.

[0262] In this specification, the entity executing each step can be the electronic device described above. Optionally, the entity executing each step can be the operating system of the electronic device. The operating system can be Android, iOS, or other operating systems; this specification does not limit this.

[0263] The electronic device described in this manual may also be equipped with a display device. This display device can be any device capable of displaying information, such as a cathode ray tube display (CR), a light-emitting diode display (LED), an e-ink screen, a liquid crystal display (LCD), or a plasma display panel (PDP). Users can use the display device on electronic device 101 to view displayed text, images, videos, and other information. The electronic device may be a smartphone, tablet computer, gaming device, AR (Augmented Reality) device, automobile, data storage device, audio playback device, video playback device, laptop, desktop computing device, or wearable device such as an electronic watch, electronic glasses, electronic helmet, electronic bracelet, electronic necklace, or electronic clothing.

[0264] exist Figure 9 In the illustrated electronic device, which can be a terminal, the processor 110 can be used to call the application stored in the memory 120 and specifically perform the following operations:

[0265] Acquire at least two color images of the target object;

[0266] Based on the first depth model, depth estimation processing is performed on each of the target color images to obtain the first depth image corresponding to each frame of the target color image;

[0267] Based on the second depth model, inter-frame depth fusion processing is performed on each of the first depth images to obtain a second depth image for the target object;

[0268] Image liveness detection processing is performed on the target object based on the second depth image and the target color image.

[0269] In one embodiment, before executing the image liveness detection method, the processor 110 also performs the following operations:

[0270] Create the initial first depth model and the initial second depth model;

[0271] Acquire at least one set of image sample data, the image sample data including at least two consecutive sample images for the same sample object;

[0272] Based on the image sample data, the initial first depth model is trained for depth estimation, and the initial second depth model is trained for inter-frame depth fusion, to obtain the trained first depth model and the second depth model.

[0273] In one embodiment, when the processor 110 performs depth estimation training on an initial first depth model and inter-frame depth fusion training on an initial second depth model based on the image sample data to obtain the trained first depth model and second depth model, it specifically performs the following operations:

[0274] Each of the sample images in the image sample data is input into the initial first depth model, which outputs a sample depth estimation map corresponding to each sample image. The initial first depth model is then trained on the sample depth estimation map until the initial first depth model is trained, resulting in the trained first depth model.

[0275] The sample depth estimation map corresponding to each sample image is input into the sample depth fusion map output by the initial second depth model, and the initial second depth model is trained by inter-frame depth fusion based on the sample depth fusion map until the initial second depth model is trained to obtain the trained second depth model.

[0276] In one embodiment, the processor 110 performs the following steps when executing the process of inputting each sample image from the image sample data into an initial first depth model and outputting a sample depth estimation map corresponding to each sample image, and training the initial first depth model based on the sample depth estimation map:

[0277] Determine the label depth map corresponding to each of the sample images;

[0278] Each of the sample images is input into the initial first depth model for depth estimation processing, and the sample depth estimation map corresponding to each of the sample images is output.

[0279] The pixel estimation loss is determined based on the aforementioned and the label depth map, and the model parameters of the initial first depth map are adjusted based on the pixel estimation loss.

[0280] In one embodiment, the processor 110 performs the following steps when determining the pixel estimation loss based on the sample depth estimation map and the label depth map:

[0281] The sample depth estimation map and the label depth map are input into the first loss calculation formula to determine the pixel estimation loss;

[0282] The first loss calculation formula satisfies the following formula:

[0283]

[0284] Among them, the Loss A Estimate the loss for each pixel. p is the estimated depth value of the i-th depth pixel in the sample depth estimation map. i γ is the label depth value of the i-th depth pixel in the label depth map, γ is the loss adaptive parameter, i is an integer, and I is the total number of pixels in the sample depth estimation map.

[0285] In one embodiment, before performing the determination of the label depth map corresponding to each of the sample images, the processor 110 further includes:

[0286] Based on the image liveness detection task, each sample image is acquired, and the image sample type of the sample image is determined;

[0287] The label depth map corresponding to the sample image is set based on the image sample type.

[0288] In one embodiment, the image sample type includes attack image type and liveness image type. The processor 110, when executing the step of setting the label depth map corresponding to the sample image based on the image sample type, specifically performs the following steps:

[0289] Obtain the first sample image corresponding to the attack image type, and set the depth pixel value of the first label depth map corresponding to the first sample image as the target depth pixel value;

[0290] Obtain the second sample image corresponding to the liveness image type, and call the target image depth service to determine the second label depth map corresponding to the second sample image.

[0291] In one embodiment, the processor 110 performs the following steps when executing the process of inputting the sample depth estimation map corresponding to each sample image into the initial second depth model to output a sample depth fusion map, and performing inter-frame depth fusion training on the initial second depth model based on the sample depth fusion map:

[0292] The sample depth estimation map corresponding to each sample image is input into the initial second depth model for depth reconstruction processing to obtain the depth map estimation feature corresponding to the sample depth estimation map and the sample depth reconstruction map corresponding to the sample depth estimation map.

[0293] Based on the depth map estimation features corresponding to each of the sample depth estimation maps, inter-frame depth fusion is performed to output a sample depth fusion map, and the model parameters of the initial second depth model are adjusted based on the sample depth reconstruction map and the sample depth estimation map.

[0294] In one embodiment, the initial second depth model includes at least a first depth encoding network, a second depth decoding network, and a third self-attention network. The processor 110 performs the following steps when executing the process of inputting the sample depth estimation maps corresponding to each sample image into the initial second depth model for depth reconstruction, obtaining depth map estimation features corresponding to the sample depth estimation maps and sample depth reconstruction maps corresponding to the sample depth estimation maps, and then performing inter-frame depth fusion based on the depth map estimation features corresponding to each sample depth estimation map to output a sample depth fusion map:

[0295] The sample depth estimation maps corresponding to each sample image of the same sample object are respectively input into the first depth coding network to obtain the depth map estimation features corresponding to each sample depth estimation map;

[0296] The depth map estimation features are input into the second depth decoding network to obtain the sample depth reconstruction map corresponding to the sample depth estimation map;

[0297] The depth map estimation features corresponding to each of the sample depth estimation maps are input into the third self-attention network to obtain a reference depth map and an inter-frame relation matrix, wherein the reference depth map is one of the sample depth estimation maps;

[0298] Inter-frame depth fusion is performed based on the inter-frame relationship matrix and the reference depth map to output a sample depth fusion map.

[0299] In one embodiment, the processor 110 performs the following steps when executing the inter-frame depth fusion based on the inter-frame relation matrix and the reference depth map to output a sample depth fusion map:

[0300] Based on the inter-frame relationship matrix and the reference depth map, a weighted fusion is performed using the second inter-frame fusion calculation formula to obtain a sample depth fusion map.

[0301] The second inter-frame fusion calculation formula satisfies the following formula:

[0302]

[0303] Wherein, the depth multi The sample depth fusion map is defined as follows: N is the total number of pixels in the reference depth map, n is an integer, and the depth is... n w is the depth pixel value of the nth pixel in the original depth map. n,j In the figure, n represents the nth row of the inter-frame relation matrix, and w n,j The label j represents the j-th column of the inter-frame relation matrix, where w... n,j The coefficient represents the relationship between the j-th pixel in the fused depth map and the n-th pixel in the reference depth map. The fused depth map is the depth map other than the reference depth map among all the depth estimation maps of the same sample object.

[0304] In one embodiment, the processor 110 performs the following steps when adjusting the model parameters of the initial second depth model based on the sample depth reconstruction map and the sample depth estimation map:

[0305] The depth estimation map and the depth reconstruction map of each sample for the same sample object are input into the third loss calculation formula to determine the depth reconstruction loss;

[0306] The model parameters of the initial second depth model are adjusted based on the depth reconstruction loss.

[0307] The third loss calculation formula satisfies the following formula:

[0308]

[0309] Where Loss B is the depth reconstruction loss, l is an integer, L is the total number of depth estimation maps corresponding to the same sample object, and I pred-l For the l-th sample depth reconstruction map, the I GT-l This is the depth estimation map for the l-th sample.

[0310] In one embodiment, after the processor 110 performs inter-frame depth fusion processing on each of the first depth images based on the second depth model to obtain a second depth image for the target object, it further performs the following steps:

[0311] The second depth image is enhanced based on the third depth model to obtain the enhanced third depth image.

[0312] The image liveness detection processing of the target object based on the second depth image and the target color image includes:

[0313] Using the third depth image as the second depth image, perform image liveness detection processing on the target object based on the second depth image and the target color image.

[0314] In one embodiment, the processor 110, while executing the image liveness detection method, also performs the following steps:

[0315] Create an initial third-depth model;

[0316] Obtain at least one sample depth fusion map corresponding to the initial second depth model of the second depth model;

[0317] The initial third depth model is trained with enhanced quality based on the depth fusion maps of the samples to obtain the trained third depth model.

[0318] In one embodiment, the processor 110 performs the following steps to perform quality enhancement training on the initial third depth model based on the depth fusion maps of each of the samples, thereby obtaining the trained third depth model:

[0319] Obtain the depth map enhancement labels corresponding to the sample depth fusion map;

[0320] Pixel perturbation processing is performed on each of the sample depth fusion maps to obtain the sample depth fusion map after pixel perturbation;

[0321] Each of the sample depth fusion maps is input into the initial third depth model for quality enhancement processing, and the sample enhanced depth map corresponding to the sample depth fusion map is output. The model parameters of the initial third depth model are adjusted based on the sample enhanced depth map and the enhancement label of the depth map until the initial third depth model completes training, and the trained third depth model is obtained.

[0322] In one embodiment, the processor 110 performs the following steps when adjusting the model parameters of the initial third depth model based on the sample enhanced depth map and the depth map enhanced labels:

[0323] The sample enhanced depth map and the depth map enhanced label are input into the fourth loss calculation formula to determine the quality enhancement loss;

[0324] The model parameters of the initial third depth model are adjusted based on the aforementioned quality enhancement loss.

[0325] The fourth loss calculation formula satisfies the following formula:

[0326]

[0327] Wherein, Loss C is the mass enhancement loss, and I re For the sample enhancement depth map, the I GT Enhance the labels on the depth map.

[0328] In one embodiment, the processor 110 performs the following steps when executing the image liveness detection processing of the target object based on the second depth image and the target color image:

[0329] The second depth image and the target color image are input into the liveness detection model, and the first liveness probability value corresponding to the second depth image and the second liveness probability value corresponding to the target color image are output.

[0330] The liveness detection type of the target object is determined based on the first liveness probability value and the second liveness probability value.

[0331] In one embodiment, the processor 110, when performing the step of determining the liveness detection type of the target object based on the first liveness probability value and the second liveness probability value, specifically executes the following steps:

[0332] The target liveness probability is determined based on the first liveness probability value and the second liveness probability value;

[0333] If the target liveness probability is greater than the target threshold, then the target object is determined to be a live object type;

[0334] If the target liveness probability is less than or equal to the target threshold, then the target object is determined to be an attack target type.

[0335] In one or more embodiments of this specification, an electronic device obtains a first depth image by estimating the depth of multiple target color images based on a first depth model, and performs depth fusion with the inter-frame depth relationships between multiple first depth images focusing on the same object by mining and focusing on a second depth model. This results in a second depth image corresponding to a higher-precision depth estimate, thereby reducing the detection requirements for image accuracy and image quality during target color image acquisition. It can resist detection interference in complex application environments and achieve a higher-precision second depth image based on a color two-dimensional image with lower image accuracy or lower image quality. Thus, image liveness detection can be performed based on the higher-precision second depth image and the target color image, improving the detection effect of image liveness detection in complex environments and low-performance hardware environments, and enhancing the liveness detection effect and robustness.

[0336] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory, or random access memory, etc.

[0337] The above-disclosed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of this application. Therefore, any equivalent variations made in accordance with the claims of this application shall still fall within the scope of this application.

Claims

1. An image living body detection method, the method comprising: obtaining at least two target color images of a target object; performing depth estimation processing on each of the target color images based on a first depth model to obtain a first depth image corresponding to each of the target color images, the first depth image being a depth estimation image of the target color image; performing inter-frame depth fusion processing on each of the first depth images based on a second depth model to obtain a second depth image of the target object, the second depth image being a depth fusion image, the inter-frame depth fusion processing comprising: obtaining, based on depth estimation image features of the same object obtained by a third self-attention network of the second depth model, a reference depth image and an inter-frame relationship matrix for representing relationships between the depth estimation images, and performing fusion based on the inter-frame relationship matrix and the reference depth image to obtain the depth fusion image, wherein each of the depth estimation image features is obtained by a first depth encoding network of the second depth model based on the depth estimation images of the same object, and the reference depth image is one of the depth estimation images; performing image living body detection processing on the target object based on the second depth image and the target color images.

2. The method of claim 1, further comprising: creating an initial first depth model and an initial second depth model; obtaining at least one set of image sample data, the image sample data comprising at least two consecutive sample images of the same sample object; performing depth estimation training on the initial first depth model and inter-frame depth fusion training on the initial second depth model based on the image sample data to obtain a trained first depth model and a trained second depth model.

3. The method of claim 2, wherein the performing depth estimation training on the initial first depth model and inter-frame depth fusion training on the initial second depth model based on the image sample data to obtain a trained first depth model and a trained second depth model comprises: inputting each of the sample images in the image sample data into the initial first depth model to output a sample depth estimation image corresponding to each of the sample images, and performing depth estimation training on the initial first depth model based on the sample depth estimation images until the initial first depth model is trained to obtain the trained first depth model; inputting the sample depth estimation image corresponding to each of the sample images into the initial second depth model to output a sample depth fusion image, and performing inter-frame depth fusion training on the initial second depth model based on the sample depth fusion image until the initial second depth model is trained to obtain the trained second depth model.

4. The method of claim 3, wherein the inputting each of the sample images in the image sample data into the initial first depth model to output a sample depth estimation image corresponding to each of the sample images, and performing depth estimation training on the initial first depth model based on the sample depth estimation images comprises: determining a label depth image corresponding to each of the sample images. inputting each of the sample images into an initial first depth model for depth estimation processing, and outputting a sample depth estimation image corresponding to each of the sample images; determining a pixel estimation loss based on the sample depth estimation image and the label depth image, and adjusting model parameters of the initial first depth model based on the pixel estimation loss.

5. The method of claim 4, wherein the determining a pixel estimation loss based on the sample depth estimation image and the label depth image comprises: inputting the sample depth estimation image and the label depth image into a first loss calculation formula to determine the pixel estimation loss; the first loss calculation formula satisfies the following formula: wherein the Loss A is a pixel estimation loss, is an estimated depth value of an i-th depth pixel in the sample depth estimation map, the pi is a label depth value of the i-th depth pixel in the label depth map, the γ is a loss adaptive parameter, the i is an integer, and the I is a total number of pixels of the sample depth estimation map.

6. The method of claim 4, wherein before the determining a label depth image corresponding to each of the sample images, the method further comprises: collecting each sample image based on an image living body detection task, and determining an image sample type of the sample image; setting the label depth image corresponding to the sample image based on the image sample type.

7. The method of claim 6, wherein the image sample type comprises an attack image type and a living body image type, the setting the label depth image corresponding to the sample image based on the image sample type comprises: obtaining a first sample image corresponding to the attack image type, and setting a depth pixel value of a first label depth image corresponding to the first sample image as a target depth pixel value; obtaining a second sample image corresponding to the living body image type, and calling a target image depth service to determine a second label depth image corresponding to the second sample image.

8. The method of claim 4, wherein the inputting the sample depth estimation image corresponding to each of the sample images into an initial second depth model to output a sample depth fusion image, and performing inter-frame depth fusion training on the initial second depth model based on the sample depth fusion image comprises: inputting the sample depth estimation image corresponding to each of the sample images into an initial second depth model for depth reconstruction processing to obtain a depth image estimation feature corresponding to the sample depth estimation image and a sample depth reconstruction image corresponding to the sample depth estimation image; performing inter-frame depth fusion based on the depth image estimation feature corresponding to each of the sample depth estimation images to output a sample depth fusion image, and adjusting model parameters of the initial second depth model based on the sample depth reconstruction image and the sample depth estimation image.

9. The method of claim 8, wherein the initial second depth model at least comprises a first depth encoding network, a second depth decoding network, and a third self-attention network, the inputting the sample depth estimation image corresponding to each of the sample images into an initial second depth model for depth reconstruction processing to obtain a depth image estimation feature corresponding to the sample depth estimation image and a sample depth reconstruction image corresponding to the sample depth estimation image, and performing inter-frame depth fusion based on the depth image estimation feature corresponding to each of the sample depth estimation images to output a sample depth fusion image comprises: inputting the sample depth estimation image corresponding to each of the sample images of the same sample object into the first depth encoding network respectively to obtain a depth image estimation feature corresponding to each of the sample depth estimation images; inputting the depth map estimation feature into the second depth decoding network to obtain a sample depth reconstruction map corresponding to the sample depth estimation map; inputting the depth map estimation feature corresponding to each of the sample depth estimation maps into a third self-attention network to obtain a reference depth map and an inter-frame relationship matrix, the reference depth map being one of the sample depth estimation maps; performing inter-frame depth fusion based on the inter-frame relationship matrix and the reference depth map to output a sample depth fusion map.

10. The method of claim 9, wherein the performing inter-frame depth fusion based on the inter-frame relationship matrix and the reference depth map to output a sample depth fusion map comprises: performing weighted fusion using a second inter-frame fusion calculation formula based on the inter-frame relationship matrix and the reference depth map to obtain a sample depth fusion map; the second inter-frame fusion calculation formula satisfies the following formula: wherein the is the sample depth fusion graph, the N is the total number of pixel points of the reference depth graph, the n is an integer, the depth n is the depth pixel value of the n th pixel point of the original depth graph, the w n,j the subscript n in the above formula is the n th row of the inter-frame relationship matrix, the w n,j the subscript j in the above formula represents the j th column of the inter-frame relationship matrix, and the w n,j represents the relationship coefficient of the j th pixel point of the fusion depth graph relative to the n th pixel point of the reference depth graph, and the fusion depth graph is a depth graph other than the reference depth graph in all sample depth estimation graphs corresponding to the same sample object.

11. The method of claim 8, wherein the adjusting model parameters of the initial second depth model based on the sample depth reconstruction map and the sample depth estimation map comprises: inputting each of the sample depth estimation maps and the sample depth reconstruction map for the same sample object into a third loss calculation formula to determine a depth reconstruction loss; adjusting model parameters of the initial second depth model based on the depth reconstruction loss; the third loss calculation formula satisfies the following formula: wherein, is the depth reconstruction loss, the is an integer, the L is a total number of the same sample object corresponding to the sample depth estimation map, the is the l-th sample depth reconstruction map, the is the l-th sample depth estimation map.

12. The method of claim 1, wherein after the performing inter-frame depth fusion processing on each of the first depth images based on a second depth model to obtain a second depth image for the target object, the method further comprises: performing quality enhancement processing on the second depth image based on a third depth model to obtain a third depth image after quality enhancement processing; the performing image living body detection processing on the target object based on the second depth image and the target color image comprises: performing the image living body detection processing on the target object based on the second depth image and the target color image using the third depth image as the second depth image.

13. The method of claim 12, further comprising: creating an initial third depth model; obtaining at least one sample depth fusion map corresponding to the initial second depth model of the second depth model; performing quality enhancement training on the initial third depth model based on each of the sample depth fusion maps to obtain a trained third depth model.

14. The method of claim 13, wherein the performing quality enhancement training on the initial third depth model based on each of the sample depth fusion maps to obtain a trained third depth model comprises: obtaining a depth map enhancement label corresponding to the sample depth fusion map; performing pixel disturbance processing on each of the sample depth fusion maps to obtain the sample depth fusion map after pixel disturbance processing; The sample depth fusion image is input into an initial third depth model for quality enhancement processing, and a sample enhanced depth image corresponding to the sample depth fusion image is output. The initial third depth model is adjusted in model parameters based on the sample enhanced depth image and the depth image enhanced label until the initial third depth model is trained to obtain a trained third depth model.

15. The method of claim 14, wherein the adjusting the initial third depth model in model parameters based on the sample enhanced depth image and the depth image enhanced label comprises: inputting the sample enhanced depth image and the depth image enhanced label into a fourth loss calculation formula to determine a quality enhancement loss; adjusting the initial third depth model in model parameters based on the quality enhancement loss; and the fourth loss calculation formula satisfies the following formula: wherein the is the quality augmentation loss, the I re is the sample augmented depth map, the is the depth map augmentation label.

16. The method of any one of claims 1-15, wherein the image living body detection processing of the target object based on the second depth image and the target color image comprises: inputting the second depth image and the target color image into a living body detection model to output a first living body probability value corresponding to the second depth image and a second living body probability value corresponding to the target color image; and determining a living body detection type of the target object based on the first living body probability value and the second living body probability value.

17. The method of claim 16, wherein the determining the living body detection type of the target object based on the first living body probability value and the second living body probability value comprises: determining a target living body probability based on the first living body probability value and the second living body probability value; determining that the target object is a living body object type if the target living body probability is greater than a target threshold value; and determining that the target object is an attack object type if the target living body probability is less than or equal to the target threshold value.

18. An image living body detection device, the device comprising: an image acquisition module configured to acquire at least two frames of target color images for a target object based on an image living body detection task; a depth estimation module configured to perform depth estimation processing on each of the target color images based on a first depth model to obtain a first depth image corresponding to each of the target color images, the first depth image being a depth estimation image of a target color image; a depth fusion module configured to perform inter-frame depth fusion processing on each of the first depth images based on a second depth model to obtain a second depth image for the target object, the second depth image being a depth fusion image, the inter-frame depth fusion processing comprising: obtaining, based on depth estimation image features of each of the depth estimation images, a reference depth image and an inter-frame relationship matrix for representing relationships between the depth estimation images by a third self-attention network of the second depth model, and performing fusion based on the inter-frame relationship matrix and the reference depth image to obtain the depth fusion image, wherein each of the depth estimation image features is obtained based on each of the depth estimation images of the same object by a first depth encoding network of the second depth model, and the reference depth image is one of the depth estimation images; and an image living body detection module configured to perform image living body detection processing on the target object based on the second depth image and the target color image. The living body detection module is configured to perform image living body detection processing on the target object based on the second depth image and the target color image.

19. A computer storage medium storing a plurality of instructions adapted to be loaded and executed by a processor to perform the method steps of any one of claims 1-17.

20. A computer program product storing at least one instruction adapted to be loaded and executed by a processor to perform the method steps of any one of claims 1-17.

21. An electronic device, comprising: A processor and a memory; wherein the memory stores a computer program adapted to be loaded and executed by the processor to perform the method steps of any one of claims 1-17.

Citation Information

Patent Citations

  • Face-in-vivo detection method, device, equipment and storage medium

    CN109034102A

  • Living body recognition method and device, electronic equipment and readable storage medium

    CN112464690A

  • Matting-based sample enhancement method, training method, system thereof and electronic equipment

    CN113012054A