Living body detection method and device and computer readable storage medium

By acquiring and analyzing facial RGB images and depth temporal images, and using facial recognition and classification models for liveness detection, the problem of low accuracy in existing technologies is solved, achieving higher liveness detection accuracy and anti-attack capabilities.

CN115311723BActive Publication Date: 2026-03-27PING AN TECH (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-16
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Current technologies for facial liveness detection have low accuracy and cannot effectively resist attacks from cyber hackers.

Method used

By acquiring RGB images of the face with blinking motion and temporal images of the face depth, facial recognition and classification models are used for analysis. The pass rate of liveness detection is determined by combining a preset threshold. A weighted average is performed using a two-fluid model to determine whether the obtained liveness detection probability is greater than the preset threshold. If it is greater, the liveness detection is determined to be successful.

Benefits of technology

It improves the accuracy of liveness detection and enhances the ability to resist external attacks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115311723B_ABST
    Figure CN115311723B_ABST
Patent Text Reader

Abstract

The application provides a living body detection method and device and a computer readable storage medium. The method comprises the following steps: obtaining face image information with a blinking action; wherein the face image information with the blinking action comprises a face RGB image and a face depth time sequence image; obtaining a first living body detection probability based on a face recognition model and a classification model according to the face RGB image and the face depth time sequence image; judging whether the first living body detection probability is greater than a preset threshold; and if the first living body detection probability is greater than the preset threshold, determining that the living body detection is passed. Through the processing of the face recognition model and the classification model on the face image with the blinking action, the living body detection probability with high accuracy is obtained, the accuracy of the living body detection is improved, and the ability of resisting external attacks is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, in particular to a living body detection method and device and a computer readable storage medium. BACKGROUND

[0002] With the continuous development of computer technology, the application scenarios of face living body detection are becoming more and more rich, for example: the face living body detection function can be used in attendance software, payment software, social software and the like, but network hackers use fake face or video splicing to complete online face recognition, which can cause very serious loss to users. Therefore, the security and credibility of the face recognition process in various scenarios become a crucial problem.

[0003] In the prior art, there are many living body detection products: action biopsy, silent biopsy, face light biopsy and the like, although they can solve the security in the process of face living body detection, but these products all have obvious shortcomings, the living body detection accuracy is low, which leads to being insufficient to resist the attack of network hackers. Therefore, it is still necessary to further improve the accuracy of living body detection. SUMMARY

[0004] The purpose of the embodiment of the present application is a living body detection method, device and computer readable storage medium, which processes the feature information of the face RGB image and the depth time sequence image through the face recognition model, the preset classification model and the double fluid model to obtain an accurate living body detection probability, so as to solve the problem of low accuracy of living body detection in the prior art.

[0005] In a first aspect, the embodiment of the present application provides a living body detection method, which comprises:

[0006] Obtaining face image information with blinking action; wherein the face image information with blinking action comprises a face RGB image and a face depth time sequence image;

[0007] According to the face RGB image and the face depth time sequence image, a first living body detection probability is obtained based on a face recognition model and a classification model;

[0008] Judging whether the first living body detection probability is greater than a preset threshold value;

[0009] If the first living body detection probability is greater than the preset threshold value, it is determined that the living body detection is passed.

[0010] In the implementation process, the face RGB image with the blinking action and the face depth time sequence image are acquired, the face RGB image and the face depth time sequence image are recognized based on a face recognition model and a classification model, a living body detection probability is obtained, and then whether the current user passes the living body detection is determined according to a pre-set living body detection probability threshold. Since the face RGB image with the blinking action and the face depth time sequence image are used simultaneously, the confusion of the living body detection by the to-be-detected person using only the image with two-dimensional features is effectively avoided, the accuracy of the living body detection is improved, and the anti-external attack capability is improved.

[0011] Optionally, the classification model comprises a preset classification model and a double-fluid model.

[0012] The first living body detection probability is obtained based on the face recognition model and the classification model according to the face RGB image and the face depth time sequence image.

[0013] The face recognition model is used to analyze the face RGB image and the face depth time sequence image, and a plurality of face images and a plurality of face depth images are obtained.

[0014] The face recognition model is used to analyze the plurality of face images, and a plurality of eye region images are obtained.

[0015] The plurality of eye region images are respectively mapped into the corresponding plurality of face depth images, and a plurality of eye region depth images are obtained.

[0016] In the implementation process, the face recognition model is used to analyze the face RGB image and the face depth image, a plurality of face images and a plurality of face depth images are obtained, the face recognition model is used to analyze the face image, a plurality of eye region images are obtained, and finally the plurality of eye region images are mapped into the plurality of face depth images, and a plurality of eye region depth images are obtained. Since the living body detection probability of the application is obtained based on the living body probability obtained by comprehensively combining the face image, the eye region image and the eye region depth image, the face RGB image, the eye region image and the eye region depth image become indispensable conditions for improving the accuracy of the living body detection probability.

[0017] Optionally, after the step of mapping the plurality of eye region images into the corresponding plurality of face depth images to obtain the plurality of eye region depth images, the method further comprises:

[0018] The plurality of face images are normalized to obtain a plurality of normalized face images.

[0019] The plurality of normalized face images are input into the preset classification model to obtain a plurality of living body detection probabilities of the face images output by the preset classification model.

[0020] The live body detection probabilities of the plurality of face graphs are weighted and averaged to obtain a second live body detection probability.

[0021] In the implementation process, the plurality of face graphs are normalized, the normalized face graphs are input into a preset classification model, live body detection probabilities of the plurality of face graphs are output, the live body detection probabilities of the plurality of different face graphs are weighted and averaged to obtain a live body detection probability based on the face graph, and the live body detection probability based on the face graph is combined with a live body detection probability based on an eye region graph and an eye region depth graph to obtain the live body detection probability. The live body detection probability based on the face graph provides an effective basis for improving the accuracy of the live body detection probability.

[0022] Optionally, after the step of weighting and averaging the live body detection probabilities of the plurality of face graphs to obtain a second live body detection probability, the method further includes:

[0023] The plurality of eye region graphs and the plurality of eye region depth graphs are normalized to obtain a plurality of normalized eye region graphs and a plurality of normalized eye region depth graphs;

[0024] The plurality of normalized eye region graphs and the plurality of normalized eye region depth graphs are input into the double-fluid model to obtain live body detection probabilities of the plurality of eye region graphs and the plurality of eye region depth graphs output by the double-fluid model;

[0025] The plurality of live body detection probabilities of the plurality of eye region graphs and the plurality of eye region depth graphs are weighted and averaged to obtain a third live body detection probability.

[0026] In the implementation process, the plurality of eye region graphs and the plurality of eye region depth graphs are input into the double-fluid model, live body detection probabilities based on the eye region graphs and the eye region depth graphs are obtained, the live body detection probabilities of the eye region graphs and the corresponding eye region depth graphs of different frames are weighted and averaged to obtain a plurality of sets of live body detection probabilities of the eye region graphs and the eye region depth graphs, and the live body detection probabilities of the plurality of sets of eye region graphs and eye region depth graphs are weighted and averaged to obtain a live body detection probability based on the eye region graphs and the eye region depth graphs. The live body detection probability based on the eye region graphs and the eye region depth graphs is combined with a live body detection probability based on a face graph to obtain the live body detection probability. The live body detection probability based on the eye region graphs and the eye region depth graphs provides an effective basis for improving the accuracy of the live body detection probability.

[0027] Optionally, the dual-stream model comprises a first stream structure and a second stream structure; the plurality of normalized eye region images and the plurality of normalized eye region depth images are input into the dual-stream model to obtain a plurality of living body detection probabilities of the eye region images and the eye region depth images output by the dual-stream model, comprising:

[0028] the plurality of eye region images are input into the first stream structure to obtain a plurality of living body detection probabilities of the eye region images output by the first stream structure;

[0029] the plurality of eye region depth images are input into the second stream structure to obtain a plurality of living body detection probabilities of the eye region depth images output by the second stream structure;

[0030] the plurality of living body detection probabilities of the eye region images and the eye region depth images are weighted and averaged to obtain the third living body detection probability.

[0031] In the above implementation process, the first stream structure and the second stream structure in the dual-stream model respectively identify the plurality of eye region images and the plurality of eye region depth images, then the living body detection probability of each frame of the eye region image output by the first stream structure is weighted and averaged with the living body detection probability of the corresponding eye region depth image output by the second stream structure to obtain the living body detection probability of each frame of the eye region image fused with the corresponding eye region depth image, and finally the living body detection probabilities of the plurality of eye region images fused with the corresponding eye region depth images are weighted and averaged to obtain the living body detection probability based on the eye region image and the eye region depth image. Since the first stream structure and the second stream structure in the dual-stream model can identify and fuse the eye region image and the eye region depth image, the living body detection probability of the face image of the to-be-detected person can be effectively determined from different dimensions.

[0032] Optionally, after the step of weighting and averaging the plurality of living body detection probabilities of the eye region images and the eye region depth images to obtain the third living body detection probability, the method further comprises:

[0033] the second living body detection probability and the third living body detection probability are weighted and averaged to obtain the first living body detection probability.

[0034] In the implementation process, the final live body detection probability is obtained by weighted average of the live body detection probability based on the face graph and the live body detection probability based on the eye region graph and the eye region depth graph, so that the live body detection probability can be comprehensively analyzed from the three aspects of the face graph, the eye region graph and the eye region depth graph, thereby improving the accuracy of the live body detection probability.

[0035] Optionally, the preset classification model is a customized neural network model; the customized neural network model comprises an SE Block module, an Adam algorithm and a cosine annealing algorithm; the SE Block module is used for identifying subtle features; and the Adam algorithm and the cosine annealing algorithm are used for optimizing internal parameter values.

[0036] In the implementation process, the algorithm identifies subtle features by adding the SE Block module to the customized neural network model, and the algorithm can iteratively update weights and accelerate the convergence speed by adding the Adam algorithm and the cosine annealing algorithm to the lightweight neural network, and then the preset classification model is obtained. Since the preset classification model comprises the SE Block module and the Adam algorithm and the cosine annealing algorithm, the features of the face graph can be better identified, thereby improving the identification effect of the face graph and further improving the accuracy of the live body detection probability based on the face graph.

[0037] Optionally, if the first live body detection probability is less than or equal to a preset threshold, it is determined that the live body detection fails.

[0038] In the implementation process, whether the current user passes the live body detection is determined according to the preset live body detection probability threshold. If the detection probability does not exceed the preset live body detection probability, it is determined that the current user does not pass the live body detection, so that the external attack is prevented, and the loss to the user is avoided.

[0039] In a second aspect, the embodiments of the present application further provide a live body detection device, comprising:

[0040] The acquisition module is configured to acquire face image information with a blinking action, wherein the face image information with the blinking action comprises a face RGB image and a face depth time sequence image.

[0041] The detection module is configured to obtain a first live body detection probability based on a face recognition model and a classification model according to the face RGB image and the face depth time sequence image.

[0042] The judgment module is configured to determine whether the first live body detection probability is greater than a preset threshold.

[0043] determining that the living body detection passes if the first living body detection probability is greater than the preset threshold.

[0044] The living body detection device provided by the above embodiments has the same beneficial effects as the living body detection method provided by the first aspect or any one of the optional implementation manners of the first aspect, which will not be repeated here.

[0045] In a third aspect, the embodiments of the present application further provide a computer readable storage medium, which stores a computer program. When the computer program is run by a processor, the above-described method is executed.

[0046] The storage medium provided by the above embodiments has the same beneficial effects as the living body detection method provided by the first aspect or any one of the optional implementation manners of the first aspect, which will not be repeated here.

[0047] To sum up, the present application provides a living body detection method. The method comprises the following steps: obtaining face image information with blinking action; wherein the face image information with blinking action comprises a face RGB image and a face depth time sequence image; obtaining a first living body detection probability based on a face recognition model and a classification model according to the face RGB image and the face depth time sequence image; determining whether the first living body detection probability is greater than a preset threshold; and determining that the living body detection passes if the first living body detection probability is greater than the preset threshold. Through the processing of the face recognition model and the classification model on the face image with blinking action, a living body detection probability with high accuracy is obtained, the accuracy of living body detection is improved, and the ability to resist external attacks is improved. BRIEF DESCRIPTION OF DRAWINGS

[0048] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments of the present application. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can also be obtained without creative labor on the basis of these drawings.

[0049] Figure 1 A structural block diagram of an electronic device provided by the embodiments of the present application is provided.

[0050] Figure 2 A flowchart of a living body detection method provided by the embodiments of the present application is provided.

[0051] Figure 3 A position diagram of key points of a face recognition model in a face RGB image provided by the embodiments of the present application is provided.

[0052] Figure 4A structural schematic diagram of a customized neural network model provided by an embodiment of the present application is shown in FIG. 1.

[0053] Figure 5 A structural schematic diagram of a double-fluid model provided by an embodiment of the present application is shown in FIG. 2.

[0054] Figure 6 A functional module schematic diagram of a living body detection device provided by an embodiment of the present application is shown in FIG. 3. DETAILED DESCRIPTION

[0055] The embodiments of the technical solutions of the present application will be described in detail below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present application, and therefore only serve as examples, and cannot limit the protection scope of the present application.

[0056] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs; the terms used herein are only for the purpose of describing specific embodiments of the present application, and are not intended to limit the present application.

[0057] In the description of the embodiments of the present application, the technical terms "first", "second", etc. are only used to distinguish different objects, and cannot be understood as indicating or implying relative importance or implicitly indicating the number, specific order or primary and secondary relationship of the indicated technical features. In the description of the embodiments of the present application, the meaning of "a plurality of" is two or more, unless otherwise explicitly specified.

[0058] In order to facilitate the understanding of the present embodiment, first, the electronic device for executing the living body detection method disclosed by the present application is described in detail.

[0059] As shown in FIG. 4, it is a block schematic diagram of an electronic device. The electronic device 100 can include a memory 111, a storage controller 112, a processor 113, a peripheral interface 114, an input / output unit 115, and a display unit 116. Those skilled in the art can understand that the structure shown in FIG. 4 is only a schematic, and does not limit the structure of the electronic device 100. For example, the electronic device 100 can include more or fewer components than those shown in FIG. 4, or have a different configuration from that shown in FIG. 4. Figure 1 Figure 1 As shown in FIG. 4, it is a block schematic diagram of an electronic device. The electronic device 100 can include a memory 111, a storage controller 112, a processor 113, a peripheral interface 114, an input / output unit 115, and a display unit 116. Those skilled in the art can understand that the structure shown in FIG. 4 is only a schematic, and does not limit the structure of the electronic device 100. For example, the electronic device 100 can include more or fewer components than those shown in FIG. 4, or have a different configuration from that shown in FIG. 4. Figure 1 Figure 1 As shown in FIG. 4, it is a block schematic diagram of an electronic device. The electronic device 100 can include a memory 111, a storage controller 112, a processor 113, a peripheral interface 114, an input / output unit 115, and a display unit 116. Those skilled in the art can understand that the structure shown in FIG. 4 is only a schematic, and does not limit the structure of the electronic device 100. For example, the electronic device 100 can include more or fewer components than those shown in FIG. 4, or have a different configuration from that shown in FIG. 4.

[0060] ​​The above-mentioned memory 111, storage controller 112, processor 113, peripheral interface 114, input / output unit 115 and display unit 116 are electrically connected to each other directly or indirectly to realize data transmission or interaction. For example, these elements can be electrically connected to each other through one or more communication buses or signal lines. The above-mentioned processor 113 is used to execute the executable modules stored in the memory.

[0061] The memory 111 can be, but is not limited to, a random access memory (RAM), a read only memory (ROM), a programmable read only memory (PROM), an erasable programmable read only memory (EPROM), an electrically erasable programmable read only memory (EEPROM), etc. The memory 111 is used to store programs, and the processor 113 executes the programs after receiving an execution instruction. The method executed by the electronic device 100 defined by the processes disclosed in any of the embodiments of the present application can be applied in the processor 113 or implemented by the processor 113.

[0062] The processor 113 can be an integrated circuit chip having a signal processing capability. The processor 113 can be a general purpose processor, including a central processing unit (CPU), a network processor (NP), etc. It can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the disclosed methods, steps and logic block diagrams in the embodiments of the present application. The general purpose processor can be a microprocessor or the processor can also be any conventional processor.

[0063] The peripheral interface 114 couples various input / output devices to the processor 113 and the memory 111. In some embodiments, the peripheral interface 114, the processor 113 and the storage controller 112 can be implemented in a single chip. In other examples, they can be implemented by independent chips respectively.

[0064] The input / output unit 115 is configured to provide the user with an input interface for interacting with the electronic device 100. The input / output unit 115 can be, but is not limited to, a mouse, a keyboard, and the like.

[0065] The display unit 116 provides an interactive interface (e.g., a user operation interface) between the electronic device 100 and the user, or is configured to display image data for the user to refer to. In this embodiment, the display unit can be a liquid crystal display or a touch display. If it is a touch display, it can be a capacitive touch screen or a resistive touch screen that supports single-point and multi-point touch operations. Supporting single-point and multi-point touch operations means that the touch display can sense a touch operation generated simultaneously at one or more positions on the touch display, and transmit the sensed touch operation to the processor for calculation and processing.

[0066] The electronic device 100 in this embodiment can be configured to execute each step in each method provided in the embodiments of the present application. The implementation process of the live body detection method will be described in detail in the following embodiments.

[0067] It should be noted that, in the prior art, the face of the user is imaged by the camera to obtain a face RGB image, and then the face RGB image is analyzed by a face recognition technology to obtain an analysis result. However, the above method cannot distinguish a synthetic face video from a real face video, so the face recognition technology in the prior art cannot meet the demand of ensuring user information security at present. In the embodiments of the present application, it is found that there is a difference between the depth time sequence diagram of a normal user's blinking action and the depth time sequence diagram of a synthetic blinking action, and the difference is that the depth time sequence diagram of the synthetic blinking action only shows a patch, while the depth time sequence diagram of the normal user's blinking action clearly shows the change of the blinking action, that is, this distinguishing feature is an effective means for distinguishing the depth time sequence diagrams in two different environments. In addition, in order to better obtain the time sequence information and considering that the synthetic blinking action cannot cause too much impact on the user and improve the user experience, the synthetic blinking action is more acceptable to the user. Therefore, the embodiments of the present application select to analyze and recognize the face RGB image and the face depth time sequence diagram with blinking action to perform live body detection.

[0068] In addition, the analysis and recognition of the face RGB image in the prior art is only two-dimensional recognition, while the analysis and recognition of the face RGB image and the depth time sequence with blinking action in the embodiments of the present application is three-dimensional recognition. Therefore, the accuracy of the live body detection method provided in the embodiments of the present application is higher than that of the live body detection method in the prior art.

[0069] Please refer to Figure 2A flowchart of a living body detection method provided by an embodiment of the application is shown.

[0070] In step S200, facial image information with a blinking action is acquired; the facial image information with the blinking action includes a facial RGB image and a facial depth time-series image.

[0071] The subject of the embodiment of the application is an electronic device 100, which can be provided with a camera and an electronic device with Lidar (Light Detection and Ranging) technology, or the camera and the electronic device with Lidar technology exist independently of the electronic device 100 but need to be connected in communication with the electronic device 100.

[0072] In one embodiment, after the electronic device 100 acquires the facial image collected by the camera and the electronic device with Lidar technology, the electronic device 100 identifies the facial image to obtain the facial RGB image with the blinking action and the facial depth time-series image.

[0073] In one embodiment, the facial image information with the blinking action includes the facial RGB image with the blinking action and the facial depth time-series image with the blinking action, which means that the facial RGB image collected by the camera and the facial depth time-series image collected by the electronic device with Lidar technology have the feature of the blinking action, that is, if the camera and the electronic device with Lidar technology collect image information at the same time, only the facial RGB image with the blinking action and the facial depth time-series image with the blinking action are sent to the electronic device 100. In addition, the facial image can be a real facial image of a user to be detected or a synthetic facial image.

[0074] Specifically, when the electronic device 100 needs to perform living body detection on facial image information, the facial image information can be real facial image information of a user to be detected or facial image information replaced or artificially synthesized by a network hacker (i.e., an external attacker), the camera and the electronic device with Lidar technology simultaneously collect and process the facial image information at the same time to obtain the facial RGB image and the facial depth time-series image, and then the electronic device 100 acquires the facial RGB image and the facial depth time-series image collected by the camera and the electronic device with Lidar technology.

[0075] In step S400, a first living body detection probability is obtained based on a facial recognition model and a classification model according to the facial RGB image and the facial depth time-series image.

[0076] Please refer to Figure 3 A schematic diagram of the positions of key points of a facial recognition model provided by an embodiment of the application in a facial RGB image is shown.

[0077] As shown in Figure 3 When the face recognition model recognizes the face RGB image, it adds multiple landmark key points (i.e., face feature points) on the face RGB image. The multiple face feature points can include face feature points for representing a face contour, face feature points for representing a right eyebrow, face feature points for representing a left eyebrow, face feature points for representing a nose, face feature points for representing a left eye, face feature points for representing a right eye, face feature points for representing an upper lip, and face feature points for representing a lower lip.

[0078] The face recognition model includes, but is not limited to, mediapipe, dlib, ssd-face, cengterface, DBFace, and the like. Compared with other face recognition models, mediapipe has a better detection effect, which is specifically embodied in that: on the one hand, mediapipe has 68 landmark points, which can more accurately capture information of a face image; on the other hand, mediapipe does not need an additional landmark algorithm, so it has a better processing efficiency. Therefore, the face recognition model can be selected according to actual application requirements, and the present application embodiment does not make specific limitations here.

[0079] The classification model includes a preset classification model and a double-fluid model. The preset classification model is a model improved on the basis of a lightweight neural network model (for example, mobileNetV3).

[0080] Specifically, after the electronic device 100 collects the face RGB image obtained by the camera and the face depth time sequence image obtained by the electronic device with Lidar technology, the face recognition model and the classification model are used to recognize and analyze the face RGB image and the face depth time sequence image, so as to obtain the live detection probability of the face image.

[0081] In one embodiment, step S400 can specifically include steps S410-S430.

[0082] In step S410, the face recognition model is used to analyze the face RGB image and the face depth time sequence image to obtain multiple frames of face images and multiple frames of face depth images.

[0083] Frame: The single image frame that affects the smallest unit in animation. A frame is a still image, and the continuous frames form animation, such as television images. Usually, the frame rate is simply the number of frames transmitted in 1 second, which can also be understood as the number of times the graphics processor can refresh per second, usually expressed in FPS (Frames Per Second). Each frame is a still image, and the rapid display of frames creates the illusion of movement. High frame rate can get smoother and more realistic animation, and the larger the FPS, the smoother the motion displayed.

[0084] It should be noted that since the face RGB image and the face depth time sequence image include not only the face region but also the region outside the face region, in order to prevent interference from the region outside the face region, the embodiment of the present application extracts only the video frames with faces when performing video frame extraction on the face RGB image and the face depth time sequence image, that is, the face image and the face depth image.

[0085] The duration of the face RGB image and the face depth time sequence image is related to the FPS value of the face RGB image and the face depth time sequence image, the number of face RGB images and face depth time sequence images containing faces, and the number of face images to be extracted, so the duration of the face RGB image and the face depth time sequence image can be set according to specific conditions and requirements, and the embodiment of the present application does not make specific limitations here. For example, if the FPS value is large, the duration of the face RGB image and the face depth time sequence image can be short, for example, 1s, 2s, etc., and if the FPS value is small, the duration of the face RGB image and the face depth time sequence image can be long, for example, 10s, 20s, etc.

[0086] Specifically, after inputting the obtained face RGB image and face depth time sequence image into the face recognition model, the face recognition model judges whether the current face image contains a face by recognizing the landmark key points of each frame of the face image in the face RGB image and the face depth time sequence image. If the face recognition model identifies that the current frame of the face image and the face depth time sequence image contains a face, the current frame of the face image and the face depth time sequence image are extracted, and finally a plurality of face images and face depth time sequence images are obtained.

[0087] Step S420, analyzing the plurality of face images by using the face recognition model to obtain a plurality of eye region images;

[0088] Specifically, after the face RGB image is input into the face recognition model, the face recognition model obtains the face key landmark points of the face in the face RGB image, then filters out the landmark points belonging to the eye region according to the key points, and finally obtains the eye region according to the landmark points belonging to the eye region, and the picture in the eye region is the eye region map.

[0089] Step S430, map the multiple frames of eye region maps into the corresponding multiple frames of face depth maps respectively to obtain multiple frames of eye region depth maps.

[0090] Specifically, after the multiple frames of eye region maps and the multiple frames of depth maps are mapped into the world coordinate system, the coordinate values of the multiple frames of eye regions are obtained, wherein the coordinates of the eye region are multiple coordinate values, then the region in the depth map corresponding to the coordinate values of the eye region is identified according to the coordinate values of the eye region, and the obtained region is the eye region depth map.

[0091] The above-mentioned living body detection method obtains multiple frames of face maps and face depth maps by analyzing the face RGB image and the face depth image by using the face recognition model, then obtains multiple frames of eye region maps by analyzing the face maps by using the face recognition model, and finally maps the multiple frames of eye region maps into the multiple frames of face depth maps to obtain multiple frames of eye region depth maps. Since the living body detection probability is obtained based on the living body probability obtained by comprehensively analyzing the face maps, the eye region maps and the eye region depth maps, the face RGB image, the eye region map and the eye region depth map become indispensable conditions for improving the accuracy of the living body detection probability.

[0092] In one embodiment, after step S430, steps S440-460 can be further included.

[0093] Step S440, normalize the multiple frames of face maps to obtain multiple frames of normalized face maps;

[0094] Normalization processing refers to transforming a numerical value into a decimal number between 0 and 1, which is to transform a dimensional expression into a dimensionless expression, and the purpose is to facilitate data processing.

[0095] It should be noted that the pixel value range of each point of the face map is a value between 0 and 255, but this value is too large for the preset classification model (improved neural network model), which will cause the calculation speed of the model to be too slow, so generally, the image data is input into the neural network model. Normalization processing is performed to improve the processing efficiency of the model for data.

[0096] Specifically, the pixel value of each point of the multi-frame face map is divided by 255, that is, normalized, and then the normalized pixel value of each point of the multi-frame face map is obtained.

[0097] In step S450, the multi-frame normalized face map is input into a preset classification model to obtain a live body detection probability of the multi-frame face map output by the preset classification model.

[0098] In step S460, the live body detection probabilities of the multi-frame face maps are weighted and averaged to obtain a second live body detection probability.

[0099] It should be noted that the live body detection of a single frame face map is not sufficient to indicate that the face image collected by the camera is a real image of the user to be detected, and therefore, the embodiments of the present application analyze and identify multiple frames of face maps, and the live body detection probabilities obtained after inputting different frames of face maps into the preset classification model are inconsistent. Finally, the probability that the face image collected by the current camera is a live body is calculated based on the live body detection probabilities of the multi-frame face maps.

[0100] In one embodiment, when the live body detection probabilities of the multi-frame face maps are weighted and averaged, a face map with a higher live body detection probability should be given a larger weight value to avoid a face map with a lower live body detection probability from dominating, thereby effectively improving the accuracy of the live body detection probability based on the face map.

[0101] Specifically, after the multi-frame face maps are normalized to obtain the pixel values of the points of the multi-frame face maps, the pixel values are input into a preset classification model. The preset classification model calculates the pixel values of the points of each frame of face map, and finally outputs a probability value of each frame of face map being a live body. Finally, the probabilities of each frame of face map being a live body are weighted and averaged to obtain a live body detection probability based on the face map.

[0102] The above live body detection method normalizes multiple frames of face maps, then inputs the normalized face maps into a preset classification model, then outputs the live body detection probabilities of the multi-frame face maps, and finally weights and averages the live body detection probabilities of the different frames of face maps to obtain a live body detection probability based on the face map. Since the accuracy of the live body detection probability of the present application is obtained by integrating the live body detection probability based on the face map, the live body detection probability of the eye region map, and the live body detection probability of the eye region depth map, the live body probability of the face map provides an effective basis for improving the accuracy of the live body detection probability.

[0103] In one embodiment, the preset classification model includes a customized neural network model; the customized neural network model includes an SE Block module, an Adam algorithm, and a cosine annealing algorithm; the SE Block module is used to identify subtle features; and the Adam algorithm and the cosine annealing algorithm are used to optimize internal parameter values.

[0104] It can be understood that, in order to meet the needs of the living body detection method of the embodiments of the present application, the initial lightweight neural network model is improved.

[0105] Please refer to Figure 4 The structure diagram of the customized neural network model provided by the embodiments of the present application is shown.

[0106] As Figure 4 shown, the customized neural network model refers to the improved model of the lightweight neural network model, and the lightweight neural network model has fewer model parameters and its performance is not worse than that of a heavier model. The lightweight neural network model includes an input layer (not shown), a convolution layer (Conv), a pooling layer (not shown), and an output layer (not shown), wherein the convolution layer includes an SE Block module, an Adam optimizer is added to adjust the learning rate of different parameters, and the learning rate is continuously reduced by a cosine annealing function to speed up the convergence speed during model training. The lightweight neural network includes but is not limited to SqueezeNet, MobileNetV3, ShuffeNet, Xception, etc., and the specific selection can be made according to the actual application requirements, which is not specifically limited in the embodiments of the present application.

[0107] The SE Block module is an image recognition structure, which strengthens important features to improve accuracy by modeling the correlation between feature channels, that is, the algorithm inside the neural network model can focus on some important feature information. The SE Block module can be set on any layer classifier in the neural network model, or can be set on a certain layer classifier according to actual needs, which is not specifically limited in the embodiments of the present application.

[0108] The Adam optimizer is a first-order optimization algorithm that can replace the traditional stochastic gradient descent process, which can iteratively update the weight values of each layer classifier in the neural network based on the training data, that is, it can adjust different learning rates for each different parameter, for example: the frequently changing parameters are updated with a smaller step, while the sparse parameters are updated with a larger step. Its advantages are high efficiency in calculation, less memory requirement; suitable for large-scale data and parameter scenarios; suitable for unstable objective functions; suitable for sparse or large noise gradient problems.

[0109] Cosine Annealing is an algorithm for decaying the learning rate, which is often used to reduce the learning rate with a cosine function, which can accelerate the convergence speed of the model and the model effect is better. The principle is: when approaching the global minimum of the loss function, the learning rate should become smaller, and as the independent variable of the cosine function increases, the cosine function value first slowly decreases, then accelerates to decrease, and then slows down. If a larger learning rate is selected at this time, the model may oscillate, so the learning rate needs to be decayed to gradually stabilize the model. The cosine function is: where η t represents the learning rate after cosine annealing, represents the minimum value of the learning rate of the i-th hot restart, represents the maximum value of the learning rate of the i-th hot restart, i represents the i-th hot restart, T cur represents the number of times of training of the neural network model, T i represents the total number of times of training of the neural network model.

[0110] Specifically, the SE Block module is added in the process of training the lightweight neural network model, so that the lightweight neural network model can pay attention to some feature information about the face graph. The Adam algorithm is added, so that the lightweight neural network model can set different learning rates for different feature information of the face graph. The cosine annealing function is also added, so that the lightweight neural network model can accelerate the convergence of the training of each facial feature information. Moreover, the initial learning rate of the lightweight neural network model is set to 0.00045. After the above training, a preset classification model is obtained, so that the preset classification model can meet the needs of the living body detection method of the embodiment.

[0111] The above living body detection method adds the SE Block module to the lightweight neural network model to enable the algorithm to recognize subtle features. The Adam algorithm and the cosine annealing algorithm are also added to the lightweight neural network to enable the algorithm to iteratively update the weights and accelerate the convergence speed, and then a preset classification model is obtained. Since the preset classification model includes the SE Block module and the Adam algorithm and the cosine annealing algorithm, it can effectively better recognize the features of the face graph, thereby improving the recognition effect of the face graph and further improving the accuracy of the living body detection probability based on the face graph.

[0112] In one embodiment, after step S460, steps S470-S490 can also be included.

[0113] Step S470, normalize the multi-frame eye region map and the multi-frame eye region depth map to obtain a multi-frame normalized eye region map and a multi-frame normalized eye region depth map.

[0114] Specifically, after the normalization processing of the images of the multi-frame eye region map and the multi-frame eye region depth map, the pixel values of each point of the multi-frame eye region map and the pixel values of each point of the multi-frame eye region depth map are obtained.

[0115] Step S480, input the multi-frame normalized eye region map and the multi-frame normalized eye region depth map into the double fluid model to obtain the live detection probability of the multi-frame eye region map and the eye region depth map output by the double fluid model;

[0116] It should be noted that since the live detection of the single-frame eye region map and the single-frame eye region map is not sufficient to indicate that the face image captured by the camera and the depth time sequence image captured by the electronic device with Lidar technology are real images of the user to be detected, the embodiments of the present application analyze and identify the multi-frame eye region map and the multi-frame eye region depth map, and the live detection probability of the fusion of the two is inconsistent after the eye region map of different frames and each frame eye region depth map are input into the double fluid model. Finally, the probability that the face image captured by the current camera and the face depth time sequence image captured by the current electronic device with Lidar technology are live is calculated based on the live detection probability of the fusion of the two.

[0117] Specifically, based on the above, the pixel values of each point of the multi-frame eye region map and the pixel values of each point of the multi-frame eye region depth map are input into the double fluid model. The spatial flow in the double fluid model calculates the pixel values of each point of the multi-frame eye region map to obtain the live detection probability of the multi-frame eye region map, and the spatial flow of the double fluid model calculates the pixel values of each point of the multi-frame eye region depth map to obtain the live detection probability of the multi-frame eye region depth map. Then, according to the softmax, the live detection probability of each frame eye region map and the corresponding frame eye region map is fused to obtain the live detection probability of the fusion of the multi-frame eye region map and the eye region depth map.

[0118] Step S490, weight average the face live detection probability of the multi-frame eye region map and the eye region depth map to obtain a third live detection probability.

[0119] Optionally, when the live detection of the multi-frame face map is weighted and averaged, a larger weight value should be given to a group of eye region maps and eye region depth maps with higher live detection probability, so as to avoid the dominant position of a group of eye region maps and eye region depth maps with lower live detection, thereby effectively improving the accuracy of the live detection probability based on the face map.

[0120] Specifically, the live body detection probabilities fused by the plurality of eye region maps and the eye region depth maps are weighted and averaged to obtain the live body detection probability based on the eye region map and the eye region depth map.

[0121] The live body detection method, by inputting the plurality of eye region maps and the eye region depth maps into the double-flow model, obtaining the live body detection probability based on the eye region map and the eye region depth map, weighting and averaging the live body detection probabilities of the eye region maps and the corresponding eye region depth maps of different frames to obtain a plurality of sets of live body detection probabilities of the eye region maps and the eye region depth maps, and weighting and averaging the live body detection probabilities of the plurality of sets of eye region maps and eye region depth maps to obtain the live body detection probability based on the eye region map and the eye region depth map, the accuracy of the live body detection probability is obtained based on the live body detection probability of the face map, the live body detection probability fused by the eye region map and the eye region depth map, so that the live body detection probability fused by the eye region map and the eye region depth map provides an effective basis for improving the accuracy of the live body detection probability.

[0122] In one embodiment, the double-flow model includes a first stream structure and a second stream structure, as shown in Figure 5 The structure of the double-flow model provided by the embodiment of the application is shown.

[0123] As shown in Figure 5 The double-flow model includes a first stream structure and a second stream structure, the first stream structure processes the delicate image frame to obtain shape information, the second stream structure processes the continuous multiple frames of dense optical flow to obtain motion information, and finally the shape information and the motion information are fused through the output layer (softmax) for classification.

[0124] Step S481, inputting the plurality of eye region maps into the first stream structure to obtain the live body detection probabilities of the plurality of eye region maps output by the first stream structure;

[0125] Step S482, inputting the plurality of eye region depth maps into the second stream structure to obtain the live body detection probabilities of the plurality of eye region depth maps output by the second stream structure;

[0126] Step S483, weighting and averaging the face live body detection probabilities of the plurality of eye region maps and the eye region depth maps to obtain a third live body detection probability.

[0127] In one embodiment, the first stream structure respectively has 5 convolutional layers (conv), 3 pooling layers (not shown), 2 fully connected layers (FC) and 1 output layer (softmax), the convolutional layers are used for feature extraction of pixel values of each point of the input eye region map, the pooling layers (not shown) are used for selection and information filtering of the features after the feature extraction by the convolutional layers, the fully connected layers are used for nonlinear combination of the selected and information filtered features to obtain an output, and the output layer is used for normalization processing of the output features, and the cumulative sum of the values of the multiple feature normalization processing is 1, and then the probability value of one of the features is output; the second stream structure respectively has 5 convolutional layers, 3 pooling layers (not shown), 2 fully connected layers (FC) and 1 softmax layer, the second stream structure has the same functions as the layers in the first stream structure, except that the second stream structure is used for processing the eye region depth map. Among them, the first stream structure and the second stream structure further include a weighted average layer (class score fusion) at the end, which is used for weighted average of the living body detection probability value of multiple eye region maps and the living body probability value of the eye region depth map to obtain the living body detection probability based on the eye region map and the eye region depth map.

[0128] The above living body detection method, through the first stream structure and the second stream structure in the double fluid model, respectively identifies the multiple frames of eye region maps and the multiple frames of eye region depth maps, then performs weighted average on the living body detection probability of each frame of eye region map output by the first stream structure and the living body detection probability of the corresponding eye region depth map output by the second stream structure to obtain the living body detection probability after fusion of each frame of eye region and the corresponding eye region depth map, and finally performs weighted average on the living body detection probabilities after fusion of the multiple frames of eye region and the corresponding eye region depth map to obtain the living body detection probability based on the eye region map and the eye region depth map. Since the first stream structure and the second stream structure in the double fluid model can identify and fuse the eye region map and the eye region depth map, the living body detection probability of the face image of the to-be-detected person can be effectively determined from different dimensions.

[0129] In one embodiment, after step S483, step S484 can also be included.

[0130] In step S484, the second living body detection probability and the third living body detection probability are weighted and averaged to obtain the first living body detection probability.

[0131] It should be noted that in order to ensure that the living body detection method of the embodiment of the present application has higher accuracy compared with the traditional living body detection technology, different aspects of facial image information need to be extracted for comprehensive analysis, so the face is combined on the basis of the eye region graph and the eye region depth graph.

[0132] In one embodiment, the weighting formula is: P (cls) = 0.454 * clsa + 0.545 * clsb, wherein P (cls) represents the first living body detection probability, clsa represents the second living body detection probability, and clsb represents the third living body detection probability. Among them, the weight value of the second living body detection probability, that is, the living body detection probability based on the face graph, is 0.454, and the weight value of the third living body detection probability, that is, the living body detection probability based on the eye region graph and the eye region depth graph, is 0.545. From the above, it can be known that the living body detection method of the embodiment of the present application is more recognized for the living body detection probability detected according to the eye region graph and the eye region depth graph. The synthesized eye region graph and eye region depth graph of the blinking action will be deformed, and the pixel values of each point of the synthesized eye region graph and eye region graph of the blinking action are inconsistent with the pixel values of each point of the normally collected eye region graph and eye region depth graph. This can clearly distinguish the eye region graph and eye region depth graph for the double-fluid model, whether it is synthesized or actually collected, so the living body detection probability of the eye region graph and the eye region depth graph is given a higher weight value.

[0133] Specifically, after the face graph is calculated and processed by the preset classification model to obtain the living body detection probability based on the face graph, and the eye region graph and the eye region depth graph are calculated and processed by the double-fluid model to obtain the living body detection probability based on the eye region graph and the depth graph, the living body detection probabilities of the two are weighted and averaged to obtain the final living body detection probability.

[0134] The above living body detection method, by weighting and averaging the living body detection probability based on the face graph and the living body detection probability based on the eye region graph and the eye region depth graph, obtains the final living body detection probability, so that the living body detection probability can be comprehensively analyzed from the three aspects of the face graph, the eye region graph and the eye region depth graph, thereby improving the accuracy of the living body detection probability.

[0135] Step S600, judging whether the first living body detection probability is greater than a preset threshold value;

[0136] The preset threshold refers to a standard probability value of the living body detection, which can be in a range of 0.5-0.7, and a specific value can be set according to actual application requirements. The present application does not make a specific limitation here. For example, if the living body detection method is applied to a bank deposit machine, in this case, a higher living body detection accuracy is definitely required, and the maximum value 0.7 can be set, so that external attacks can be better resisted. If it is a simple authentication of identity information on a personal user's mobile phone, in this case, a higher living body detection accuracy is not required, and the minimum value 0.5 can be set to facilitate the improvement of user experience.

[0137] Specifically, after the face RGB image and the face depth time sequence image are analyzed and processed by the face recognition model, the preset classification model and the double fluid model, the living body detection probability based on the face RGB image and the face depth time sequence image is obtained, and then it is determined whether the living body detection probability is greater than the preset standard probability of the living body detection.

[0138] In step S800, if the first living body detection probability is greater than the preset threshold, it is determined that the living body detection is passed.

[0139] Specifically, when the living body detection probability based on the face RGB image and the face depth time sequence image is greater than the preset standard probability of the living body detection, it is indicated that the camera and the electronic device with Lidar technology collect the face image information of the real user to be detected, rather than the face image information replaced or artificially synthesized by a network hacker (i.e. an external attacker), that is, the face image information obtained by the electronic device passes the living body detection.

[0140] The above living body detection method obtains the face RGB image with blinking action and the face depth time sequence image, identifies the face RGB image and the face depth time sequence image based on the face recognition model and the classification model, obtains the living body detection probability, and then determines whether the current user passes the living body detection according to the preset living body detection probability threshold. Since the face RGB image with blinking action and the face depth time sequence image are used at the same time, the confusion of the living body detection by the to-be-detected personnel using only the image with two-dimensional features is effectively avoided, the accuracy of the living body detection is improved, and the ability to resist external attacks is further improved.

[0141] In addition, in one or more embodiments of the present application, the process of living body detection does not require user participation, and the user experience is completed without user awareness, thereby improving the user experience.

[0142] In one embodiment, after step S800, the method can further include:

[0143] In step S900, if the first living body detection probability is less than or equal to the preset threshold, it is determined that the living body detection fails.

[0144] Specifically, when the living body detection probability based on the face RGB image and the face depth time sequence image is less than or equal to the preset standard probability of living body detection, it is indicated that the camera and the electronic device with the Lidar technology collect face image information replaced or artificially synthesized by a network hacker (i.e., an external attacker), that is, the face image information obtained by the electronic device fails the living body detection.

[0145] The living body detection method determines whether the current user passes the living body detection according to the preset living body detection probability threshold. If the detection probability does not exceed the preset living body detection probability, it is indicated that the current user does not pass the living body detection, so that external attacks are prevented, and loss to the user is avoided.

[0146] Please refer to Figure 6 The living body detection device 200 provided by the embodiment of the application is shown in a structural schematic diagram. The living body detection device includes:

[0147] The acquisition module 110 is configured to acquire face image information with a blinking action. The face image information with the blinking action includes a face RGB image and a face depth time sequence image.

[0148] The detection module 120 is configured to obtain a first living body detection probability based on the face RGB image, the face depth time sequence image, a face recognition model, and a classification model.

[0149] The judgment module 130 is configured to determine whether the first living body detection probability is greater than a preset threshold.

[0150] The determination module 140 is configured to determine that the living body detection passes if the first living body detection probability is greater than the preset threshold.

[0151] Optionally, the classification model includes a preset classification model and a double-fluid model. The detection module 120 can be further configured to:

[0152] The face recognition model is used to analyze the face RGB image and the face depth time sequence image to obtain a plurality of frames of face images and a plurality of frames of face depth images.

[0153] The face recognition model is used to analyze the plurality of frames of face images to obtain a plurality of frames of eye region images.

[0154] The plurality of frames of eye region images are respectively mapped into the corresponding plurality of frames of face depth images to obtain a plurality of frames of eye region depth images.

[0155] Optionally, the detection module 120 can be further configured to:

[0156] normalize the plurality of face images to obtain a plurality of normalized face images;

[0157] input the plurality of normalized face images into the preset classification model to obtain a plurality of live detection probabilities of the plurality of face images output by the preset classification model;

[0158] weight-average the plurality of live detection probabilities of the plurality of face images to obtain a second live detection probability.

[0159] Optionally, the detection module 120 can also be configured to:

[0160] normalize the plurality of eye region images and the plurality of eye region depth images to obtain a plurality of normalized eye region images and a plurality of normalized eye region depth images;

[0161] input the plurality of normalized eye region images and the plurality of normalized eye region depth images into the double fluid model to obtain a plurality of live detection probabilities of the plurality of eye region images and the plurality of eye region depth images output by the double fluid model;

[0162] weight-average the plurality of live detection probabilities of the plurality of eye region images and the plurality of eye region depth images to obtain a third live detection probability.

[0163] Optionally, the double fluid model includes a first stream structure and a second stream structure; the detection module 120 can also be configured to:

[0164] input the plurality of eye region images into the first stream structure to obtain a plurality of live detection probabilities of the plurality of eye region images output by the first stream structure;

[0165] input the plurality of eye region depth images into the second stream structure to obtain a plurality of live detection probabilities of the plurality of eye region depth images output by the second stream structure;

[0166] weight-average the plurality of live detection probabilities of the plurality of eye region images and the plurality of eye region depth images to obtain a third live detection probability.

[0167] Optionally, the detection module 120 can also be configured to:

[0168] weight-average the second live detection probability and the third live detection probability to obtain a first live detection probability.

[0169] Optionally, the preset classification model comprises a customized neural network model; the customized neural network model comprises an SE Block module, an Adam algorithm and a cosine annealing algorithm; the SE Block module is used for identifying subtle features; and the Adam algorithm and the cosine annealing algorithm are used for optimizing internal parameter values.

[0170] Optionally, the determining module 140 can also be configured to:

[0171] If the first living body detection probability is less than or equal to a preset threshold, it is determined that the living body detection fails.

[0172] It should be understood that the apparatus corresponds to the living body detection method embodiments described above, and can perform each step involved in the above method embodiments. The specific functions of the apparatus can be referred to the description above. To avoid repetition, the detailed description is appropriately omitted here. The apparatus comprises at least one software function module stored in the memory in the form of software or firmware or solidified in the operating system (OS) of the apparatus.

[0173] In addition to the above embodiments, the present application also provides a storage medium having a computer program stored thereon, and the computer program is executed by the processor 113 to perform the above method.

[0174] The storage medium can be implemented by any type of volatile or nonvolatile storage devices or a combination thereof, such as a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic storage, a flash memory, a magnetic disk or an optical disk.

[0175] It should be understood that the disclosed apparatus and method can also be implemented in other manners. The embodiments described above are merely exemplary embodiments of the present application. In the embodiments of the present application, the described apparatus embodiments are merely schematic, and the functions of the flowcharts and the block diagrams can be implemented in other manners. For example, the flowcharts and the block diagrams can be implemented by using a computer program, a digital signal processor, an application-specific integrated circuit, a field programmable gate array, a processor, a controller, another hardware device, or a combination thereof. In this case, the disclosed apparatus and method can be implemented in a form of a computer program product. The computer program product is directly downloadable from a network, or stored in a computer-readable storage medium, and includes a plurality of instructions. When the instructions are executed by a processor, the processor performs the method according to the embodiments of the present application.

[0176] In addition, each functional module in each of the embodiments of the present application can be integrated together to form a separate part, or each module can exist independently, or two or more modules can be integrated to form a separate part.

[0177] The above description is merely optional implementation of the embodiments of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical range disclosed by the embodiments of the present application, which should be covered within the protection scope of the embodiments of the present application.

Claims

1. A method of detecting living matter, characterized by, The living body detection method comprises: obtaining face image information with a blinking action; wherein the face image information with the blinking action comprises a face RGB image and a face depth time sequence image; obtaining a first living body detection probability based on a face recognition model and a classification model according to the face RGB image and the face depth time sequence image; wherein the classification model comprises a preset classification model and a double fluid model, the preset classification model comprises a customized neural network model; the customized neural network model comprises an SE Block module, an Adam algorithm and a cosine annealing algorithm; the SE Block module is used for identifying subtle features; the Adam algorithm and the cosine annealing algorithm are used for optimizing internal parameter values; determining whether the first living body detection probability is greater than a preset threshold value; if the first living body detection probability is greater than the preset threshold value, determining that the living body detection is passed; the step of obtaining the first living body detection probability based on the face recognition model and the classification model according to the face RGB image and the face depth time sequence image comprises: analyzing the face RGB image and the face depth time sequence image by using the face recognition model to obtain a plurality of face images and a plurality of face depth images; analyzing a plurality of the face images by using the face recognition model to obtain a plurality of eye region images; mapping the plurality of eye region images into corresponding plurality of face depth images respectively to obtain a plurality of eye region depth images; normalizing a plurality of the face images to obtain a plurality of normalized face images; inputting the plurality of normalized face images into the preset classification model to obtain a plurality of living body detection probabilities of the face images output by the preset classification model; weighting and averaging the plurality of living body detection probabilities of the face images to obtain a second living body detection probability.

2. The liveness detection method according to claim 1, characterized in that, after the step of weighting and averaging the plurality of living body detection probabilities of the face images to obtain the second living body detection probability, the method further comprises: normalizing a plurality of the eye region images and a plurality of the eye region depth images to obtain a plurality of normalized eye region images and a plurality of normalized eye region depth images; inputting the plurality of normalized eye region images and the plurality of normalized eye region depth images into the double fluid model to obtain a plurality of living body detection probabilities of the eye region images and the eye region depth images output by the double fluid model; weighting and averaging a plurality of face living body detection probabilities of the eye region images and the eye region depth images to obtain a third living body detection probability.

3. The method of claim 2, wherein the step of detecting the living body is performed by detecting a change in the intensity of the reflected light. wherein, the double fluid model comprises a first stream structure and a second stream structure; the step of inputting the plurality of normalized eye region images and the plurality of normalized eye region depth images into the double fluid model to obtain a plurality of living body detection probabilities of the eye region images and the eye region depth images output by the double fluid model comprises: input the multiple frames of the eye region images into the first stream structure to obtain a living body detection probability of multiple eye region images output by the first stream structure; input the multiple frames of the eye region depth images into the second stream structure to obtain a living body detection probability of multiple eye region depth images output by the second stream structure; perform weighted average on the face living body detection probabilities of the multiple eye region images and eye region depth images to obtain the third living body detection probability.

4. The method according to claim 2 or 3, wherein the living body is detected by the method. The method further comprises: perform weighted average on the second living body detection probability and the third living body detection probability to obtain the first living body detection probability.

5. The method of claim 1, wherein the step of detecting the living body is characterized by: The method further comprises:

6. A living body detecting apparatus characterized by comprising: if the first living body detection probability is less than or equal to a preset threshold, determining that the living body detection fails. The device comprises: an acquisition module configured to acquire face image information with a blinking action; wherein the face image information with the blinking action comprises a face RGB image and a face depth time-series image; a detection module configured to obtain a first living body detection probability based on a face recognition model and a classification model according to the face RGB image and the face depth time-series image; wherein the classification model comprises a preset classification model and a double-fluid model, the preset classification model comprises a customized neural network model, the customized neural network model comprises an SE Block module, an Adam algorithm and a cosine annealing algorithm, the SE Block module is configured to identify subtle features, and the Adam algorithm and the cosine annealing algorithm are configured to optimize internal parameter values; a judgment module configured to judge whether the first living body detection probability is greater than a preset threshold; a determination module configured to, if the first living body detection probability is greater than the preset threshold, determine that the living body detection passes. The detection module is further configured to analyze the face RGB image and the face depth time-series image by using the face recognition model to obtain multiple frames of face images and multiple frames of face depth images, analyze the multiple frames of face images by using the face recognition model to obtain multiple frames of eye region images, and map the multiple frames of eye region images into corresponding multiple frames of face depth images respectively to obtain multiple frames of eye region depth images.

7. A computer-readable storage medium, characterized in that, The detection module is further configured to perform normalization processing on the multiple frames of face images to obtain multiple frames of normalized face images, input the multiple frames of normalized face images into the preset classification model to obtain a living body detection probability of multiple frames of face images output by the preset classification model, and perform weighted average on the living body detection probabilities of the multiple frames of face images to obtain a second living body detection probability. The computer program is stored on the computer readable storage medium and is run by the processor to execute the method according to any one of claims 1 to 5. The computer program is stored on the computer readable storage medium and is run by the processor to execute the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Human face living body detection method, system and device and readable storage medium

    CN111209820A

  • Human face living body detection method and device and computer equipment

    CN111626163A

  • Bimodal living body detection method based on facial expression unit and eye movement

    CN112381050A