Recognition method, recognition system, and electronic device
By acquiring depth and grayscale images using a single TOF camera module, and combining ambient light estimation and automatic exposure control, the complexity of multi-camera module systems and the impact of environmental factors are solved, achieving efficient and low-cost liveness detection and face recognition.
Patent Information
- Application Number
- CN202111635738.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-29
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2041-12-29
AI Technical Summary
Existing liveness detection and facial recognition technologies require multiple camera modules, resulting in complex and costly system structures that are susceptible to environmental influences, leading to inconsistent recognition accuracy.
A single TOF camera module is used to acquire depth and grayscale images, including confidence and intensity maps. These images are used for liveness detection and face recognition, combined with ambient light estimation and automatic exposure control to adapt to different lighting environments and scenes.
It simplifies the system structure, reduces costs, and improves the accuracy of liveness detection and face recognition, making it suitable for various application scenarios.
Smart Images

Figure CN116434286B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of camera modules, and more particularly to an identification method, an identification system and an electronic device. BACKGROUND
[0002] Face is an important biological feature, and live body detection and face recognition technology have been widely applied in security monitoring, financial payment, health care and other fields. However, at present, live body detection and face recognition technology in practical application all need to configure multiple camera modules (for example, structured light camera module, RGB camera module, infrared camera module), the system structure is complex, and the cost is high. Moreover, the live body detection and face recognition results are easily affected by the environment, and are difficult to adapt to various application scenarios. Accordingly, in different scenarios, the accuracy of live body detection and face recognition is inconsistent, in other words, the accuracy of live body detection and face recognition results is easy to change with the change of environment, and it is difficult to maintain at a high level.
[0003] Therefore, a new identification scheme is expected to be applied to different application scenarios. SUMMARY
[0004] One advantage of the present application is to provide an identification method, an identification system and an electronic device, wherein the identification method can adjust the live body detection scheme according to different application scenarios to adapt to different application scenarios.
[0005] Another advantage of the present application is to provide an identification method, an identification system and an electronic device, wherein the identification method can be applied in different light environments to reduce the interference of ambient light on live body detection and face recognition.
[0006] Still another advantage of the present application is to provide an identification method, an identification system and an electronic device, wherein the identification method can adjust the input data for live body detection and face recognition according to different application scenarios, so as to reduce the influence of the difference between different application scenarios on live body detection and face recognition from the data source, so as to improve the accuracy of identification.
[0007] Still another advantage of the present application is to provide an identification method, an identification system and an electronic device, wherein the identification method reduces the influence of the difference between different application scenarios on live body detection and face recognition from the data source, so as to relatively reduce the data processing difficulty of live body detection and face recognition.
[0008] Still another advantage of the present application is to provide an identification method, an identification system and an electronic device, wherein in the identification system, a single TOF camera module can meet the requirement of image data integrity for live body detection and face recognition without the assistance of other camera modules, so as to simplify the system structure and reduce the industrial cost.
[0009] To achieve at least one of the above advantages or other advantages and an object, according to an aspect of the present application, there is provided an identification method, comprising:
[0010] obtaining a depth map and a grayscale map of a target object, the grayscale map comprising a confidence map and an intensity map; and
[0011] performing a living body detection based on the confidence map, the intensity map and the depth map.
[0012] In the identification method according to the present application, performing the living body detection based on the confidence map, the intensity map and the depth map comprises: determining foreground regions and background regions in the confidence map based on a comparison between pixel values of each pixel point in the confidence map and a preset threshold; determining foreground regions and background regions of the intensity map based on positions of the foreground regions of the confidence map in the confidence map; determining a global transformation curve of the intensity map based on the foreground regions of the intensity map through a color estimation model, and performing a global transformation on the intensity map with the global transformation curve of the intensity map to obtain a first globally transformed grayscale map; and performing the living body detection based on the first globally transformed grayscale map and the depth map.
[0013] In the identification method according to the present application, the identification method further comprises: in response to a result of the living body detection being that a face of the target object is a living body face, performing a face recognition based on the first globally transformed grayscale map and the depth map.
[0014] Determining the foreground regions and the background regions in the confidence map based on the comparison between the pixel values of each pixel point in the confidence map and the preset threshold comprises: counting the pixel values of the pixel points in each region in the confidence map to obtain an adaptive threshold value of each region in the confidence map; and comparing the pixel values of the pixel points in each region in the confidence map with the adaptive threshold value to determine the foreground regions and the background regions in the confidence map.
[0015] In the identification method according to the present application, obtaining the depth map and the grayscale map of the target object comprises: obtaining an original image of the target object through a TOF camera module; and analyzing the original image to generate the confidence map and the intensity map.
[0016] In the identification method according to the present application, the original image is parsed to generate the confidence map and the intensity map, including: determining a plurality of first pixel values based on the sum of the photoelectric conversion data of the first detection node and the photoelectric conversion data of the second detection node of the original image in a plurality of preset phase frames, and determining a plurality of second pixel values based on the difference between the photoelectric conversion data of the first detection node and the photoelectric conversion data of the second detection node of the original image in a plurality of preset phase frames; and obtaining the intensity map based on the plurality of first pixel values, and obtaining the confidence map based on the plurality of second pixel values.
[0017] In the identification method according to the present application, the identification method further includes determining an ambient light estimation result based on the photoelectric conversion data of the original image.
[0018] In the identification method according to the present application, the living body detection is performed based on the confidence map, the intensity map and the depth map, including: in response to the ambient light estimation result being strong light, performing living body detection based on the confidence map, the intensity map and the depth map.
[0019] In the identification method according to the present application, the identification method further includes: in response to the ambient light estimation result being weak light, performing living body detection based on the confidence map and the depth map.
[0020] In the identification method according to the present application, in response to the ambient light estimation result being weak light, performing living body detection based on the confidence map and the depth map, including: determining a foreground region and a background region in the confidence map based on a comparison between the pixel value of each pixel point in the confidence map and a preset threshold; determining a global transformation curve of the confidence map based on the foreground region of the confidence map through a color estimation model, and performing global transformation on the confidence map with the global transformation curve of the confidence map to obtain a second global transformation grayscale image; and performing living body detection based on the second global transformation grayscale image and the depth map.
[0021] In the identification method according to the present application, the identification method further includes: in response to the result of the living body detection being that the face of the photographed target belongs to a living body face, performing face recognition based on the second global transformation grayscale image and the depth map.
[0022] In the identification method according to the present application, the ambient light estimation result is determined based on the photoelectric conversion data of the original image, including: counting the photoelectric conversion data of the original image in a preset phase frame to obtain a photoelectric conversion statistical value of the original image; and determining an ambient light estimation result based on the photoelectric conversion statistical value of the original image and a preset photoelectric conversion threshold.
[0023] In the identification method according to the present application, the identification method further comprises: automatically adjusting exposure of the TOF camera module.
[0024] In the identification method according to the present application, the automatic exposure adjustment of the TOF camera module comprises: in response to the absence of a face in the grayscale image, counting the number of all overexposed pixel points in the grayscale image, and adjusting the exposure time based on a comparison between the number of all overexposed pixel points in the grayscale image and a first preset value; and in response to the presence of a face in the grayscale image, counting the average value of pixels in the face region of the grayscale image, and adjusting the exposure time based on a comparison between the average value of pixels in the face region of the grayscale image and the second preset value.
[0025] According to another aspect of the present application, an identification system is provided, comprising:
[0026] a source data acquisition unit configured to acquire a depth map and a grayscale map of a target object, the grayscale map comprising a confidence map and an intensity map; and
[0027] a first detection unit configured to perform living body detection based on the confidence map, the intensity map and the depth map.
[0028] In the identification system according to the present application, the first detection unit is further configured to: determine a foreground region and a background region in the confidence map based on a comparison between the pixel value of each pixel point in the confidence map and a preset threshold; determine a foreground region and a background region of the intensity map based on the position of the foreground region of the confidence map in the confidence map; determine a global transformation curve of the intensity map based on the foreground region of the intensity map through a color estimation model, and perform global transformation on the intensity map with the global transformation curve of the intensity map to obtain a first globally transformed grayscale map; and perform living body detection based on the first globally transformed grayscale map and the depth map.
[0029] In the identification system according to the present application, the identification system further comprises: a face recognition unit configured to perform face recognition based on the first globally transformed grayscale map and the depth map in response to the result of the living body detection being that the face of the target object belongs to a living body face.
[0030] In the identification system according to the present application, the source data acquisition unit is further configured to: acquire an original image of the target object through a TOF camera module; and analyze the original image to generate the confidence map and the intensity map; the identification system further comprises: an ambient light estimation unit configured to determine an ambient light estimation result based on photoelectric conversion data of the original image; and the first detection unit is further configured to perform living body detection based on the confidence map, the intensity map and the depth map in response to the ambient light estimation result being strong light.
[0031] In the recognition system according to the present application, the recognition system further comprises a second detection unit configured to perform living body detection based on the confidence map and the depth map; and the second detection unit is further configured to perform living body detection based on the confidence map and the depth map in response to the ambient light estimation result being weak light.
[0032] In the recognition system according to the present application, the second detection unit is further configured to: determine a foreground region and a background region in the confidence map based on a comparison between a pixel value of each pixel point in the confidence map and a preset threshold; determine a global transformation curve of the confidence map based on the foreground region of the confidence map by a color estimation model, and perform global transformation on the confidence map by using the global transformation curve of the confidence map to obtain a second globally transformed grayscale map; and perform living body detection based on the second globally transformed grayscale map and the depth map.
[0033] In the recognition system according to the present application, the face recognition unit is further configured to perform face recognition based on the second globally transformed grayscale map and the depth map in response to a result of the living body detection being that a face of the measured target belongs to a face of a living body.
[0034] According to yet another aspect of the present application, an electronic device is provided, comprising:
[0035] a memory; and
[0036] a processor, wherein the memory stores computer program instructions which, when executed by the processor, cause the processor to perform any of the above-mentioned recognition methods.
[0037] The further objects and advantages of the present application will be more fully apparent from the following description and the accompanying drawings.
[0038] These and other objects, features and advantages of the present application will become apparent from the following detailed description of the application, the appended claims and the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS
[0039] These and / or other aspects and advantages of the present application will become apparent and more readily appreciated from the following detailed description, taken in conjunction with the accompanying drawings, in which:
[0040] Figure 1 FIG. 1 illustrates a flowchart of a recognition method according to an embodiment of the present application.
[0041] Figure 2 FIG. 2 illustrates a flowchart of performing living body detection based on the confidence map, the intensity map and the depth map in the recognition method according to an embodiment of the present application.
[0042] Figure 3 FIG. 6 illustrates a flow diagram of a process of performing liveness detection based on the confidence map and the depth map in response to the ambient light estimation result being weak light, according to an embodiment of the present application.
[0043] Figure 4 FIG. 7 illustrates a flow diagram of a process of automatic exposure regulation in the recognition method, according to an embodiment of the present application.
[0044] Figure 5 FIG. 8 illustrates a flow diagram of a process of liveness detection in the recognition method, according to an embodiment of the present application.
[0045] Figure 6 FIG. 9 illustrates a flow diagram of a process of face recognition in the recognition method, according to an embodiment of the present application.
[0046] Figure 7 FIG. 10 illustrates a flow diagram of a process of another face recognition in the recognition method, according to an embodiment of the present application. DETAILED DESCRIPTION
[0047] The following description is presented to enable any person skilled in the art to practice the present application as claimed. Stated examples are provided for purposes of illustration and description, and various modifications to the preferred embodiments can be practiced with the benefits that are derived from the description while not departing from the spirit and scope of the application. The basic principles defined in the following description can be applied to other embodiments, variations, improvements, equivalents, and other technical solutions without departing from the spirit and scope of the present application.
[0048] SUMMARY
[0049] As described above, both liveness detection and face recognition technologies require multiple camera modules (e.g., a structured light camera module, an RGB camera module, an infrared camera module) to be configured in practical applications, and the system structure is complex and the cost is high. Moreover, the liveness detection and face recognition results are easily affected by the environment and are difficult to adapt to various application scenarios. Accordingly, the accuracy of liveness detection and face recognition is inconsistent in different scenarios, in other words, the accuracy of liveness detection and face recognition results is easily changed with the change of the environment, and it is difficult to maintain at a high level.
[0050] Specifically, in liveness detection and face recognition, multiple aspects of information of the photographed target need to be obtained to improve the accuracy of the detection result, and accordingly, the characteristics of different camera modules can be used to obtain multiple aspects of information (e.g., depth information, grayscale information) of the photographed target. Moreover, the characteristics of different camera modules can be used to compensate for the respective functional deficiencies. However, such a design scheme not only increases the complexity of the system structure, but also makes the data processing process relatively cumbersome.
[0051] In different application scenarios, the imaging effect will be affected by environmental conditions, and then the results of living body detection and face recognition are affected. In the traditional living body detection and face recognition results, the same type of data is used for living body detection and face recognition for different scenes. However, the same type of data and data processing method is usually only applicable to a specific scene, and the detection result obtained in this scene has high accuracy, while the detection result obtained in other scenes has relatively low accuracy.
[0052] Correspondingly, to solve the above problems, the inventors of the present application propose that, on the one hand, the image acquisition system in living body detection and face recognition is simplified to reduce the complexity of the system structure. On the other hand, different data is applied for living body detection and face recognition according to different application scenarios, that is, the influence of the difference between different application scenarios on living body detection and face recognition is reduced from the data source to improve the recognition accuracy.
[0053] Specifically, in the present application, the gray-scale image is processed to obtain a confidence map removing environmental light information and retaining active light information and an intensity map containing active light and environmental light information. In a normal indoor environment, the confidence map can meet the requirements of image quality for living body detection and face recognition, and can balance the relationship between detection speed and image quality. In an outdoor strong light environment, the signal-to-noise ratio and image quality of the confidence map containing only active light information are far inferior to those of the intensity map, and the intensity map can be used for living body detection and face recognition.
[0054] In particular, in the process of dividing the foreground region and the background region of the intensity map, the confidence map can be used as a mask to obtain the foreground region and the background region of the intensity map. Specifically, in an outdoor strong light environment, since the confidence map removes environmental light information, the background region is almost not imaged, so the foreground region and the background region in the confidence map are more easily distinguished. Since the pixels of the confidence map and the intensity map are aligned, the confidence map can be used as a mask to obtain the foreground region and the background region of the intensity map, thereby improving the accuracy of region division and the accuracy of living body detection and face recognition.
[0055] Based on this, the application provides a recognition method, which comprises: acquiring a depth map and a grayscale map of a target, the grayscale map comprising a confidence map and an intensity map; and performing living body detection based on the confidence map, the intensity map and the depth map; wherein performing living body detection based on the confidence map, the intensity map and the depth map comprises: determining foreground and background regions in the confidence map based on comparison between pixel values of each pixel point in the confidence map and a preset threshold; determining foreground and background regions of the intensity map based on positions of the foreground region of the confidence map in the confidence map; determining a global transformation curve of the intensity map based on the foreground region of the intensity map through a color estimation model, and performing global transformation on the intensity map with the global transformation curve of the intensity map to obtain a first globally transformed grayscale map; and performing living body detection based on the first globally transformed grayscale map and the depth map.
[0056] Also, the application further provides a recognition system, which comprises: a source data acquisition unit configured to acquire a depth map and a grayscale map of a target, the grayscale map comprising a confidence map and an intensity map; and a first detection unit configured to perform living body detection based on the confidence map, the intensity map and the depth map, the first detection unit being further configured to: determine foreground and background regions in the confidence map based on comparison between pixel values of each pixel point in the confidence map and a preset threshold; determine foreground and background regions of the intensity map based on positions of the foreground region of the confidence map in the confidence map; determine a global transformation curve of the intensity map based on the foreground region of the intensity map through a color estimation model, and perform global transformation on the intensity map with the global transformation curve of the intensity map to obtain a first globally transformed grayscale map; and perform living body detection based on the first globally transformed grayscale map and the depth map.
[0057] The application further provides an electronic device, which comprises: a memory; and a processor, wherein computer program instructions are stored in the memory, and the computer program instructions, when executed by the processor, cause the processor to perform the recognition method as described above.
[0058] After introducing the basic principles of the application, various non-limiting embodiments of the application will be specifically introduced below with reference to the accompanying drawings.
[0059] Exemplary identification method
[0060] As shown in FIG. 1, the recognition method according to an embodiment of the application is illustrated. As shown in FIG. 2, the recognition system according to an embodiment of the application is illustrated. Figures 1-7 Figure 1 As shown, the identification method comprises: S110, acquiring a depth map and a grayscale map of a photographed target, the grayscale map comprising a confidence map and an intensity map; and S120, performing living body detection based on the confidence map, the intensity map and the depth map.
[0061] In step S110, a depth map and a grayscale map of a photographed target are acquired. Specifically, in the embodiments of the present application, the depth map and the grayscale map can be acquired directly, or the original image of the photographed target can be acquired through the TOF camera module, and then the original image is analyzed to obtain the depth map and the grayscale map, and different types of grayscale maps can be obtained through different analysis methods. Accordingly, step S110 of acquiring the depth map and the grayscale map of the photographed target comprises: acquiring the original image of the photographed target through the TOF camera module; and analyzing the original image to generate the confidence map and the intensity map.
[0062] Specifically, taking an indirect time-of-flight camera module (iTOF camera module) as an example, the data of each frame of original image (i.e., raw image) acquired through the TOF camera module comprises photoelectric conversion data of 4 phase (0°, 90°, 180°, 270°) frames. The current-assisted photodetector demodulator (CAPD) in each photosensitive chip pixel generates an alternating electric field between two detection nodes through alternating voltage, so as to guide the electrons generated by the photodiode to the two detection nodes respectively. Therefore, the photoelectric conversion data of each phase frame further comprises photoelectric conversion data A and B of the two detection nodes. The confidence map adopts an analysis method of (A-B) to obtain the pixel value of each pixel point, which offsets the common environmental infrared light component of the two detection nodes, while retaining the phase information of the modulated infrared light, and the intensity map adopts an analysis method of (A+B) to obtain the pixel value of each pixel point, which retains all the environmental infrared light components, but does not contain the phase information of the modulated infrared light.
[0063] It is worth mentioning that, here, the analysis method of (A-B) refers to an analysis method of using a function related to the difference between the photoelectric conversion data of the two detection nodes of each phase frame for analysis, that is, the pixel value of the confidence map is related to the difference between the photoelectric conversion data of the two detection nodes of each phase frame of the original image. The analysis method of (A+B) refers to an analysis method of using a function related to the sum of the photoelectric conversion data of the two detection nodes of each phase frame for analysis, that is, the pixel value of the intensity map is related to the sum of the photoelectric conversion data of the two detection nodes of each phase frame of the original image.
[0064] Accordingly, in some embodiments of the present application, in the process of analyzing the original image to generate the confidence map and the intensity map, first, a plurality of first pixel values are determined based on the sum of the photoelectric conversion data of the first detection node and the photoelectric conversion data of the second detection node of the original image at a plurality of preset phase frames (for example, four phases of 0°, 90°, 180°, and 270°), and a plurality of second pixel values are determined based on the difference between the photoelectric conversion data of the first detection node and the photoelectric conversion data of the second detection node of the original image at a plurality of preset phase frames; then, the intensity map is obtained based on the plurality of first pixel values, and the confidence map is obtained based on the plurality of second pixel values.
[0065] In one specific example of the present application, the pixel value of each pixel point of the intensity map (i.e., the plurality of first pixel values) is positively correlated with the sum of the photoelectric conversion data of the first detection node and the photoelectric conversion data of the second detection node of the original image at each phase frame. Accordingly, the pixel value of each pixel point of the intensity map can be obtained by first determining the sum data of each preset phase frame based on the photoelectric conversion data of the first detection node and the photoelectric conversion data of the second detection node of the original image at a plurality of preset phase frames, wherein the sum data of each preset phase frame is the sum of the photoelectric conversion data of the first detection node and the photoelectric conversion data of the second detection node of the each preset phase frame; then determining the pixel value of each pixel point of the intensity map based on the plurality of sum data of the plurality of preset phase frames by a first function positively correlated with the sum data of each preset phase frame of the plurality of preset phase frames.
[0066] In this specific example, the pixel value of each pixel point of the confidence map (i.e., the plurality of second pixel values) is positively correlated with the difference value of the difference data of each two mutually spaced preset phase frames, wherein the difference data of each preset phase frame is the difference between the photoelectric conversion data of the first detection node and the photoelectric conversion data of the second detection node of the each preset phase frame, and the difference value of the difference data of each two mutually spaced preset phase frames is the difference between the difference data of the preset phase frame and the difference data of the preset phase frame spaced from the preset phase frame. Accordingly, the photoelectric conversion data of the confidence map can be obtained by first determining the difference data of each preset phase frame based on the photoelectric conversion data of the first detection node and the photoelectric conversion data of the second detection node of the original image at a plurality of preset phase frames, and then determining the pixel value of each pixel point of the confidence map by a second function positively correlated with the difference value of the difference data of each two mutually spaced phase frames of the plurality of preset phase frames.
[0067] It is worth mentioning that in the identification method, a single TOF camera module can meet the requirement of image data integrity for live body detection and face recognition without the assistance of other camera modules (for example, an RGB camera module, a structured light camera module, and an infrared camera module), that is, multiple camera modules are not required to meet the requirement of image data integrity for live body detection and face recognition.
[0068] The TOF camera module can emit near-infrared light, and the grayscale image obtained by the TOF camera module can also be referred to as an infrared image or an IR image. Specifically, the TOF camera module includes a circuit board, a photosensitive chip, a lens seat, a light filtering element, and an optical lens, and the laser projected by the TOF camera module is modulated into near-infrared light by the light filtering element before being emitted.
[0069] The TOF camera module can be an indirect time-of-flight camera module (iTOF camera module) or a direct time-of-flight camera module (dTOF). Accordingly, the time of flight can be obtained by measuring the phase shift of the laser when the laser is emitted and when the laser reflected by the photographed target is received, and then the depth data of the photographed target can be obtained, or the time taken from emitting the laser to receiving the laser reflected by the photographed target can be directly measured, and then the depth of the photographed target can be obtained.
[0070] It is also worth mentioning that in order to adapt to the application requirements of multiple scenes, the image quality of the depth map and the grayscale image can be regulated by controlling the exposure parameters to improve the accuracy of identification. Specifically, the TOF camera module can be automatically exposed and regulated according to whether a face exists in the depth map and the grayscale image.
[0071] More specifically, as Figure 4As shown, first, the face detection result of the gray image of the current frame and the gray image of the previous frame adjacent to the current frame can be acquired. Then, if the current frame is the first frame, the face detection result of the previous frame is false by default, i.e., there is no face in the gray image. Then, if the current frame is the first frame, the face detection result of the previous frame is false by default, or the current frame is not the first frame, and the face detection result of the previous frame is false, the number of all overexposed pixel points of the gray image (confidence map or intensity map, or other types of gray images) is counted. If the number of all overexposed pixel points (N) in the gray image is less than the first preset value (Nr), the exposure time is extended, and if the number of all overexposed pixel points (N) in the gray image is greater than the first preset value (Nr), the exposure time is shortened, wherein the exposure time is not less than a first threshold (Tmin) and not greater than a second threshold (Tmax). If the current frame is not the first frame, and the face detection result of the previous frame is true, i.e., there is a face, the average value of the pixels in the face region of the gray image is counted. If the average value of the pixels in the face region (Im) of the gray image is less than a second preset value (Ir), the exposure time is extended, and if the average value of the pixels in the face region of the gray image is greater than the second preset value, the exposure time is shortened.
[0072] It is worth mentioning that in the process of acquiring the face detection result, the gray image obtained by analyzing the original image can be subjected to preliminary face detection to obtain the face detection result, or the optimized gray image can be subjected to face detection to obtain the face detection result, which is not limited by the present application. Similarly, in the process of determining the face region, the gray image obtained by analyzing the original image can be subjected to face region recognition and cutting, or the optimized gray image can be subjected to face region recognition and cutting, which is not limited by the present application.
[0073] The above automatic exposure control method combines the advantages of global automatic exposure and local automatic exposure. When no face is detected, global automatic exposure can ensure the quality of the gray image under various lighting conditions, thereby meeting the needs of face detection. Once a face is detected, local automatic exposure can further adjust the exposure for the face, thereby improving the depth and gray quality of the face region, and further improving the accuracy of living body detection and face recognition. The step size of exposure extension and shortening can be adjusted according to the requirements of the application scene for exposure convergence speed.
[0074] Accordingly, the recognition system further comprises: automatically adjusting exposure of the TOF camera module. And the automatically adjusting exposure of the TOF camera module comprises: in response to the absence of a face in the grayscale image, counting the number of all overexposed pixel points in the grayscale image, and adjusting the exposure time based on a comparison between the number of all overexposed pixel points in the grayscale image and a first preset value; and in response to the presence of a face in the grayscale image, counting the average value of pixels in the face region of the grayscale image, and adjusting the exposure time based on a comparison between the average value of pixels in the face region of the grayscale image and the second preset value.
[0075] Adjusting the exposure time based on a comparison between the number of all overexposed pixel points in the grayscale image and the first preset value comprises: in response to the number of all overexposed pixel points in the grayscale image being less than the first preset value, lengthening the exposure time; and in response to the number of all overexposed pixel points in the grayscale image being greater than the first preset value, shortening the exposure time.
[0076] Adjusting the exposure time based on a comparison between the average value of pixels in the face region of the grayscale image and the second preset value comprises: in response to the average value of pixels in the face region of the grayscale image being less than the second preset value, lengthening the exposure time; and in response to the average value of pixels in the face region of the grayscale image being greater than the second preset value, shortening the exposure time.
[0077] In some embodiments of the present application, after the depth map, confidence map and intensity map are obtained, ambient light intensity estimation can be performed to select source data (input data) for live body detection according to the light conditions in the actual application. In some embodiments of the present application, the light conditions are determined based on photoelectric conversion data in the original image.
[0078] Accordingly, in some embodiments of the present application, the recognition method, before step 120 of performing live body detection based on the confidence map, the intensity map and the depth map, further comprises: determining an ambient light estimation result based on photoelectric conversion data of the original image. Specifically, the photoelectric conversion data of the original image at a preset phase frame is counted, and the count value or related data of the count value is compared with a preset threshold value to determine the ambient light intensity condition through the comparison result.
[0079] That is, in some embodiments of the present application, determining an ambient light estimation result based on photoelectric conversion data of the original image comprises: counting the photoelectric conversion data of the original image at a preset phase frame to obtain a photoelectric conversion count value of the original image; and determining an ambient light estimation result based on the photoelectric conversion count value of the original image and a preset photoelectric conversion threshold value.
[0080] In one specific example of the present application, the photoelectric conversion statistical value of the original image is obtained by counting the sum of the photoelectric conversion data of the first detection node and the photoelectric conversion data of the second detection node of the original image in at least one preset phase frame. In other embodiments of the present application, the photoelectric conversion statistical value of the original image can be obtained in other ways, for example, by counting the photoelectric conversion data (including the ambient infrared light component) of the first detection node (or the second detection node) of the original image in at least one preset phase frame, and the present application is not limited thereto.
[0081] In one specific embodiment of the present application, the 99th percentile (P99) of the sum of the photoelectric conversion data of the first detection node and the photoelectric conversion data of the second detection node of the original image in at least one preset phase frame is taken as the photoelectric conversion statistical value, and the ambient light intensity condition is determined by comparing the 99th percentile (P99) with a preset threshold value (Pt). When the 99th percentile (P99) is greater than the preset threshold value (Pt), it indicates that the ambient light intensity is strong enough, otherwise, it indicates that the ambient light intensity does not reach the preset intensity. Of course, the photoelectric conversion statistical value can also be obtained in other ways, and the present application is not limited thereto.
[0082] It is worth mentioning that the ambient light intensity condition can also be determined by comparing the statistical value of the pixel value of the grayscale image with a preset threshold value. In a variant embodiment of the embodiment of the present application, in the process of determining the ambient light estimation result, first, the pixel values of each pixel point of the grayscale image are counted to obtain the pixel statistical value of the grayscale image; then, the ambient light estimation result is determined based on the pixel statistical value of the grayscale image and a preset pixel threshold value.
[0083] In the embodiment of the present application, the confidence map is an active light imaging without ambient light interference, and the intensity map is an imaging of the superposition of active light and ambient light. That is, the confidence map only contains information corresponding to active light, and the intensity map contains information corresponding to both active light and ambient light. In a normal indoor scene, the active light component accounts for the majority, and the image quality of the confidence map can meet the needs of living body detection and face recognition. However, in an outdoor strong light scene, the ambient light component (mainly including the infrared light component in sunlight) accounts for the majority, which easily interferes with the active light of the TOF camera module, and the signal-to-noise ratio and image quality of the confidence map containing only information corresponding to active light are far inferior to those of the intensity map.
[0084] Accordingly, the depth map, the intensity map and the confidence map can be used as input data for the living body detection in a strong light environment (i.e., under a condition of high illuminance of ambient light), and the depth map and the confidence map can be used as input data (source data) for the living body detection in a weak light environment (i.e., under a condition of low illuminance of ambient light). In this way, the input data for the living body detection and the face recognition is adjusted according to different application scenarios, so as to reduce the influence of the difference between different application scenarios on the living body detection and the face recognition from the data source, improve the recognition accuracy, and relatively reduce the data processing difficulty of the living body detection and the face recognition.
[0085] Accordingly, in some embodiments of the present application, the recognition method comprises: step S120, performing living body detection based on the confidence map, the intensity map and the depth map; and step S130, performing living body detection based on the confidence map and the depth map. Step S120 comprises: in response to the ambient light estimation result being strong light, performing living body detection based on the confidence map, the intensity map and the depth map. Step S130 comprises: in response to the ambient light estimation result being weak light, performing living body detection based on the confidence map and the depth map.
[0086] In order to further improve the image quality of the depth map and the grayscale map to meet the application requirements of various scenarios, the depth map and the grayscale map can be optimized. For example, the processing of the depth map includes but is not limited to smoothing filtering, hole filling, invalid point removal, point cloud conversion and various combinations thereof. The processing of the grayscale map includes but is not limited to histogram equalization, smoothing filtering, image enhancement, curve stretching and various combinations thereof, and the present application is not limited thereto.
[0087] In a specific example of the present application, when the ambient light estimation result is strong light, the living body detection is performed by using the intensity map (containing ambient infrared light information). In this process, the following optimization can be performed on the intensity map. First, the foreground region and the background region of the intensity map are determined. Then, the grayscale transformation curve of the global image is determined, and the grayscale transformation is performed on the global image according to the transformation curve. Then, the detailed outline (such as the face outline, eyes, nose wings, lips, etc.) in the image is extracted and the detailed enhancement is performed.
[0088] In particular, in a strong light environment, the pixel value of the foreground region of the confidence map and the pixel value of the background region are quite different, and the background region of the confidence map is almost not imaged (mainly because the confidence map cancels the information corresponding to the ambient light in the analysis process, and the background region of the confidence map has weak reflection of the modulated infrared light), so the foreground region and the background region of the confidence map can be conveniently divided by the preset threshold. The intensity map retains the information corresponding to the ambient light and the active light, and the quality of the intensity map in an outdoor strong light environment is obviously higher than that of the confidence map. However, due to the influence of the ambient light, the foreground region and the background region of the intensity map are not easy to divide.
[0089] Based on this, the inventors of the present application propose that the pixels of the intensity map and the pixels of the confidence map are aligned, the confidence map with the foreground region and the background region that can be relatively easily distinguished is used as a mask to determine the foreground region of the intensity map, so as to facilitate subsequent more targeted improvement of the image quality of the foreground region of the intensity map, and subsequent live body detection and face recognition based on the intensity map. That is, in the process of determining the foreground region and the background region of the intensity map, the foreground region and the background region of the confidence map are first determined, and then the foreground region and the background region of the intensity map are obtained by using the confidence map as a mask. In this way, the accuracy of the region division of the intensity map is improved, and the accuracy of the live body detection and the face recognition is improved.
[0090] Correspondingly, in step S120, live body detection is performed based on the confidence map, the intensity map and the depth map, including: in step S121, determining the foreground region and the background region in the confidence map based on comparison between the pixel value of each pixel point in the confidence map and a preset threshold. Specifically, in the embodiments of the present application, since the object of the live body detection and the face recognition is the face in the foreground region of the image, first distinguishing the foreground region and the background region of the grayscale map can more targetedly improve the image quality of the foreground region of the grayscale map. The preset threshold can be determined according to experience, or can be determined by other methods, for example, iterative selection of threshold, maximum inter-class variance method, and the present application is not limited thereto.
[0091] In one specific example of the present application, the foreground region and the background region in the confidence map are determined by comparing the pixel values of the pixels in the confidence map with an adaptive threshold. Accordingly, in this specific example, the step S121 of determining the foreground region and the background region in the confidence map based on the comparison between the pixel values of the pixels in the confidence map and the preset threshold comprises: counting the pixel values of the pixels in each region in the confidence map to obtain an adaptive threshold for each region in the confidence map; and comparing the pixel values of the pixels in each region in the confidence map with the adaptive threshold to determine the foreground region and the background region in the confidence map. Specifically, the counting of the pixel values of the pixels in each region in the confidence map can obtain a statistical value of the pixel values of the pixels in each region as the adaptive threshold. The statistical value of the pixel values of the pixels in each region can be the average value of the pixel values of the pixels in the region, or a weighted sum, and the weight can be pre-set or obtained through learning.
[0092] Further, the step S120 of performing the living body detection based on the confidence map, the intensity map and the depth map further comprises: a step S122 of determining the foreground region and the background region of the intensity map based on the position of the foreground region in the confidence map. Specifically, since the confidence map and the intensity map are pixel-by-pixel aligned (i.e., the pixels of the confidence map and the pixels of the intensity map correspond one-to-one), the foreground region and the background region of the intensity map corresponding to the foreground region of the confidence map can be determined.
[0093] Since the dynamic range of the TOF camera module is limited, and the illuminance of the ambient light in different scenes (e.g., indoor and outdoor) can be quite different, the pixel values of the face region in the foreground region of the intensity map can be too large or too small to meet the requirements of subsequent living body detection and face recognition. Therefore, based on the determination of the foreground region of the intensity map, the global gray scale transformation curve of the intensity map can be determined according to the foreground region of the intensity map and the gray scale transformation is performed to make the pixel value of the face region in the foreground region of the intensity map meet the requirements.
[0094] Accordingly, the step S120 of performing the living body detection based on the confidence map, the intensity map and the depth map further comprises: a step S123 of determining the global transformation curve of the intensity map based on the foreground region of the intensity map through a color estimation model, and performing global transformation on the intensity map with the global transformation curve to obtain a first globally transformed gray scale map, as shown in Figure 2Specifically, first, a gray mean value M of a foreground region of the intensity image is calculated, and then a global transformation curve of the intensity image is determined by a color estimation model (CEM).
[0095] It is worth mentioning that, in the embodiments of the present application, the global transformation curve of the intensity image is determined by the color estimation model, and the global transformation curve of the intensity image can also be determined by other manners, and the present application is not limited thereto.
[0096] Extracting and enhancing the details (facial features, contours) of the image helps to improve the accuracy of the living body detection and face recognition. Accordingly, step S120 further comprises: extracting local features of the first globally transformed gray image, and performing enhancement processing on the local features to obtain a locally enhanced gray image. It is worth mentioning that the details feature enhancement should be performed after the global transformation of the gray image, and if the order is reversed, it is more difficult to extract the detail features of the dark part of the gray image, and at the same time, the noise of the dark part of the gray image will be amplified.
[0097] After the intensity image and / or the depth image are optimized, the living body detection can be performed based on the depth image and the optimized intensity image. Accordingly, step S120 further comprises: step S124, performing living body detection based on the first globally transformed gray image and the depth image.
[0098] When the ambient light estimation result is weak light, the living body detection is performed by using the confidence map, and in this process, the confidence map can be optimized as follows: first, a foreground region of the confidence map is obtained; then, a gray transformation curve of the global image thereof is determined, and the global image is subjected to gray transformation according to the transformation curve, and then the details contours (such as face contours, eyes, nose wings, lips, etc.) in the image are extracted and enhanced. After the confidence map is optimized, the living body detection can be performed based on the depth image and the optimized confidence map.
[0099] Accordingly, step 130, performing living body detection based on the confidence map and the depth image, comprises: S131, determining a foreground region and a background region in the confidence map based on a comparison between the pixel value of each pixel point in the confidence map and a preset threshold; S132, determining a global transformation curve of the confidence map by a color estimation model based on the foreground region of the confidence map, and performing global transformation on the confidence map by using the global transformation curve of the confidence map to obtain a second globally transformed gray image; and S133, performing living body detection based on the second globally transformed gray image and the depth image, as shown in Figure 3 .
[0100] In the process of performing liveness detection based on the first global transformed grayscale image (or, the second global transformed grayscale image) and the depth map, this application does not limit the specific liveness detection method (liveness recognition method). In a specific example of this application, such as Figure 5 As shown, firstly, face detection needs to be performed on the grayscale image (either an unoptimized grayscale image or an optimized grayscale image) to determine whether a face exists. If no face is detected, the liveness detection result indicates that the subject is not alive; otherwise, the corresponding face region can be cropped from the grayscale image and the depth image. Since the depth image and the grayscale image obtained by the TOF camera module are pixel-wise aligned, the face region in the depth image corresponds to the face region in the grayscale image, eliminating the need for dual / multi-camera extrinsic calibration. Next, liveness detection is performed on the face regions in the depth image and the grayscale image respectively. If both results indicate liveness, the subject is determined to be alive; otherwise, it is determined to be not alive. This application simultaneously uses the depth image and the grayscale image obtained by the TOF camera module, integrating infrared 2D and 3D information, and cross-compares the results, thereby making the liveness detection result more accurate.
[0101] Accordingly, in a specific example of this application, step S124, performing liveness detection based on the first global transformation grayscale image and the depth map, includes: performing face detection based on the first global transformation grayscale image to obtain a face detection result; in response to the face detection result indicating that there is no face in the first global transformation grayscale image, determining the final liveness detection result as the subject being non-live; and in response to the face detection result indicating that there is a face in the first global transformation grayscale image, performing a first liveness detection and a second liveness detection based on the depth map and the first global transformation grayscale image respectively, and in response to the detection results of both the first liveness detection and the second liveness detection being live, determining the final liveness detection result as the subject being live.
[0102] In a specific example of this application, step S133, performing liveness detection based on the second global transformation grayscale image and the depth map, includes: performing face detection based on the second global transformation grayscale image to obtain a face detection result; in response to the face detection result indicating that there is no face in the first global transformation grayscale image, determining the final liveness detection structure as the target being non-live; and in response to the face detection structure indicating that there is a face in the second global transformation grayscale image, performing a third liveness detection and a fourth liveness detection based on the depth map and the second global transformation grayscale image respectively, and in response to the detection structures of the third liveness detection and the fourth liveness detection both being live, determining the final liveness detection structure as the target being live.
[0103] In the embodiments of the present application, the face detection can be performed based on the first globally transformed grayscale image (or the second globally transformed grayscale image), and subsequent data processing can be performed based on the face detection result obtained after the face detection is performed on the first globally transformed grayscale image (or the second globally transformed grayscale image). It should be understood that the face detection can also be performed based on the grayscale image obtained after other optimization processing, and subsequent data processing can be performed based on the face detection result obtained after the face detection is performed on the grayscale image obtained after other optimization processing, and the present application is not limited in this respect.
[0104] It is worth mentioning that, in the embodiments of the present application, the living body detection is described by taking a human as an example, and in actual applications, the living body detection of other animals (for example, dogs and cats) can be performed based on the grayscale image and the depth image.
[0105] It is worth mentioning that, in the variant embodiments of the embodiments of the present application, in a weak light environment, the image quality difference between the intensity image and the confidence image is relatively small, and the depth image, the confidence image and the intensity image can be used as source data (input data) to perform the living body detection and subsequent data processing, so that the step of estimating the ambient light intensity can be omitted.
[0106] The living body detection technology is applied to the face recognition, so that the identity recognition can be performed based on the facial feature information of a human on the premise that the photographed target is a living body, and in this way, the face recognition accuracy can be improved by resisting face photo, face sculpture, mask and other face authentication attacks.
[0107] Correspondingly, in the embodiments of the present application, after the photographed object is identified as a living body, that is, the face of the photographed target belongs to a living body face, the face recognition can be further performed based on the depth image and the grayscale image.
[0108] The recognition method further includes: step S140, in response to the result of the living body detection that the face of the photographed target belongs to a living body face, performing the face recognition based on the first globally transformed grayscale image and the depth image; and step S150, in response to the result of the living body detection that the face of the photographed target belongs to a living body face, performing the face recognition based on the second globally transformed grayscale image and the depth image.
[0109] The present application is not limited to a specific face recognition method. For example, as shown in FIG. 6, the face recognition can be performed based on the first globally transformed grayscale image and the depth image, and the face recognition can also be performed based on the second globally transformed grayscale image and the depth image. Figure 6As shown, first, the face region of the gray image (the gray image without optimization processing or the gray image with optimization processing) and the face region of the depth image (the depth image without optimization processing or the depth image with optimization processing) are fused in two channels, then the face feature of the fused face region is extracted, and the extracted face feature is compared with the face feature in the database (which can be 1:1 comparison or 1:N comparison), so as to obtain the recognition result.
[0110] Correspondingly, in one specific example of the present application, in step S140, in response to the result of the living body detection that the face of the photographed target belongs to the face of a living body, face recognition is performed based on the first globally transformed gray image and the depth image, including: based on the first globally transformed gray image and the depth image, determining the face region of the first globally transformed gray image and the face region of the depth image; fusing the face region of the first globally transformed gray image and the face region of the depth image to obtain a fused face region; extracting the face feature in the fused face region to obtain the face extraction feature of the fused face region; and based on the comparison between the face extraction feature of the fused face region and the face feature in the database, determining the face recognition result.
[0111] In step S150, in response to the result of the living body detection that the face of the photographed target belongs to the face of a living body, face recognition is performed based on the second globally transformed gray image and the depth image, including: based on the second globally transformed gray image and the depth image, determining the face region of the second globally transformed gray image and the face region of the depth image; fusing the face region of the second globally transformed gray image and the face region of the depth image to obtain a fused face region; extracting the face feature in the fused face region to obtain the face extraction feature of the fused face region; and based on the comparison between the face extraction feature of the fused face region and the face feature in the database, determining the face recognition result.
[0112] For another example, as shown in the following table, first, the face feature extraction and comparison are performed on the face region of the gray image (the gray image without optimization processing or the gray image with optimization processing) and the depth image (the depth image without optimization processing or the depth image with optimization processing) respectively, and the consistency comparison is performed on the two recognition results, and then the final recognition result is outputted, and the result of face recognition can be used for subsequent data processing. Figure 7 As shown, first, the face feature extraction and comparison are performed on the face region of the gray image (the gray image without optimization processing or the gray image with optimization processing) and the depth image (the depth image without optimization processing or the depth image with optimization processing) respectively, and the consistency comparison is performed on the two recognition results, and then the final recognition result is outputted, and the result of face recognition can be used for subsequent data processing.
[0113] Accordingly, in another specific example of the present application, in response to the result of the living body detection being that the photographed target face belongs to a living body face, the face recognition based on the first globally transformed grayscale image and the depth image in step S140 comprises: determining a face region of the first globally transformed grayscale image and a face region of the depth image based on the first globally transformed grayscale image and the depth image; extracting face features in the face region of the first globally transformed grayscale image and face features in the face region of the depth image respectively to obtain face extracted features of the first globally transformed grayscale image and face extracted features of the depth image; determining a first face recognition result based on a comparison between the face extracted features of the first globally transformed grayscale image and face features in a database, and determining a second face recognition result based on a comparison between the face extracted features of the depth image and face features in the database; and determining a final face recognition result based on consistency of the first face recognition result and the second face recognition result.
[0114] In response to the result of the living body detection being that the photographed target face belongs to a living body face, the face recognition based on the second globally transformed grayscale image and the depth image in step S150 comprises: determining a face region of the second globally transformed grayscale image and a face region of the depth image based on the second globally transformed grayscale image and the depth image; extracting face features in the face region of the second globally transformed grayscale image and face features in the face region of the depth image respectively to obtain face extracted features of the second globally transformed grayscale image and face extracted features of the depth image; determining a third face recognition result based on a comparison between the face extracted features of the second globally transformed grayscale image and face features in a database, and determining a fourth face recognition result based on a comparison between the face extracted features of the depth image and face features in the database; and determining a final face recognition result based on consistency of the third face recognition result and the fourth face recognition result.
[0115] In this specific example, it should be understood that the face recognition based on the first globally transformed grayscale image (or the second globally transformed grayscale image) and the depth image can also be based on grayscale images and depth images obtained after other optimization processing, and the present application is not limited in this regard.
[0116] In summary, the recognition method is clarified, which can adjust input data for living body detection and face recognition according to different application scenarios, reduce the influence of differences between different application scenarios on living body detection and face recognition from the data source, and improve the accuracy of recognition.
[0117] Exemplary identification system
[0118] According to another aspect of the present application, there is also provided an identification system 10 according to embodiments of the present application, which comprises a source data acquisition unit 11 and a first detection unit 12.
[0119] In particular, the source data acquisition unit 11 is configured to acquire a depth map and a grayscale map of a target object, the grayscale map comprising a confidence map and an intensity map. The first detection unit 12 is configured to perform a liveness detection based on the confidence map, the intensity map and the depth map.
[0120] In one specific example of the present application, the first detection unit 12 is further configured to determine foreground regions and background regions in the confidence map based on a comparison between pixel values of each pixel point in the confidence map and a preset threshold; determine foreground regions and background regions in the intensity map based on positions of the foreground regions in the confidence map; determine a global transformation curve of the intensity map based on the foreground regions in the intensity map through a color estimation model, and perform a global transformation on the intensity map with the global transformation curve of the intensity map to obtain a first globally transformed grayscale map; and perform a liveness detection based on the first globally transformed grayscale map and the depth map.
[0121] In one specific example of the present application, the first detection unit 12 is further configured to count pixel values of pixel points in each region in the confidence map to obtain an adaptive threshold value of each region in the confidence map; and determine foreground regions and background regions in the confidence map based on a comparison between pixel values of each pixel point in the confidence map and the adaptive threshold value.
[0122] It is worth mentioning that, in one specific example of the present application, the source data acquisition unit 11 is further configured to acquire an original image of the target object through a TOF camera module; and analyze the original image to generate the confidence map and the intensity map.
[0123] Specifically, taking an indirect time-of-flight camera module (iTOF camera module) as an example, the data of each frame of raw image (i.e., raw image) obtained by the TOF camera module includes photoelectric conversion data of four phase (0°, 90°, 180°, 270°) frames. The current-assisted photodetector demodulator (CAPD) in each photosensitive chip pixel generates an alternating electric field between two detection nodes through alternating voltage, so as to guide the electrons generated by the photodiode to the two detection nodes respectively. Therefore, the photoelectric conversion data of each phase frame further includes photoelectric conversion data A and B of the two detection nodes. The confidence map adopts an analytical method of (A-B) to obtain the pixel value of each pixel point, offsets the common ambient infrared light component of the two detection nodes, and retains the phase information of the modulated infrared light, while the intensity map adopts an analytical method of (A+B) to obtain the pixel value of each pixel point, retains all the ambient infrared light components, but does not include the phase information of the modulated infrared light.
[0124] It is worth mentioning that, here, the analytical method of (A-B) refers to an analytical method of using a function related to the difference between the photoelectric conversion data of the two detection nodes of each phase frame for analysis, that is, the pixel value of the confidence map is related to the difference between the photoelectric conversion data of the two detection nodes of each phase frame of the raw image. The analytical method of (A+B) refers to an analytical method of using a function related to the sum of the photoelectric conversion data of the two detection nodes of each phase frame for analysis, that is, the pixel value of the intensity map is related to the sum of the photoelectric conversion data of the two detection nodes of each phase frame of the raw image.
[0125] The confidence map is active light imaging for removing ambient light interference, and the intensity map is imaging of superposition of active light and ambient light. That is, the confidence map only includes information corresponding to active light, and the intensity map includes information corresponding to both active light and ambient light. In a normal indoor scene, the active light component accounts for the majority, and the image quality of the confidence map can meet the needs of living body detection and face recognition. However, in an outdoor strong light scene, the ambient light component (mainly including the infrared light component in sunlight) accounts for the majority, which easily interferes with the active light of the TOF camera module, and the signal-to-noise ratio and image quality of the confidence map which only includes information corresponding to active light are far inferior to those of the intensity map.
[0126] Accordingly, the depth map and the confidence map can be used as input data (source data) for the living body detection in a weak light environment (under the condition that the illuminance of the ambient light is low), and the depth map, the intensity map and the confidence map can be used as input data for the living body detection in a strong light environment (under the condition that the illuminance of the ambient light is high). In this way, the input data for the living body detection and the face recognition is adjusted according to different application scenarios, so as to reduce the influence of the difference between different application scenarios on the living body detection and the face recognition from the data source, improve the recognition accuracy, and relatively reduce the data processing difficulty of the living body detection and the face recognition.
[0127] It is worth mentioning that, in the embodiment of the present application, the recognition system 10 can collect source data required for subsequent living body detection and face recognition through a single TOF camera module, without the assistance of other camera modules (for example, an RGB camera module, a structured light camera module, an infrared camera module), that is, without the cooperation of multiple camera modules to meet the requirement of image data integrity for living body detection and face recognition. In this way, the system structure can be simplified, and the industrial cost can be reduced.
[0128] The TOF camera module can emit near-infrared light, and the grayscale map obtained by the TOF camera module can also be called an infrared map or an IR map. Specifically, the TOF camera module includes a circuit board, a photosensitive chip, a lens seat, a light filtering element and an optical lens, and the laser projected by the TOF camera module is modulated into near-infrared light by the light filtering element before being emitted.
[0129] The TOF camera module can be an indirect time-of-flight type camera module (iTOF camera module), or a direct time-of-flight type camera module (dTOF). Accordingly, the time of flight can be obtained by measuring the phase shift of the laser when the laser is emitted and when the laser reflected by the photographed target is received, and then the depth data of the photographed target can be obtained, or the time taken from emitting the laser to receiving the laser reflected by the photographed target can be directly measured, and then the depth of the photographed target can be obtained.
[0130] In one specific example of the present application, the recognition system 10 further comprises an ambient light estimation unit, which is configured to determine an ambient light estimation result based on the photoelectric conversion data of the original map.
[0131] In one specific example of the present application, the ambient light estimation unit is further configured to: count the photoelectric conversion data of the original map in a preset phase frame to obtain a photoelectric conversion statistical value of the original map; and determine an ambient light estimation result based on the photoelectric conversion statistical value of the original map and a preset photoelectric conversion threshold.
[0132] In one specific example of the present application, the recognition system 10 further comprises an ambient light estimation unit configured to determine an ambient light estimation result based on pixel values of the grayscale image.
[0133] In one specific example of the present application, the first detection unit 12 is further configured to, in response to the ambient light estimation result being strong light, perform liveness detection based on the confidence map, the intensity map and the depth map.
[0134] In one specific example of the present application, the recognition system 10 further comprises a second detection unit 13 configured to perform liveness detection based on the confidence map and the depth map, the second detection unit 13 being further configured to, in response to the ambient light estimation result being weak light, perform liveness detection based on the confidence map and the depth map.
[0135] In one specific example of the present application, the second detection unit 13 is further configured to determine foreground regions and background regions in the confidence map based on a comparison between pixel values of each pixel point in the confidence map and a preset threshold; determine a global transformation curve of the confidence map based on the foreground regions of the confidence map by a color estimation model, and perform global transformation on the confidence map by the global transformation curve of the confidence map to obtain a second globally transformed grayscale image; and perform liveness detection based on the second globally transformed grayscale image and the depth map.
[0136] In one specific example of the present application, the recognition system 10 further comprises a face recognition unit 14 configured to, in response to a result of the liveness detection being that a face of the target belongs to a face of a living body, perform face recognition based on the first globally transformed grayscale image and the depth map.
[0137] In one specific example of the present application, the face recognition unit 14 is further configured to, in response to a result of the liveness detection being that a face of the target belongs to a face of a living body, perform face recognition based on the second globally transformed grayscale image and the depth map.
[0138] In one specific example of the present application, the first detection unit 12 is further configured to perform face detection based on the first globally transformed grayscale image to obtain a face detection result, determine that the final live body detection result is that the photographed target is a non-live body in response to the face detection result being that there is no face in the first globally transformed grayscale image, and perform first live body detection and second live body detection based on the depth map and the first globally transformed grayscale image respectively in response to the face detection result being that there is a face in the first globally transformed grayscale image, and determine that the final live body detection result is that the photographed target is a live body in response to the detection results of the first live body detection and the second live body detection both being live bodies.
[0139] In one specific example of the present application, the second detection unit 13 is further configured to perform live body detection based on the second globally transformed grayscale image and the depth map, including performing face detection based on the second globally transformed grayscale image to obtain a face detection result, determining that the final live body detection structure is that the photographed target is a non-live body in response to the face detection result being that there is no face in the first globally transformed grayscale image, and performing third live body detection and fourth live body detection based on the depth map and the second globally transformed grayscale image respectively in response to the face detection structure being that there is a face in the second globally transformed grayscale image, and determining that the final live body detection structure is that the photographed target is a live body in response to the detection structures of the third live body detection and the fourth live body detection both being live bodies.
[0140] In one specific example of the present application, the face recognition unit 14 is further configured to determine a face region of the first globally transformed grayscale image and a face region of the depth map based on the first globally transformed grayscale image and the depth map, fuse the face region of the first globally transformed grayscale image and the face region of the depth map to obtain a fused face region, extract face features in the fused face region to obtain face extraction features of the fused face region, and determine a face recognition result based on a comparison between the face extraction features of the fused face region and face features in a database.
[0141] In one specific example of the present application, the face recognition unit 14 is further configured to determine a face region of the second globally transformed grayscale image and a face region of the depth map based on the second globally transformed grayscale image and the depth map, fuse the face region of the second globally transformed grayscale image and the face region of the depth map to obtain a fused face region, extract face features in the fused face region to obtain face extraction features of the fused face region, and determine a face recognition result based on a comparison between the face extraction features of the fused face region and face features in a database.
[0142] In one specific example of the present application, the face recognition unit 14 is further configured to determine a face region of the first globally transformed grayscale image and a face region of the depth map based on the first globally transformed grayscale image and the depth map; extract face features in the face region of the first globally transformed grayscale image and face features in the face region of the depth map respectively to obtain face extracted features of the first globally transformed grayscale image and face extracted features of the depth map; determine a first face recognition result based on a comparison between the face extracted features of the first globally transformed grayscale image and face features in a database, and determine a second face recognition result based on a comparison between the face extracted features of the depth map and face features in the database; and determine a final face recognition result based on consistency of the first face recognition result and the second face recognition result.
[0143] In one specific example of the present application, the face recognition unit 14 is further configured to determine a face region of the second globally transformed grayscale image and a face region of the depth map based on the second globally transformed grayscale image and the depth map; extract face features in the face region of the second globally transformed grayscale image and face features in the face region of the depth map respectively to obtain face extracted features of the second globally transformed grayscale image and face extracted features of the depth map; determine a third face recognition result based on a comparison between the face extracted features of the second globally transformed grayscale image and face features in a database, and determine a fourth face recognition result based on a comparison between the face extracted features of the depth map and face features in the database; and determine a final face recognition result based on consistency of the third face recognition result and the fourth face recognition result.
[0144] In the embodiments of the present application, the result of face recognition obtained by the face recognition unit 14 can be used for subsequent data processing, or can be transmitted to other systems through various communication modes (such as USB, network, serial port, etc.) as input data of the other systems.
[0145] In one specific example of the present application, the recognition system 10 further comprises an automatic exposure unit 15 configured to perform automatic exposure control on the TOF camera module.
[0146] In one specific example of the present application, the automatic exposure unit 15 is further configured to, in response to the absence of a face in the grayscale image, count the number of all overexposed pixel points in the grayscale image, and adjust the exposure time based on a comparison between the number of all overexposed pixel points in the grayscale image and a first preset value; and in response to the presence of a face in the grayscale image, count the average value of pixels in the face region in the grayscale image, and adjust the exposure time based on a comparison between the average value of pixels in the face region in the grayscale image and a second preset value.
[0147] In one specific example of the present application, the automatic exposure unit 15 is further configured to, in response to the number of all overexposed pixel points in the grayscale image being less than the first preset value, lengthen the exposure time; and in response to the number of all overexposed pixel points in the grayscale image being greater than the first preset value, shorten the exposure time.
[0148] In one specific example of the present application, the automatic exposure unit 15 is further configured to, in response to the average value of pixels in the face region in the grayscale image being less than the second preset value, lengthen the exposure time; and in response to the average value of pixels in the face region in the grayscale image being greater than the second preset value, shorten the exposure time.
[0149] Here, the specific functions of the various units of the recognition system 10 have been described in detail above with reference to the description of the recognition method Figures 1-7 schematically, and thus repeated descriptions thereof will be omitted.
[0150] In summary, the recognition system 10 is illustrated, which can adjust the input data for live body detection and face recognition according to different application scenarios, reduce the influence of differences between different application scenarios on live body detection and face recognition from the data source, and improve the recognition accuracy.
[0151] Exemplary electronic device
[0152] According to yet another aspect of the present application, an electronic device 80 is also provided, which includes a memory 81 and a processor 82, and computer program instructions are stored in the memory 81, which, when executed by the processor 82, cause the processor 82 to perform the recognition method Figures 1-7 schematically. Here, the recognition method has been described in detail above with reference to the description of the recognition method Figures 1-7 schematically, and thus repeated descriptions thereof will be omitted.
[0153] In summary, the electronic device 80 is illustrated, which can perform an optimized identification method to reduce the influence of the difference between different application scenarios on the live detection and face recognition, so as to improve the accuracy of identification.
[0154] It should be understood by those skilled in the art that the above description and the embodiments of the present application shown in the drawings are only examples and do not limit the present application. The purpose of the present application has been fully and effectively achieved. The function and structural principle of the present application has been shown and explained in the embodiments, and the embodiments of the present application can be any modification or modification without departing from the principle.
Claims
1. A method of identification, characterized in that, The method comprises: obtaining a depth map and a grayscale map of a target object, the grayscale map comprising a confidence map and an intensity map; wherein the depth map and the grayscale map of the target object are obtained by: obtaining an original image of the target object by a TOF camera module; and analyzing the original image to generate the confidence map and the intensity map; determining an ambient light estimation result based on photoelectric conversion data of the original image; and in response to the ambient light estimation result being weak light, performing living body detection based on the confidence map and the depth map; in response to the ambient light estimation result being strong light, performing living body detection based on the confidence map, the intensity map and the depth map.
2. The identification method of claim 1, wherein the living body detection based on the confidence map, the intensity map and the depth map comprises: determining a foreground region and a background region in the confidence map based on a comparison between pixel values of each pixel point in the confidence map and a preset threshold value; determining a foreground region and a background region of the intensity map based on a position of the foreground region of the confidence map in the confidence map; determining a global transformation curve of the intensity map based on the foreground region of the intensity map by a color estimation model, and performing global transformation on the intensity map by the global transformation curve of the intensity map to obtain a first globally transformed grayscale map; and performing living body detection based on the first globally transformed grayscale map and the depth map. in response to a result of the living body detection being that a face of the target object belongs to a living body, performing face recognition based on the first globally transformed grayscale map and the depth map.
3. The identification method of claim 2, further comprising: determining a foreground region and a background region in the confidence map based on a comparison between pixel values of each pixel point in the confidence map and a preset threshold value comprises:
4. The identification method according to claim 2, wherein, statistically analyzing pixel values of each region in the confidence map to obtain an adaptive threshold value of each region in the confidence map; and comparing the pixel values of each region in the confidence map with the adaptive threshold value to determine the foreground region and the background region in the confidence map. analyzing the original image to generate the confidence map and the intensity map comprises:
5. The identification method of claim 1, wherein, determining a plurality of first pixel values based on a sum of photoelectric conversion data of a first detection node and photoelectric conversion data of a second detection node of a plurality of preset phase frames of the original image, and determining a plurality of second pixel values based on a difference between the photoelectric conversion data of the first detection node and the photoelectric conversion data of the second detection node of the plurality of preset phase frames of the original image; and obtaining the intensity map based on the plurality of first pixel values, and obtaining the confidence map based on the plurality of second pixel values. in response to the ambient light estimation result being weak light, performing living body detection based on the confidence map and the depth map comprises:
6. The identification method of claim 1, wherein, determining a foreground region and a background region in the confidence map based on a comparison between pixel values of each pixel point in the confidence map and a preset threshold value; determining a global transformation curve of the confidence map based on the foreground region of the confidence map by a color estimation model, and performing global transformation on the confidence map by the global transformation curve of the confidence map to obtain a second globally transformed grayscale map; and Perform liveness detection based on the second globally transformed grayscale image and the depth map.
7. The identification method of claim 6, further comprising: Perform face recognition based on the second globally transformed grayscale image and the depth map in response to a result of the liveness detection being that the face of the photographed target is a live face.
8. The identification method of claim 1, wherein, Determine an ambient light estimation result based on photoelectric conversion data of the original image, including: Statistically determine photoelectric conversion data of the original image at a preset phase frame to obtain a photoelectric conversion statistical value of the original image; and Determine an ambient light estimation result based on the photoelectric conversion statistical value of the original image and a preset photoelectric conversion threshold.
9. The identification method of claim 1, further comprising: Perform automatic exposure regulation on the TOF camera module.
10. The identification method according to claim 9, wherein, Perform automatic exposure regulation on the TOF camera module, including: In response to the grayscale image not containing a face, statistically determine the number of all overexposed pixel points in the grayscale image, and adjust the exposure time based on a comparison between the number of all overexposed pixel points in the grayscale image and a first preset value; and In response to the grayscale image containing a face, statistically determine the average value of pixels in the face region of the grayscale image, and adjust the exposure time based on a comparison between the average value of pixels in the face region of the grayscale image and a second preset value.
11. An identification system characterized by Include: A source data acquisition unit configured to acquire a depth map and a grayscale image of a photographed target, the grayscale image including a confidence map and an intensity map; further configured to: acquire an original image of the photographed target through a TOF camera module; and analyze the original image to generate the confidence map and the intensity map; An ambient light estimation unit configured to determine an ambient light estimation result based on photoelectric conversion data of the original image; A first detection unit configured to perform liveness detection based on the confidence map, the intensity map, and the depth map in response to the ambient light estimation result being strong light; A second detection unit configured to perform liveness detection based on the confidence map and the depth map in response to the ambient light estimation result being weak light.
12. The identification system of claim 11, wherein, The first detection unit is further configured to: Determine foreground and background regions in the confidence map based on a comparison between pixel values of each pixel point in the confidence map and a preset threshold; Determine foreground and background regions of the intensity map based on the location of the foreground region of the confidence map in the confidence map; Determine a globally transformed curve of the intensity map based on the foreground region of the intensity map through a color estimation model, and globally transform the intensity map based on the globally transformed curve of the intensity map to obtain a first globally transformed grayscale image; and Perform liveness detection based on the first globally transformed grayscale image and the depth map. The recognition system further includes a face recognition unit configured to perform face recognition based on the first globally transformed grayscale image and the depth map in response to a result of the liveness detection being that the face of the photographed target is a live face.
13. The identification system of claim 12, wherein, The second detection unit is further configured to:
14. The identification system of claim 11, wherein, Determine foreground and background regions in the confidence map based on a comparison between pixel values of each pixel point in the confidence map and a preset threshold; determine a global transformation curve of the confidence map based on a foreground region of the confidence map through a color estimation model, and perform global transformation on the confidence map with the global transformation curve of the confidence map to obtain a second globally transformed grayscale map; and perform liveness detection based on the second globally transformed grayscale map and the depth map.
15. The identification system of claim 14, wherein, The recognition system further comprises a face recognition unit, which is further configured to, in response to a result of the liveness detection being that a face of the measured target belongs to a live face, perform face recognition based on the second globally transformed grayscale map and the depth map.
16. An electronic device, comprising: comprise: a memory; and a processor, wherein the memory stores computer program instructions, and the computer program instructions, when executed by the processor, cause the processor to perform the recognition method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Recognition and location method for apple on tree based on TOF camera
CN106951905A
Target extraction method and device based on image processing and terminal device
CN109978890A