Face depth image processing method and device, electronic equipment and storage medium

By performing face detection on grayscale images and employing super-resolution technology on speckle images, and dynamically generating irregularly shaped windows for parallax matching, the problems of low accuracy and significant loss of detail information in face depth images acquired by structured light depth cameras are solved, thereby improving the performance of face liveness detection, face recognition, and 3D face modeling.

CN117058067BActive Publication Date: 2025-12-12ZHEJIANG SUNNY INTELLIGENT OPTICAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210482524.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-05
Publication Date
2025-12-12
Estimated Expiration
2042-05-05

AI Technical Summary

Technical Problem

Existing structured light depth cameras acquire face depth images with low accuracy and significant loss of detail, which affects the performance of face liveness detection, face recognition, and 3D face modeling.

Method used

By performing face detection on grayscale images, dynamically generating irregularly shaped windows for sub-region matching, and using super-resolution technology on speckle images to improve the accuracy per unit disparity, combined with a disparity matching algorithm, high-precision face depth images are obtained.

Benefits of technology

It improves the depth accuracy and detail information restoration of the face region, and enhances the algorithm performance of face liveness detection, face recognition and face 3D modeling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117058067B_ABST
    Figure CN117058067B_ABST
Patent Text Reader

Abstract

The present application provides a kind of face depth image processing method, device, electronic equipment and storage medium, it can solve the problem that the precision of face depth image obtained by structured light is low and the problem that the loss of detail information is big.The face depth image processing method includes the following steps: obtaining the gray image and speckle image of target scene;Face detection is carried out on the gray image to obtain the position of face region and multiple face sub-regions;According to the characteristics of each face sub-region in the gray image, dynamically generate the matching window corresponding to each face sub-region;The face region image cut from the speckle image is subjected to super-resolution operation to obtain the face region image to be matched;And by the matching window corresponding to each face sub-region, the corresponding face sub-region on the face region image to be matched is subjected to disparity matching to obtain the actual disparity value of each pixel on the face region image, and then the face depth image is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of machine vision, in particular to a face depth image processing method and device based on structured light, electronic equipment and storage medium. BACKGROUND

[0002] As a representation of face data, each pixel point of the depth image corresponds to the distance (depth) information from the collector to each point in the scene, so that the face data based on the depth image can be applied in the technical fields of face liveness detection, face recognition, face 3D modeling, etc. The existing face depth image is generally obtained by using a depth camera, and the depth camera mainly includes a binocular camera, a structured light depth camera, and a TOF camera based on the light flight time principle, etc. Since the structured light depth camera has the advantages of small size, high resolution and low power consumption compared with other depth camera technologies, the structured light depth camera has been widely applied and researched in the technical fields of face liveness detection, face recognition, face 3D modeling, etc.

[0003] As known, the high degree of restoration and high precision of the detail information in the face depth image are of great significance to the algorithm performance improvement of face liveness detection and face recognition. However, although the structured light technology can collect the face depth image in the scene, the collected face depth image has the problems of low precision and large loss of detail information due to the block matching principle of the parallax matching algorithm, which further causes the subsequent face liveness detection, face recognition, face 3D modeling, etc. to be difficult to obtain ideal performance, and even unable to be smoothly carried out. SUMMARY

[0004] An advantage of the present application is to provide a face depth image processing method, device, electronic equipment and storage medium, which can solve the problems of low precision and large loss of detail information of the face depth image obtained by structured light, and help to improve the algorithm performance of the application based on the face depth image.

[0005] Another advantage of the present application is to provide a face depth image processing method, device, electronic equipment and storage medium, wherein in an embodiment of the present application, the face depth image processing method can divide the face region into sub-regions through face detection technology on the gray image or speckle image, and match the corresponding sub-regions by using different shaped windows, so as to improve the restoration degree of the detail information in the face image, and help to improve the depth details of the face region.

[0006] Another advantage of the present application is to provide a face depth image processing method, device, electronic equipment and storage medium, wherein in one embodiment of the present application, the face depth image processing method can use super resolution to increase the virtual focal length, improve the accuracy of unit parallax, and improve the depth accuracy of the face region.

[0007] Another advantage of the present application is to provide a face depth image processing method, device, electronic equipment and storage medium, wherein in order to achieve the above-mentioned purpose, in the present application, complex logic or complex structure is not required. Therefore, the present application successfully and effectively provides a solution, not only provides a simple face depth image processing method, device, electronic equipment and storage medium, but also increases the practicability and reliability of the face depth image processing method, device, electronic equipment and storage medium.

[0008] In order to achieve the above-mentioned at least one advantage or other advantages and purposes of the present application, the present application provides a face depth image processing method, comprising the steps of:

[0009] Obtaining a gray image and a speckle image;

[0010] Performing face detection on the gray image to obtain a face region position and a plurality of face sub-regions in the gray image;

[0011] According to the characteristics of each face sub-region in the gray image, a matching window corresponding to each face sub-region is dynamically generated;

[0012] According to the face region position, an image corresponding to the face region is cropped from the speckle image to obtain a face region image to be matched; and

[0013] By using the matching window corresponding to each face sub-region, the actual parallax value of each pixel in the face region image is obtained by performing parallax matching on the corresponding face sub-region in the face region image to be matched, and a face depth image is obtained.

[0014] According to one embodiment of the present application, the step of obtaining a gray image and a speckle image comprises the steps of:

[0015] Obtaining a scene speckle image and a gray image of a target scene collected by a monocular structured light camera; and

[0016] Calling a plurality of reference speckle images from a memory, wherein the reference speckle images are obtained by shooting a plane at different distances by the monocular structured light camera.

[0017] According to one embodiment of the present application, the step of obtaining the gray image and the speckle image comprises the steps of:

[0018] cropping a face speckle image corresponding to the face region from the scene speckle image according to the face region position;

[0019] cropping face reference images corresponding to the face region from the plurality of reference speckle images respectively according to the face region position and a reserved disparity search range; and

[0020] performing super-resolution operations on the face speckle image and the plurality of face reference images respectively to obtain a to-be-matched face speckle image and a plurality of to-be-matched face reference images as the to-be-matched face region image.

[0021] According to one embodiment of the present application, the step of obtaining the gray image and the speckle image comprises the steps of:

[0022] obtaining a binocular gray image and a binocular speckle image of a target scene collected by a binocular structured light camera, wherein when a left-eye gray image in the binocular gray image is subjected to face detection, a left-eye speckle image in the binocular speckle image is taken as a scene speckle image, and a right-eye speckle image in the binocular speckle image is taken as a reference speckle image.

[0023] According to one embodiment of the present application, the step of obtaining the gray image and the speckle image comprises the steps of:

[0024] cropping a face speckle image corresponding to the face region from the left-eye speckle image according to the face region position;

[0025] cropping face reference images corresponding to the face region from the right-eye speckle image respectively according to the face region position and a reserved disparity search range; and

[0026] performing super-resolution operations on the face speckle image and the face reference images respectively to obtain a to-be-matched face speckle image and a to-be-matched face reference image as the to-be-matched face region image.

[0027] According to one embodiment of the present application, the step of obtaining the gray image and the speckle image comprises the steps of:

[0028] detecting the face region in the gray image by a deep learning model to obtain the face region position in the target scene;

[0029] Keypoint detection is performed on the face region in the grayscale image to obtain the locations of facial keypoints in the grayscale image; and

[0030] Based on the location of facial key points, the face in the grayscale image is divided into multiple facial sub-regions.

[0031] According to one embodiment of this application, in the step of dividing the face in the grayscale image into multiple face sub-regions based on the location of facial key points: the multiple face sub-regions include a forehead region, an eye region, a nose region, a mouth region, and a cheek region.

[0032] According to one embodiment of this application, the step of dynamically generating a matching window corresponding to each face sub-region based on the characteristics of each face sub-region in the grayscale image includes the following steps:

[0033] Based on the depth characteristics of the human face, dynamically select a matching window of the corresponding shape; and

[0034] The size of the matching window is adjusted in reverse based on the distance to the face region.

[0035] According to one embodiment of this application, the step of dynamically selecting a matching window of a corresponding shape based on the facial depth characteristics includes the following steps:

[0036] Based on the relatively flat depth of the forehead area, a square matching window is selected;

[0037] Based on the relatively elongated horizontal shape of the eye and mouth areas, a horizontally rectangular matching window is selected; and

[0038] Based on the fact that the nose and cheek areas are relatively long and thin in the vertical direction, a vertically rectangular matching window is selected.

[0039] According to one embodiment of this application, the step of performing disparity matching on the corresponding face sub-regions in the face region image to be matched through matching windows corresponding to each face sub-region, so as to obtain the actual disparity value of each pixel in the face region image, and then obtaining the face depth image, includes the following steps:

[0040] Based on the generated matching window, each pixel of the corresponding face sub-region in the face speckle image to be matched is translated left and right within a predetermined disparity range on the face reference image to be matched, so as to obtain the disparity value and similarity array of each pixel of the face sub-region through disparity matching.

[0041] Sort the similarity array by size, and use the disparity value corresponding to the maximum similarity value as the actual disparity value of the corresponding pixel; and

[0042] According to the actual parallax value of each pixel on the face region image, the actual depth value corresponding to each pixel is calculated to obtain the face depth image.

[0043] According to another aspect of the present application, the present application further provides a face depth image processing device, comprising:

[0044] The acquisition module is configured to acquire the grayscale image and the speckle image.

[0045] The detection module is configured to perform face detection on the grayscale image to obtain a face region position and a plurality of face sub-regions in the grayscale image.

[0046] The generation module is configured to dynamically generate a matching window corresponding to each face sub-region according to the characteristics of each face sub-region in the grayscale image.

[0047] The cropping module is configured to crop an image corresponding to the face region from the speckle image according to the face region position to obtain a face region image to be matched; and

[0048] The matching module is configured to perform parallax matching on the corresponding face sub-region in the face region image to be matched through the matching window corresponding to each face sub-region to obtain the actual parallax value of each pixel on the face region image, and further obtain the face depth image.

[0049] According to another aspect of the present application, the present application further provides a face recognition device, comprising:

[0050] The structured light camera module is configured to acquire image data; and

[0051] The data processing module is communicatively connected to the structured light camera module and is configured to perform the steps of any of the face depth image processing methods described above based on the image data from the structured light camera module to perform depth calculation.

[0052] According to another aspect of the present application, the present application further provides an electronic device, comprising:

[0053] A processor; and

[0054] A storage medium having a computing program instruction stored therein, the computing program instruction, when executed by the processor, causing the processor to perform any of the face depth image processing methods described above.

[0055] According to another aspect of the present application, the present application further provides a storage medium having stored thereon computer program instructions which, when executed by a computing device, are operable to perform the method for processing face depth image according to any one of the above aspects. BRIEF DESCRIPTION OF DRAWINGS

[0056] Figure 1 is a flowchart of the method for processing face depth image according to an embodiment of the present application;

[0057] Figure 2A shows a first example of the image obtaining step in the method for processing face depth image according to the above embodiment of the present application;

[0058] Figure 2B shows a second example of the image obtaining step in the method for processing face depth image according to the above embodiment of the present application;

[0059] Figure 3 shows a flowchart of the face detecting step in the method for processing face depth image according to the above embodiment of the present application;

[0060] Figure 4 shows a flowchart of the shape selecting step in the method for processing face depth image according to the above embodiment of the present application;

[0061] Figure 5 shows a flowchart of the size adjusting step in the method for processing face depth image according to the above embodiment of the present application;

[0062] Figure 6A shows a first example of the image operating step in the method for processing face depth image according to the above embodiment of the present application;

[0063] Figure 6B shows a second example of the image operating step in the method for processing face depth image according to the above embodiment of the present application;

[0064] Figure 7 shows a flowchart of the disparity matching step in the method for processing face depth image according to the above embodiment of the present application;

[0065] Figure 8A shows a block diagram of the structured light camera according to the first example of the present application;

[0066] Figure 8B shows a block diagram of the structured light camera according to the second example of the present application;

[0067] Figure 9A schematic diagram showing the principle of super-resolution operation in the processing method of the face depth image according to the above embodiment of the present application is shown.

[0068] Figure 10 An example of face region division in the processing method of the face depth image according to the above embodiment of the present application is shown.

[0069] Figure 11 An example of the shape of the matching window in the processing method of the face depth image according to the above embodiment of the present application is shown.

[0070] Figure 12 An example of the size of the matching window in the processing method of the face depth image according to the above embodiment of the present application is shown.

[0071] Figure 13 An example of the depth of the nose region in the processing method of the face depth image according to the above embodiment of the present application is shown.

[0072] Figure 14 An example of the processing method of the face depth image according to the above embodiment of the present application is shown.

[0073] Figure 15 A block diagram schematic of a face depth image processing apparatus according to an embodiment of the present application is shown.

[0074] Figure 16 A block diagram schematic of a face recognition apparatus according to an embodiment of the present application is shown.

[0075] Figure 17 A block diagram schematic of an electronic device according to an embodiment of the present application is shown.

[0076] Main element symbol explanation: 10, structured light camera device; 11, speckle projector; 12, floodlight projector; 13, infrared camera; 13L, left eye camera; 13R, right eye camera; 14, projection driving module; 15, processor; 16, memory; 20, face depth image processing apparatus; 21, acquisition module; 211, image obtaining module; 212, image calling module; 22, detection module; 221, region detection module; 222, key point detection module; 223, region division module; 23, generation module; 231, shape selection module; 232, size adjustment module; 24, operation module; 241, cropping module; 242, super-resolution module; 25, matching module; 251, translation module; 252, sorting module; 253, calculation module; 30, face recognition apparatus; 31, data processing module; 40, electronic device; 41, processor; 42, storage medium; 43, input device; 44, output device.

[0077] The above main element symbol description is further explained in detail in combination with the drawings and specific embodiments. DETAILED DESCRIPTION

[0078] The following description is presented to enable any person skilled in the art to practice the present application as claimed. The preferred embodiments disclosed herein are merely examples of the best mode of the application and thus are not intended to limit the scope of the application. The principles described herein can be applied to other embodiments, modifications, improvements, equivalents, and other technical solutions without departing from the spirit and scope of the present application.

[0079] It should be understood by those skilled in the art that in the disclosure of the present application, the terms "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. indicate the orientation or positional relationship shown in the drawings, which are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the devices or elements referred to must have a particular orientation, be constructed and operated in a particular orientation, therefore the above terms cannot be understood as a limitation of the present application.

[0080] In the present application, the term "one" in the claims and the description should be understood as "one or more", that is, in one embodiment, the number of one element can be one, and in another embodiment, the number of the element can be multiple. Unless it is explicitly shown in the disclosure of the present application that the number of the element is only one, the term "one" cannot be understood as unique or single, and the term "one" cannot be understood as a limitation on the number.

[0081] In the description of the present application, it should be understood that "first", "second", etc. are only for the purpose of description, and cannot be understood as indicating or implying relative importance. In the description of the present application, it should be noted that unless otherwise explicitly specified and limited, "connected", "connected" should be understood in a broad sense, for example, it can be fixedly connected, or it can be detachably connected or integrally connected; it can be mechanically connected, or it can be electrically connected; it can be directly connected, or it can be indirectly connected through a medium. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.

[0082] In the description of the present specification, the description referring to the terms "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" and the like means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the present specification, the illustrative description of the above terms does not necessarily refer to the same embodiment or example. Also, the specific features, structures, materials or characteristics described can be combined in any appropriate manner in any one or more embodiments or examples. In addition, the person skilled in the art can combine and combine the different embodiments or examples described in the present specification and the features of the different embodiments or examples without contradiction.

[0083] Considering that the face surface is a three-dimensional surface, the positions of the facial features will show obvious depth changes on the depth image; and the problems of low precision and large loss of detail information of the face depth image obtained by the structured light are mainly caused by the coupling of the two factors of low proportion of the same depth area in the matching window and low single-parallax precision at a long distance. Therefore, in view of these reasons, the present application creatively proposes a face depth image processing method, device, electronic equipment and storage medium, which can solve the problems of low precision and large loss of detail information of the face depth image obtained by the structured light, and help to improve the performance of the application algorithm based on the face depth image.

[0084] Specifically, referring to the drawings of the specification of the present application Figure 1 According to an embodiment of the present application, a face depth image processing method is provided, which can include the following steps:

[0085] S100: obtaining a gray image and a speckle image;

[0086] S200: performing face detection on the gray image to obtain a face region position and a plurality of face sub-regions in the gray image;

[0087] S300: dynamically generating a matching window corresponding to each face sub-region according to the characteristics of each face sub-region in the gray image;

[0088] S400: cropping an image corresponding to the face region from the speckle image according to the face region position to obtain a face region image to be matched; and

[0089] S500: performing parallax matching on the corresponding face sub-region on the face region image to be matched through the matching window corresponding to each face sub-region to obtain the actual parallax value of each pixel on the face region image, and further obtaining a face depth image.

[0090] It is worth noting that the face depth image processing method of the present application can not only use matching windows with different lengths and widths for parallax matching for different face sub-regions detected on the grayscale image to improve the proportion of facial feature images in the matching window, so as to improve the depth details of the face region, but also can use precision improvement methods such as super-resolution to increase the virtual focal length for the speckle image of the face region to improve the precision of unit parallax, so as to improve the depth precision of the face region, so as to extract the best face depth image, which helps to improve the performance of application algorithms based on face depth images, such as face liveness detection algorithms, face recognition algorithms or face 3D modeling algorithms.

[0091] More specifically, the speckle image of the present application can include but is not limited to a scene speckle image and a reference speckle image. In other words, the scene speckle image and the grayscale image can be obtained by shooting the target scene with a monocular structured light camera. At this time, the reference speckle image can be obtained by calling from the memory of the monocular structured light camera.

[0092] Exemplarily, in the first example of the present application, as shown in Figure 8A The structured light camera 10 can include a speckle projector 11, a floodlight projector 12, an infrared camera 13, a projection driving module 14, a processor 15 and a memory 16 which are communicatively connected to each other. The processor 15 controls the projection driving module 14 to alternately light up the speckle projector 11 and the floodlight projector 12, so that the speckle projector 11 projects random irregular speckles to the object in the target scene, and the floodlight projector 12 projects uniform infrared illumination to the object in the target scene. The infrared camera 13 collects images of the target scene. The processor 15 receives images from the infrared camera 13, and alternately receives speckle images and grayscale images according to the patterns projected by the projectors. At the same time, when shooting planes at different distances, the processor 15 will store the received speckle images as reference speckle images in the memory 16.

[0093] In other words, according to the above first example of the present application, as shown in Figure 2A The step S100 of the face depth image processing method can include the steps of:

[0094] S110: obtaining a scene speckle image and a grayscale image of a target scene collected by a monocular structured light camera; and

[0095] S120: calling a plurality of reference speckle images from the memory, wherein the reference speckle images are obtained by shooting planes at different distances with the monocular structured light camera.

[0096] It is worth noting that in other examples of the present application, the speckle image can also include a pair of scene speckle images. In other words, the scene speckle image and the grayscale image can be acquired by taking a target scene by a binocular structured light camera. At this time, the scene speckle image includes a left-eye speckle image and a right-eye speckle image.

[0097] Exemplarily, in a second example of the present application, as shown in Figure 8B The structured light camera 10 can include a speckle projector 11, a floodlight projector 12, a left-eye camera 13L, a right-eye camera 13R, a projection driving module 14, a processor 15, and a memory 16, which are communicatively connected to each other. The processor 15 controls the projection driving module 14 to alternately light up the speckle projector 11 and the floodlight projector 12, so that the speckle projector 11 projects random irregular speckles to an object in a target scene, and the floodlight projector 12 projects uniform infrared illumination to the object in the target scene. The left-eye camera 13L and the right-eye camera 13R collect images of the target scene. The processor 15 receives images from the left-eye camera 13L and the right-eye camera 13R, and alternately receives speckle images and grayscale images in the scene according to the patterns projected by the projectors.

[0098] In other words, according to the above-mentioned second example of the present application, as shown in Figure 2B The step S100 of the face depth image processing method can include the following steps:

[0099] S110': obtaining binocular grayscale images and binocular speckle images of a target scene collected by a binocular structured light camera, wherein when performing face detection on a left-eye grayscale image in the binocular grayscale images, a left-eye speckle image in the binocular speckle images is taken as a scene speckle image, and a right-eye speckle image in the binocular speckle images is taken as a reference speckle image. It can be understood that in other examples of the present application, if face detection is performed on the right-eye grayscale image, the right-eye speckle image is taken as the scene speckle image, and the left-eye speckle image is taken as the reference speckle image, which will not be described herein.

[0100] According to the above-mentioned embodiments of the present application, as shown in Figure 3 The step S200 of the face depth image processing method can include the following steps:

[0101] S210: performing face region detection on the grayscale image by a deep learning model to obtain a face region position in the target scene;

[0102] S220: performing key point detection on the face region in the grayscale image to obtain a face key point position in the grayscale image; and

[0103] S230: dividing the face region in the grayscale image into a plurality of face sub-regions according to the face key point position.

[0104] Optionally, the deep learning model used in the present application can be implemented as a Caffe (English: Convolutional Architecture for Fast Feature Embedding; Chinese: Convolutional Neural Network Framework), Tensflow (an open source software library for numerical computation using data flow graphs), Keras (an open source artificial neural network library written in Python), etc. framework model, as long as the position of the face region in the target scene can be obtained, and the present application will not be repeated.

[0105] Optionally, the face key point position of the present application can be implemented as the position of the eyes, nose and mouth in the face region, and the face region can be divided into five face sub-regions according to the position. For example, as shown in Figure 10 , the face sub-region of the present application can include a forehead region a, an eye region b, a nose region c, a mouth region d and a cheek region e.

[0106] According to the above embodiments of the present application, as shown in Figure 4 , the step S300 of the face depth image processing method can include a step S310 of dynamically selecting a matching window corresponding to the shape according to the face depth feature.

[0107] Exemplarily, as shown in Figure 4 and Figure 11 , the step S310 of the face depth image processing method of the present application can include steps:

[0108] S311: according to the feature that the depth of the forehead region is relatively flat, selecting a square matching window;

[0109] S312: according to the feature that the depth of the eye region and the mouth region is relatively elongated in the horizontal direction, selecting a horizontal rectangular matching window; and

[0110] S313: according to the feature that the depth of the nose region and the cheek region is relatively elongated in the vertical direction, selecting a vertical rectangular matching window.

[0111] It is worth noting that, as the distance of the face changes, the proportion of the face sub-region (i.e. the face feature region) in the same matching window will change inversely; that is, as the face distance increases, the proportion of the face sub-region in the matching window will decrease; and as the face distance decreases, the proportion of the face sub-region in the matching window will increase. Therefore, as shown in Figure 5 the step S300 of the face depth image processing method of the present application can further include a step S320 of inversely adjusting the size of the matching window according to the distance of the face region.

[0112] Exemplarily, as shown in Figure 5 and Figure 12 the step S320 of the face depth image processing method of the present application can include steps:

[0113] S321: when the distance of the face region increases, the size of the matching window is reduced; and

[0114] S322: when the distance of the face region decreases, the size of the matching window is increased.

[0115] In this way, the matching window generated by the step S300 of the present application will be a special-shaped window with unequal length and width, and not only the shape of the generated matching window is different for different face sub-regions, but also the size of the matching window used is different for different face distances, for example, the specific shape and size can refer to the matching window as shown in Figure 11 and Figure 12 In addition, taking the nose region as an example, after matching using a matching window with unequal length and width, the nose region depth matching result is as shown in Figure 13 It can be easily known that: the nose region depth obtained by using the face depth image processing method of the present application is closer to the real depth of the nose region; compared with the existing original scheme, the accuracy is significantly improved.

[0116] It is worth noting that in the above first example of the present application, since the shooting scenes of the speckle image and the gray image shot by the monocular structured light camera are consistent, the position information of the face region and the position information of the face sub-region obtained in the step S200 are also applicable to the speckle image, so that the present application can cut out the face speckle image corresponding to the face region from the scene speckle image and the face reference image from the reference speckle image according to the position information of the face region. In particular, when cutting out the reference speckle image, it is necessary to reserve pixels that can cover the disparity search range on the left and right. It can be understood that the face region image in the step S400 of the present application can include the face speckle image cut out from the scene speckle image and the face reference image cut out from the reference speckle image; at the same time, the face region image to be matched in the present application also includes the face speckle image to be matched and the face reference image to be matched.

[0117] In addition, the present application further enlarges the length and width of the cut-out face speckle image and the plurality of face reference images to increase the virtual focal length, thereby improving the accuracy of the unit disparity so as to obtain more accurate face depth information subsequently.

[0118] For example, in the above first example of the present application, as shown in Figure 6A the step S400 of the face depth image processing method of the present application can include the steps of:

[0119] S410: cutting out a face speckle image corresponding to the face region from the scene speckle image according to the face region position;

[0120] S420: cutting out face reference images corresponding to the face region from a plurality of reference speckle images respectively according to the face region position and the reserved disparity search range; and

[0121] S430: performing super-resolution operation on the face speckle image and the plurality of face reference images to obtain a face speckle image to be matched and a plurality of face reference images to be matched as the face region image to be matched.

[0122] It is worth noting that in the above first example of the present application, since the shooting scenes of the speckle image and the gray image shot by the monocular structured light camera are consistent, the position information of the face region and the position information of the face sub-region obtained in the step S200 are also applicable to the speckle image, so that the present application can cut out the face speckle image corresponding to the face region from the scene speckle image and the face reference image from the reference speckle image according to the position information of the face region. In particular, when cutting out the reference speckle image, it is necessary to reserve pixels that can cover the disparity search range on the left and right. It can be understood that the face region image in the step S400 of the present application can include the face speckle image cut out from the scene speckle image and the face reference image cut out from the reference speckle image; at the same time, the face region image to be matched in the present application also includes the face speckle image to be matched and the face reference image to be matched.

[0123] In particular, if disparity matching is performed with the right-eye speckle image as the reference, when the right-eye speckle image is cropped, pixels capable of covering the disparity search range need to be reserved on the left and right sides. It can be understood that the face region image in the step S400 of the method for processing a face depth image according to the present application can include a face speckle image cropped from the left-eye speckle image and a face reference image cropped from the right-eye speckle image; meanwhile, the face region image to be matched according to the present application also includes a face speckle image to be matched and a face reference image to be matched.

[0124] In addition, the face speckle image and the plurality of face reference images cropped according to the present application are further enlarged in length and width to increase the virtual focal length, so as to improve the accuracy of unit disparity, so as to obtain more accurate face depth information subsequently.

[0125] Exemplarily, in the above-mentioned second example of the present application, as shown in Figure 6B the step S400 of the method for processing a face depth image according to the present application can include the steps of:

[0126] S410': cropping a face speckle image corresponding to the face region from the left-eye speckle image according to the face region position;

[0127] S420': respectively cropping face reference images corresponding to the face region from the right-eye speckle image according to the face region position and the reserved disparity search range; and

[0128] S430': performing super-resolution operation on the face speckle image and the face reference images to obtain a face speckle image to be matched and a face reference image to be matched as the face region image to be matched.

[0129] It is worth noting that the reserved disparity search range according to the present application can refer to the number of pixels reserved on the left and right sides of the face region when the reference speckle image is cropped in order to cover the disparity search range; for example, the reserved pixel threshold can be implemented as 50 to 100 pixels.

[0130] In addition, in the step S430 of the present application, by performing super-resolution operation on the face speckle image and the plurality of face reference images, that is, by enlarging the resolution (i.e. the length and width of the image) and the virtual focal length, the accuracy of unit disparity can be improved, so as to obtain more accurate face depth information. In other words, the super-resolution technology is essentially a process of enlarging the resolution (W / H) and the focal length f by the same proportion; and the unit disparity accuracy formula is as follows: Figure 9

[0131]

[0132] In the formula, S is the distance between the projector and the camera lens, i.e., the baseline length; d is the distance from the projector to the reference platform; u is the pixel size represented by each pixel, i.e., the actual distance represented by each pixel; f is the lens focal length; and Offset is the parallax corresponding to the pixel. It can be understood that Δd' / ΔOffset is the actual distance represented by the unit parallax, and the smaller the distance represented by the unit parallax, the higher the accuracy.

[0133] From the above formula, it can be deduced that if f and Offset are increased in proportion, then and (f*S-(Offset+1)*u*d) are increased, so Δd' / ΔOffset will be smaller, i.e., the accuracy is higher. In other words, according to the above formula, as long as the feature point rules of the enlarged image are basically consistent with the original image, the image length and width and the virtual focal length are increased in proportion, which will improve the unit parallax accuracy, i.e., the depth resolution capability is improved.

[0134] According to the above embodiments of the present application, as Figure 7 shown, the step S500 of the face depth image processing method can include the steps of:

[0135] S510: According to the generated matching window, each pixel in the corresponding face sub-region of the face speckle image to be matched is respectively translated left and right within a predetermined parallax range on the plurality of face reference images to be matched, so as to obtain the parallax value and the similarity array of each pixel in the face sub-region through parallax matching;

[0136] S520: The similarity array is sorted by size, and the parallax value corresponding to the maximum similarity value is taken as the actual parallax value of the corresponding pixel; and

[0137] S530: According to the actual parallax value of each pixel on the face region image, the actual depth value corresponding to each pixel is calculated to obtain the face depth image.

[0138] It is worth noting that in the step S510 of the present application, for each pixel in the corresponding face sub-region of the face speckle image to be matched, the corresponding matching window is used to perform left and right translation within a certain parallax range on the plurality of face reference images to be matched, and the best parallax and similarity result matched are output. In this way, after repeatedly matching the plurality of face reference images, a binary array (including parallax value and similarity value) with a length of M will be obtained, where M is the number of face reference images.

[0139] After that, the binary array is sorted according to the size of the similarity value to obtain the disparity value under the maximum similarity value. In this way, the obtained disparity value is the final integer disparity result (i.e. the actual depth value) of the pixel. Finally, through the monocular theoretical calculation model and sub-pixel interpolation, the corresponding high-resolution real face depth, i.e. the face depth image, can be obtained.

[0140] In summary, in one example of the present application, as shown in Figure 14 the processing method of the face depth image can include the following steps: first, the speckle image and the infrared image (i.e. the gray image) of the target scene are obtained respectively; then, after the face detection is performed on the infrared image and the face region position is obtained, on the one hand, the face speckle pattern and the face reference image in the face region are cropped according to the face region position information, and the face speckle pattern and the face reference image are subjected to super-resolution operation; on the other hand, the key points in the face region are located, the face region is segmented according to the face key points, and is divided into the forehead region, the eye region, the nose region, the cheek region and the mouth region to obtain multiple face sub-regions; then, according to the size of each face sub-region, the matching window size of each face sub-region is dynamically generated in combination with the face depth characteristics; finally, the disparity matching is performed on each face sub-region in the face speckle pattern and the face reference image according to the generated matching window to generate the disparity value and the similarity array of each pixel in the face sub-region, and then the similarity array is sorted according to the size to obtain the disparity value under the maximum similarity value, so as to calculate the actual depth value corresponding to each pixel point according to the disparity value and the theoretical formula.

[0141] It is worth mentioning that, as Figure 15 shown in FIG. 2, the processing device 20 of the face depth image according to one embodiment of the present application can include the following components which are communicatively connected to each other:

[0142] The acquisition module 21 is configured to acquire the gray image and the speckle image.

[0143] The detection module 22 is configured to perform face detection on the gray image to obtain the face region position and multiple face sub-regions in the gray image.

[0144] The generation module 23 is configured to dynamically generate the matching window corresponding to each face sub-region according to the characteristics of each face sub-region in the gray image.

[0145] The operation module 24 is configured to crop the image corresponding to the face region from the speckle image according to the face region position to obtain the face region image to be matched.

[0146] The matching module 25 is configured to perform disparity matching on the corresponding face sub-region in the face region image to be matched through a matching window corresponding to each face sub-region, so as to obtain the actual disparity value of each pixel in the face region image, and further obtain the face depth image.

[0147] In one example of the present application, as shown in Figure 15 The obtaining module 21 can include an image obtaining module 211 and an image calling module 212 which are communicatively connected, the image obtaining module 211 is configured to obtain a scene speckle image and a grayscale image of a target scene captured by a monocular structured light camera; and the image calling module 212 is configured to call a plurality of reference speckle images from a memory, wherein the reference speckle images are obtained by shooting a plane at different distances through the monocular structured light camera.

[0148] It is worth noting that in other examples of the present application, the obtaining module 21 can also only include the image obtaining module 211, which is configured to obtain a binocular grayscale image and a binocular speckle image of a target scene captured by a binocular structured light camera, wherein when performing face detection on a left-eye grayscale image in the binocular grayscale image, a left-eye speckle image in the binocular speckle image is used as a scene speckle image, and a right-eye speckle image in the binocular speckle image is used as a reference speckle image.

[0149] In one example of the present application, as shown in Figure 15 The detection module 22 can include a region detection module 221, a key point detection module 222 and a region division module 223 which are communicatively connected, the region detection module 221 is configured to perform face region detection on the grayscale image through a deep learning model to obtain the position of the face region in the target scene; the key point detection module 222 is configured to perform key point detection on the face region in the grayscale image to obtain the position of the face key point in the grayscale image; and the region division module 223 is configured to divide the face region in the grayscale image into a plurality of face sub-regions according to the position of the face key point.

[0150] In one example of the present application, as shown in Figure 15 The generation module 23 can include a shape selection module 231 and a size adjustment module 232 which are communicatively connected, the shape selection module 231 is configured to dynamically select a matching window with a corresponding shape according to the face depth characteristics; and the size adjustment module 232 is configured to inversely adjust the size of the matching window according to the distance of the face region.

[0151] It is worth noting that the shape selection module 231 can be further configured to: select a square matching window according to the feature that the depth of the forehead region is relatively flat; select a horizontal long rectangular matching window according to the feature that the depth of the eye region and the mouth region is relatively long in the horizontal direction; and select a vertical long rectangular matching window according to the feature that the depth of the nose region and the cheek region is relatively long in the vertical direction.

[0152] In addition, the size adjustment module 232 can be further configured to: adjust the size of the matching window to be smaller when the distance of the face region becomes larger; and adjust the size of the matching window to be larger when the distance of the face region becomes smaller.

[0153] In an example of the present application, as shown in Figure 15 The operation module 24 can include a cropping module 241 and a super-resolution module 242 which are communicatively connected to each other. The cropping module 241 is configured to crop a face region speckle image corresponding to the face region from the scene speckle image according to the face region position, and crop face region reference images corresponding to the face region from the plurality of reference speckle images according to the face region position and the reserved disparity search range, respectively. The super-resolution module 242 is configured to perform super-resolution operation on the face region speckle image and the plurality of face region reference images to obtain a face region speckle image to be matched and a plurality of face region reference images to be matched as the face region image to be matched.

[0154] It is worth noting that in other examples of the present application, the operation module 24 can include a cropping module 241 and a super-resolution module 242 which are communicatively connected to each other. The cropping module 241 is configured to crop a face speckle image corresponding to the face region from the left eye speckle image according to the face region position, and crop face reference images corresponding to the face region from the right eye speckle image according to the face region position and the reserved disparity search range, respectively. The super-resolution module 242 is configured to perform super-resolution operation on the face speckle image and the face reference images to obtain a face speckle image to be matched and a face reference image to be matched as the face region image to be matched.

[0155] In an example of the present application, as shown in Figure 15As shown, the matching module 25 may include a translation module 251, a sorting module 252, and a calculation module 253 that are communicatively connected to each other. The translation module 251 is used to perform left and right translations within a predetermined disparity range on the multiple face reference images to be matched, according to the generated matching window, for each pixel of the corresponding face sub-region in the face speckle image to be matched, so as to obtain the disparity value and similarity array of each pixel of the face sub-region through disparity matching. The sorting module 252 is used to sort the similarity array by size, so as to take the disparity value corresponding to the maximum similarity value as the actual disparity value. The calculation module 253 is used to calculate the actual depth value corresponding to each pixel based on the actual disparity value of each pixel in the face region image, so as to obtain the face depth image.

[0156] It is worth mentioning that, attached Figure 16 A face recognition device 30 according to an embodiment of this application is shown, which may include a structured light camera device 10 and a data processing module 31 that are communicatively connected to each other. The structured light camera device 10 is used to acquire image data input to the data processing module 31, and the data processing module 31 is used to perform the steps in the above-described face depth image processing method based on the image data to perform depth calculation.

[0157] Indicative electronic products

[0158] Below, for reference Figure 17 To describe the electronic device according to embodiments of the present invention ( Figure 9 A block diagram of an electronic device according to an embodiment of the present invention is shown. Figure 17 As shown, the electronic device 40 includes one or more processors 41 and storage medium 42.

[0159] The processor 41 may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 40 to perform desired functions.

[0160] The storage medium 42 may include one or more computing program products, which may include various forms of computing-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computing program instructions may be stored on the computing-readable storage medium, and the processor 41 may execute the program instructions to implement the methods of the various embodiments of the present invention described above and / or other desired functions.

[0161] In one example, as shown in Figure 17 The electronic device 40 can further include an input device 43 and an output device 44, which are interconnected via a bus system and / or other form of connection mechanism (not shown). For example, the input device 43 can be a sensor or the like for monitoring hardware parameters of the plurality of display devices.

[0162] The output device 44 can output various information externally, including a display screen and the like. The output device 44 can include a display and a communication network and a remote output device connected thereto, and the like.

[0163] Of course, in order to simplify, Figure 17 Only some of the components of the electronic device 40 related to the present application are shown in the above-described "example method" section, and components such as buses, input / output interfaces, and the like are omitted. In addition to this, the electronic device 40 can include any other appropriate components according to the specific application.

[0164] Illustrative computing program product

[0165] In addition to the above-described methods and devices, embodiments of the present application can also be a computing program product, such as a storage medium, which includes computing program instructions that, when executed by a processor, cause the processor to perform the steps of the methods according to various embodiments of the present application described in the above "example method" section of the specification.

[0166] The computing program product can be written in any combination of one or more programming languages, including an object-oriented programming language such as Java, C++, and the like, and conventional procedural programming languages, such as the C programming language or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0167] In addition, embodiments of the present application can also be a computing readable storage medium, which stores computing program instructions, which, when executed by a processor, cause the processor to perform the steps of the above-described methods.

[0168] The computer readable storage medium can employ any type of suitable media to store the data and / or instructions for use by a computer. Examples of suitable media include, but are not limited to, RAM, read-only memory (ROM), tape, magnetic disk storage or other memory on hard disk drives, optical disk storage, etc. The computer readable storage medium can be distributed among computer systems connected through a network and data and / or instructions stored by any one computer system can be shared with other computer systems as appropriate for the load balancing.

[0169] The technical features of the above embodiments can be combined in any manner. For the sake of brevity, not all possible combinations of the technical features of the above embodiments are described, however, any combination of the technical features should be considered as within the scope of the present disclosure as long as the combination does not result in a contradiction.

[0170] The above embodiments merely express several implementation manners of the present application, and the description is relatively specific and detailed, but should not be understood as a limitation on the scope of the patent. It should be noted that, for those skilled in the art, several modifications and improvements can be made without departing from the concept of the present application, and these are within the scope of the present application. Therefore, the scope of the patent of the present application should be subject to the appended claims.

Claims

1. A method of processing a human face depth image, characterized in that, The method comprises the steps of: obtaining a grayscale image and a speckle image; performing face detection on the grayscale image to obtain a face region position and a plurality of face sub-regions in the grayscale image; dynamically generating a matching window corresponding to each face sub-region according to the characteristics of each face sub-region in the grayscale image; cropping an image corresponding to the face region from the speckle image according to the face region position to obtain a face region image to be matched; and performing disparity matching on the corresponding face sub-region on the face region image to be matched through the matching window corresponding to each face sub-region to obtain the actual disparity value of each pixel on the face region image, and further obtain a face depth image. The step of dynamically generating a matching window corresponding to each face sub-region according to the characteristics of each face sub-region in the grayscale image comprises the steps of: dynamically selecting a matching window of a corresponding shape according to the depth characteristics of the face; and adjusting the size of the matching window reversely according to the distance of the face region.

2. The method of claim 1, wherein, The step of obtaining a grayscale image and a speckle image comprises the steps of: obtaining a scene speckle image and a grayscale image of a target scene collected by a monocular structured light camera; and calling a plurality of reference speckle images from a memory, wherein the reference speckle images are obtained by shooting a plane at different distances through the monocular structured light camera.

3. The method of claim 2, wherein, The step of cropping an image corresponding to the face region from the speckle image according to the face region position to obtain a face region image to be matched comprises the steps of: cropping a face speckle image corresponding to the face region from the scene speckle image according to the face region position; cropping face reference images corresponding to the face region from the plurality of reference speckle images respectively according to the face region position and a reserved disparity search range; and performing super-resolution operation on the face speckle image and the plurality of face reference images respectively to obtain a face speckle image to be matched and a plurality of face reference images to be matched as the face region image to be matched. The step of obtaining a grayscale image and a speckle image comprises the steps of:

4. The method of claim 1, wherein, obtaining a binocular grayscale image and a binocular speckle image of a target scene collected by a binocular structured light camera, wherein when performing face detection on a left grayscale image in the binocular grayscale image, a left speckle image in the binocular speckle image is taken as a scene speckle image, and a right speckle image in the binocular speckle image is taken as a reference speckle image. The step of cropping an image corresponding to the face region from the speckle image according to the face region position to obtain a face region image to be matched comprises the steps of:

5. The method of claim 4, wherein, cropping a face speckle image corresponding to the face region from the left speckle image according to the face region position; cropping face reference images corresponding to the face region from the right speckle image respectively according to the face region position and a reserved disparity search range; and performing super-resolution operation on the face speckle image and the face reference image to obtain a face speckle image to be matched and a face reference image to be matched as the face region image to be matched. ​ ​ 6. The method of claim 1, wherein, The step of performing face detection on the grayscale image to obtain a face region position and a plurality of face sub-regions in the grayscale image comprises the steps of: performing face region detection on the grayscale image by a deep learning model to obtain a face region position in a target scene; performing key point detection on the face region in the grayscale image to obtain a face key point position in the grayscale image; and dividing the face in the grayscale image into a plurality of face sub-regions according to the face key point position. In the step of dividing the face in the grayscale image into a plurality of face sub-regions according to the face key point position, the plurality of face sub-regions comprises a forehead region, an eye region, a nose region, a mouth region, and a cheek region.

7. The method of claim 6, wherein, The step of dynamically selecting a matching window of a corresponding shape according to a face depth feature comprises the steps of:

8. The method of claim 1 to 7, wherein, selecting a square matching window according to the feature that the depth of the forehead region is relatively flat; selecting a horizontal long rectangular matching window according to the feature that the depth of the eye region and the mouth region is relatively long in the horizontal direction; and selecting a vertical long rectangular matching window according to the feature that the depth of the nose region and the cheek region is relatively long in the vertical direction. The step of performing disparity matching on a corresponding face sub-region in the face region image to be matched by a matching window corresponding to each face sub-region to obtain an actual disparity value of each pixel in the face region image, and then obtaining a face depth image comprises the steps of: 9.The method of claim 3 or 5, wherein, performing left-right translation within a predetermined disparity range on each pixel in the corresponding face sub-region in the face speckle image to be matched in the face reference image to be matched respectively to obtain a disparity value and a similarity array of each pixel in the face sub-region by disparity matching according to the generated matching window; performing size sorting on the similarity array to take the disparity value corresponding to the maximum similarity value as the actual disparity value of the corresponding pixel; and calculating an actual depth value corresponding to each pixel according to the actual disparity value of each pixel in the face region image to obtain the face depth image. The system comprises the following which are communicatively connected to each other: an acquisition module configured to acquire a grayscale image and a speckle image; 10. An apparatus for processing a human face depth image, characterized by a detection module configured to perform face detection on the grayscale image to obtain a face region position and a plurality of face sub-regions in the grayscale image; a generation module configured to dynamically generate a matching window corresponding to each face sub-region according to a feature of each face sub-region in the grayscale image; an operation module configured to crop an image corresponding to the face region from the speckle image according to the face region position to obtain a face region image to be matched; and a matching module configured to perform disparity matching on a corresponding face sub-region in the face region image to be matched by a matching window corresponding to each face sub-region to obtain an actual disparity value of each pixel in the face region image, and then obtain a face depth image. ​ ​ ​ The generating module comprises a shape selecting module and a size adjusting module which are communicatively connected, the shape selecting module is used for dynamically selecting a matching window with a corresponding shape according to the face depth feature; and the size adjusting module is used for reversely adjusting the size of the matching window according to the distance of the face region.

11. A face recognition apparatus characterized by comprising: Comprise: a structured light camera module, configured to acquire image data; and a data processing module, communicatively connected to the structured light camera module, configured to perform the steps of the face depth image processing method according to any one of claims 1 to 9 based on the image data from the structured light camera module to perform depth calculation.

12. An electronic device, characterized by Comprise: a processor; and a storage medium, in which a computing program instruction is stored, the computing program instruction, when executed by the processor, causes the processor to perform the face depth image processing method according to any one of claims 1 to 9.

13. Storage medium, characterized in that The storage medium has a computing program instruction stored thereon, and the computing program instruction is operable to perform the face depth image processing method according to any one of claims 1 to 9 when executed by a computing device.

Citation Information

Patent Citations

  • Stereo vision matching method and system

    CN109658443A

  • Depth imaging method and device, electronic equipment and computer readable storage medium

    CN114387324A