A portrait recognition method, system, device and medium based on depth map self-adaption
Patent Information
- Application Number
- CN202310613442.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-29
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-05-29
AI Technical Summary
[0041]本发明利用第一预设距离参数参图像进行初步处理,大大减少了背景信息的干扰,减少了数据处理量,有利于提高检测精度,并提高效率。
Smart Images

Figure CN119068374B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of video processing technology, and more specifically, to a depth map-based adaptive human face recognition method, system, device, and medium. Background Technology
[0002] Image matting (background removal) refers to extracting the foreground image of interest from an image or video (image sequence) while filtering out the background. An image can be simply viewed as consisting of two parts: the foreground and the background. In short, image matting separates the foreground and background of a given image, and it is one of the key techniques in many image editing applications. Research on this technique has a history of over 20 years. The mathematical definition of image background removal was first proposed by Porter and Duff in 1984. They first introduced the concept of an alpha channel to control the linear interpolation ratio of the foreground and background colors when blending them. Ultimately, the image matting problem was defined as the task of estimating the alpha value (foreground-background color ratio) for each pixel in the image. In this task, the input is the original image, and the output is the alpha value of each pixel. Since the known information is the color information of each pixel, and the unknown information is the alpha, foreground color, and background color, the number of unknown variables is greater than the number of known variables, making it a mathematically challenging problem with insufficient constraints. Most extinction methods rely on user-provided guidance or prior assumptions to constrain the problem and obtain good estimates of unknown variables. Using guidance and prior assumptions inherently leads to weak generalization ability and poor prediction accuracy under specific conditions. Therefore, providing high-confidence prior information during the matting process can significantly improve its performance. A classic approach is interactive matting technology called Bayesian-based matting, where the user participates in the process, manually labeling the foreground and background to extract the foreground. In contrast, this invention does not rely on other information; the algorithm can autonomously determine and separate the foreground and background.
[0003] Furthermore, current common image matting algorithms, such as those provided by Tencent Meeting and Zoom, can only meet the needs of portrait matting, i.e., extracting the foreground of the portrait. Other foreground elements associated with the portrait, such as handheld objects, attached tables, and computers, cannot be fully extracted. For users who need to display items in live streaming or meetings, existing methods cannot meet their expectations. However, this invention can extract the foreground of the portrait or the portrait plus objects based on a depth map, according to the user's needs, while fully extracting the foreground of the portrait. The depth image, also known as a range image, refers to an image that uses the distance (depth) from the image acquisition device to various points in the scene as pixel values. It directly reflects the geometry of the visible surface of the scene.
[0004] The above background information is provided only to aid in understanding the inventive concept and technical solution of this invention. It does not necessarily belong to the prior art of this patent application. In the absence of clear evidence that the above information was disclosed on the filing date of this patent application, the above background information should not be used to evaluate the novelty and inventiveness of this application. Summary of the Invention
[0005] To this end, the present invention uses a first preset distance parameter to obtain a first human image range, and then uses the UNet network and Skip-Connection with different dimensional features, combined with temporal features, to identify a second human image. The second human image range is adaptively adjusted according to the depth distribution, which can completely and effectively identify human bodies and interactive objects, remove weakly related object information, and generate high-quality, high-resolution images, enabling real-time video processing.
[0006] In a first aspect, the present invention provides a face recognition method based on depth map adaptation, characterized by comprising the following steps:
[0007] Step S1: Acquire the first frame image from the video; wherein the first frame image includes a first RGB image and a first depth image;
[0008] Step S2: On the first depth image, the first portrait range is obtained using the first preset distance parameter;
[0009] Step S3: Based on the UNet network structure and Skip-Connection with different dimensional features, and combined with the temporal features of the first frame image, the first portrait range is identified to obtain the second portrait range;
[0010] Step S4: Calculate the depth value distribution range of the second human image range, and obtain the second preset distance parameter and the third preset distance parameter based on the confidence interval a;
[0011] Step S5: Adjust the second portrait range using the second preset distance parameter and the third preset distance parameter to obtain the third portrait range.
[0012] Optionally, the depth map-adaptive facial recognition method is characterized in that step S2 includes:
[0013] Step S21: Filter the first depth image to obtain the second depth image;
[0014] Step S22: Extract the first portrait range from the second depth image according to the first preset distance parameter;
[0015] Step S23: Perform an erosion operation on the first portrait area, and mark the eroded area as 1, the eroded area as 0.5, and other areas as 0.
[0016] Optionally, the depth map-based adaptive human face recognition method is characterized in that the erosion operation is determined according to the pixel offset between the second depth map and the first RGB image.
[0017] Optionally, the aforementioned depth map-based adaptive face recognition method is characterized in that step S3 includes:
[0018] Step S31: Use an encoder to extract features and condense semantic information for the first portrait area;
[0019] Step S32: Use a decoder to restore image pixels;
[0020] Step S33: Supplement information using Skip-Connections with different dimensions of features;
[0021] Step S34: Obtain the second portrait range by iterating through the video temporal correlation loss function.
[0022] Optionally, the depth map-adaptive human face recognition method is characterized in that, in step S32, the decoder includes a GRU for inter-frame information transmission.
[0023] Optionally, the aforementioned depth map-based adaptive face recognition method is characterized by further comprising:
[0024] Step S6: Identify the hair region within the third person's image range and adjust the hair region according to the temporal features.
[0025] Optionally, the depth map-adaptive facial recognition method is characterized in that step S6 includes:
[0026] Step S61: Identify the hair region within the third person's image area and classify the hair type;
[0027] Step S62: Select a motion trajectory for the hair according to the hair type;
[0028] Step S63: Adjust the hair region according to the motion trajectory and the temporal characteristics.
[0029] Secondly, the present invention provides a depth map-adaptive facial recognition system for implementing the depth map-adaptive facial recognition method described in any of the above claims, characterized in that it includes:
[0030] The acquisition module is used to acquire the first frame image in the video; wherein the first frame image includes a first RGB image and a first depth image;
[0031] The preprocessing module is used to obtain the first portrait range on the first depth image using a first preset distance parameter;
[0032] The image matting module is used to identify the first portrait range based on the UNet network structure and Skip-Connection with different dimensional features, combined with the temporal features of the first frame image, to obtain the second portrait range;
[0033] The confidence module is used to calculate the depth value distribution range of the second human image range and obtain the second preset distance parameter and the third preset distance parameter based on the confidence interval a;
[0034] The adjustment module is used to adjust the second portrait range using the second preset distance parameter and the third preset distance parameter to obtain the third portrait range.
[0035] Thirdly, the present invention provides a depth map-adaptive facial recognition device, characterized in that it includes:
[0036] processor;
[0037] A memory in which executable instructions of the processor are stored;
[0038] The processor is configured to perform the steps of any of the above-described depth map-adaptive facial recognition methods by executing the executable instructions.
[0039] Fourthly, the present invention provides a computer-readable storage medium for storing a program, characterized in that, when the program is executed, it implements the steps of the depth map-based adaptive human face recognition method described in any one of the preceding claims.
[0040] Compared with the prior art, the present invention has the following beneficial effects:
[0041] This invention utilizes a first preset distance parameter image for preliminary processing, which greatly reduces interference from background information, reduces the amount of data processing, and helps improve detection accuracy and efficiency.
[0042] This invention improves the stability of video matting by leveraging the temporal features of images and integrates human segmentation functionality. Through a multi-task training strategy, the matting algorithm achieves both refined foreground prediction and the integrity of the foreground and robustness of the matting. It can realize real-time foreground matting based on RGB and depth images, while simultaneously meeting the needs of matting both people and objects.
[0043] This invention analyzes the depth value distribution of the second human image range, adjusts the data using the confidence level of the depth values, and obtains the final cutout result. It effectively removes weakly correlated interactive objects and abnormal data interference such as flying points, thereby improving the cutout quality. Attached Figure Description
[0044] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort. Other features, objects, and advantages of the present invention will become more apparent by reading the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0045] Figure 1 This is a flowchart illustrating the steps of a depth map-adaptive facial recognition method in an embodiment of the present invention.
[0046] Figure 2 This is a schematic diagram of a depth value distribution range in an embodiment of the present invention;
[0047] Figure 3 This is a schematic diagram of another depth value distribution range in an embodiment of the present invention;
[0048] Figure 4 This is a flowchart illustrating one step in obtaining the first portrait range according to an embodiment of the present invention;
[0049] Figure 5 This is a schematic diagram of the human image range in an embodiment of the present invention;
[0050] Figure 6 This is a flowchart illustrating one step in obtaining the range of a second human image according to an embodiment of the present invention;
[0051] Figure 7 This is a flowchart illustrating the steps of another depth map-adaptive facial recognition method in an embodiment of the present invention.
[0052] Figure 8 This is a schematic diagram of the hair area in an embodiment of the present invention;
[0053] Figure 9 This is a flowchart illustrating the steps of adjusting a hair area in an embodiment of the present invention;
[0054] Figure 10 This is a schematic diagram of the structure of a depth map-adaptive facial recognition system according to an embodiment of the present invention;
[0055] Figure 11 This is a schematic diagram of the structure of a depth map-adaptive facial recognition device according to an embodiment of the present invention; and
[0056] Figure 12 This is a schematic diagram of the structure of a computer-readable storage medium in an embodiment of the present invention. Detailed Implementation
[0057] The present invention will now be described in detail with reference to specific embodiments. These embodiments will help those skilled in the art to further understand the present invention, but do not limit the invention in any way. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention. These all fall within the scope of protection of the present invention.
[0058] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0059] This invention provides a depth map-adaptive facial recognition method, which aims to solve the problems existing in the prior art.
[0060] The technical solutions of the present invention and how they solve the above-mentioned technical problems will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present invention will now be described with reference to the accompanying drawings.
[0061] This invention uses a first preset distance parameter to obtain a first human image range, and then uses a UNet network and Skip-Connections with different dimensional features, combined with temporal features, to identify a second human image. The second human image range is adaptively adjusted according to the depth distribution, which can completely and effectively identify human bodies and interactive objects, remove weakly related object information, and generate high-quality, high-resolution images, enabling real-time video processing.
[0062] Figure 1 This is a flowchart illustrating the steps of a depth map-adaptive facial recognition method in an embodiment of the present invention.
[0063] like Figure 1 As shown, the steps of a depth map-adaptive facial recognition method in this embodiment of the invention include:
[0064] Step S1: Obtain the first frame image from the video.
[0065] In this step, the first frame image includes a first RGB image and a first depth image. In some application scenarios, the first RGB image and the first depth image may be described as RGBD images, which serve the same purpose as in this embodiment. The first RGB image and the first depth image in this step are pixel-aligned; describing them separately is to more clearly illustrate the image data processing procedure. The first frame image is acquired by a depth camera. Subsequent steps are all performed on a mask to obtain the final portrait region range. Figure 5 'a' represents an RGB image. Figure 5 b is the depth image.
[0066] Step S2: On the first depth image, the first portrait range is obtained using the first preset distance parameter.
[0067] In this step, the first preset distance parameter is a constant used to extract the effective area. The range smaller than the first preset distance parameter constitutes the first portrait range. For example, if the first preset distance parameter is 1.2m, then all pixels smaller than 1.2m constitute the first portrait range. For the same video, the first preset distance parameter is a fixed value. For different videos, the first preset distance parameter can have different values. The setting of the first preset distance parameter can be selected based on the depth camera and the specific application scenario.
[0068] Step S3: Based on the UNet network structure and Skip-Connection with different dimensional features, and combined with the temporal features of the first frame image, the first portrait range is identified to obtain the second portrait range.
[0069] In this step, the UNet structure consists of three parts: an encoder, a decoder, and in-layer Skip-Connections. The encoder performs downsampling. For example, if the input image size is 572*572, the feature map size after two 3*3 convolutions is 568*568, and the output size after 22 max pooling is 284*284. The decoder performs upsampling. For example, the decoder consists of four sets of 2*2 transposed convolutions, 3*3 convolutions, and a ReLU activation function, with an additional 1*1 convolution added to the final output layer. Finally, the in-layer Skip-Connections fold the outputs of each downsampling layer and connect them to the upsampling layer for fusion. This step also incorporates the temporal features of the first frame to process the image, ensuring the stability of human image recognition across different frames. To some extent, temporal features can also be considered a form of deep supervision. The second human image identified in this step encompasses the entire human body, including the head, body, clothing, and objects that interact with the human body.
[0070] Step S4: Calculate the depth value distribution range of the second portrait range, and obtain the second preset distance parameter and the third preset distance parameter based on the confidence interval a.
[0071] In this step, the depth values are analyzed and processed to obtain the second and third preset distance parameters. For example... Figure 2 As shown, the depth range of the target object can be intuitively observed through the depth value distribution. Figure 2 A depth value of 0 is the result after resetting related flying points, noise, etc. Figure 2 yes Figure 5 Depth value distribution map in b.
[0072] Figure 3 yes Figure 5 A depth value distribution map in d. (Example) Figure 3As shown, after resetting the background area (depth value 1500-2500) to 0, the number of pixels with a depth value of 0 increases significantly, and the depth value range of the second portrait area becomes (300, 600). Due to the characteristics of portraits, the depth values exhibit a continuous distribution. The second preset distance parameter is located to the left of the depth value range of the second portrait area, and the third preset distance parameter is located to the right of the depth value range of the second portrait area. The two endpoints of confidence interval a are the second preset distance parameter and the third preset distance parameter, respectively. The confidence level of confidence interval a is not less than 90%. Confidence interval a is not fixed but is determined based on the continuity of the depth value distribution and the number of depth values. On the depth value distribution map, the confidence level of depth values closer to the edges is lower; the confidence level of discontinuous parts with depth values in the middle is even lower; the fewer pixels with the same depth value, the lower the confidence level. When determining confidence interval a, the continuity of the depth value distribution is maintained, and the confidence level is maximized as much as possible.
[0073] Step S5: Adjust the second portrait range using the second preset distance parameter and the third preset distance parameter to obtain the third portrait range.
[0074] In this step, unlike the filtering performed in step S2, a portion of the depth value range is reset to optimize the second portrait range and obtain the third portrait range. This step sets pixels with depth values at the second and third preset distance parameters to "display," i.e., sets them to 1 on the mask. The third portrait range is larger than the second portrait range.
[0075] Figure 4 This is a flowchart illustrating one step in obtaining the first portrait range according to an embodiment of the present invention. Figure 4 As shown, one step in obtaining the range of a first human image in an embodiment of the present invention includes:
[0076] Step S21: Filter the first depth image to obtain the second depth image.
[0077] In this step, the filtering process includes flicker removal, flypoint removal, and missing region filling. Figure 5 b is a TOF depth map obtained from a 3D camera. As can be seen from the image, the depth map contains flying spots and missing regions. Differences between different frames can cause flickering and other artifacts. This step, by processing a single frame and the images preceding it, improves the quality of the depth image while ensuring the integrity of the image content, resulting in more complete foreground extraction.
[0078] Step S22: Extract the first portrait range from the second depth image according to the first preset distance parameter.
[0079] In this step, the first preset distance parameter is a constant. By comparing the depth value of each pixel with the preset distance parameter, all pixels with depth values less than the preset distance parameter can be identified, forming the first portrait range. The first portrait range is the set of all pixels with depth values less than the preset distance parameter.
[0080] Step S23: Perform an erosion operation on the first portrait area, and mark the eroded area as 1, the eroded area as 0.5, and other areas as 0.
[0081] In this step, the image is labeled in three categories based on the first portrait area. An erosion operation is performed on the first portrait area to obtain the core region, which is labeled as 1, indicating a high degree of confidence in the foreground. The eroded portion of the first portrait area is labeled as 0.5, indicating a lower degree of confidence in the foreground. Other areas are labeled as 0, representing the background. The portion labeled 0.5 could be either foreground or background, requiring further determination in subsequent steps.
[0082] In some embodiments, the erosion operation is determined based on the pixel offset between the second depth map and the first RGB image. Unlike the erosion operation in the prior art that uses a convolutional template for a single RGB image, this embodiment uses the offset between the depth map and the RGB image for the erosion operation. By utilizing the depth change, the confidence level of the human figure can be further confirmed, which can greatly improve the range and accuracy of the erosion. This is of great significance for obtaining a definite human figure and can greatly improve the recognition effect.
[0083] In some embodiments, the region filtered in step S21 is also marked as 0.5. The region processed in step S21 is not the original data obtained, but a low-confidence region, and is also marked as 0.5. Subsequent steps will then make further judgments on it to obtain more accurate results.
[0084] In this embodiment, the image is filtered, and then the range of the first portrait is extracted according to the preset distance parameter. Then, according to the erosion operation, the image area is marked as 1, 0.5 and 0 respectively. The depth value is used to determine part of the foreground and background, which reduces the amount of computation of subsequent algorithms, improves efficiency, and can avoid the determined foreground or background being misidentified by the algorithm, thus ensuring quality.
[0085] Figure 6 This is a flowchart illustrating one step in obtaining the range of a second human image according to an embodiment of the present invention. Figure 6 As shown, one step in obtaining the range of a second human image in an embodiment of the present invention includes:
[0086] Step S31: Use an encoder to extract features and condense semantic information for the first portrait area.
[0087] In this step, the encoder consists of a downsampling module composed of two 3x3 convolutional layers (RELU) and a 2x2 maxpooling layer. The low-resolution information obtained after multiple downsampling steps yields human portrait features and provides contextual semantic information for segmenting the human portrait within the entire image.
[0088] Step S32: Use a decoder to restore image pixels.
[0089] In this step, the decoder consists of an upsampled convolutional layer (deconvolutional layer), a feature concatenation layer, and two 3x3 convolutional layers (ReLU).
[0090] Step S33: Use Skip-Connections with different dimensional features to supplement information.
[0091] In this step, a U-Net with a skipconnection structure is added, enabling the network to fuse feature maps at corresponding encoder positions across channels during upsampling at each level. By fusing low-level and high-level features, the network retains more high-resolution detail information contained in the high-level feature maps, thereby improving image segmentation accuracy. While enhancing feature information transfer between layers through cross-layer feature reconstruction, the rich detail information in high-level convolutional feature layers is further utilized, maximizing the utilization rate of feature information in each layer of the network. This step introduces feature information at corresponding scales into the upsampling or deconvolution process, providing multi-scale and multi-level information for subsequent image segmentation, resulting in more refined segmentation effects.
[0092] Step S34: Obtain the second portrait range by iterating through the video temporal correlation loss function.
[0093] In this step, information from the previously processed multiple images is compared with the first frame image, and iterative optimization is performed using a loss function to obtain the final second portrait range.
[0094] In some embodiments, the decoder includes a GRU for inter-frame information propagation. The GRU takes three-dimensional variables as input, and the input variables of the conv layers change, making it well-suited for inter-frame information propagation. Furthermore, by predicting more consistent results, this significantly reduces artifacts and improves perceptual quality. Simultaneously, the GRU also improves extinction robustness, better guessing boundaries even when individual frames may be blurry.
[0095] This embodiment uses UNet to process the first portrait range and combines it with Skip-Connection for supplementation. It uses the video temporal correlation loss function for iterative optimization, so that the final second portrait range can obtain a high-accuracy and high-precision image on a single frame, and can ensure the continuity of the portrait object in the video, preventing problems such as video jitter.
[0096] In some embodiments, step S3 uses a pre-trained model, and the pre-trained model is trained by alternating between single-frame images and continuous video. The input for training is the image and video processed in step S2.
[0097] Figure 7 This is a flowchart illustrating the steps of another depth map-based adaptive facial recognition method in an embodiment of the present invention. Figure 7 As shown, compared to the aforementioned embodiments, another depth map-adaptive facial recognition method in this embodiment of the invention further includes:
[0098] Step S6: Identify the hair region within the third person's image range and adjust the hair region according to the temporal features.
[0099] In this step, the identification of the hair region within the third person's image area can be performed simultaneously with the aforementioned steps, or it can be performed separately in this step. When the hair region is identified in the aforementioned steps, the same method is used, but this step does not identify the hair region separately. When the hair region is identified in this step, the same model as in the aforementioned steps can be used, or a dedicated hair region identification model can be used; this embodiment does not impose any restrictions on this. Various schemes can be used when adjusting the hair region based on temporal features. For example, such as... Figure 8 As shown, the hair area is divided into fixed and moving areas. The fixed area remains stationary when adjusting the hair. The moving area is the area that needs to be repositioned. For example... Figure 8 As shown, the position of the motion area can be changed from position 1 to position 2, thereby adjusting the hair area. Different effects can be achieved by adjusting the hair area, such as increasing dynamism through hair movement or overcoming environmental interference by keeping the hair within a certain range. If this embodiment is combined with a background replacement, the hair movement can be adapted to the background, achieving a better blend between the foreground and the replaced background.
[0100] This embodiment identifies and adjusts the hair area of a person, enabling the hair to achieve the desired swaying effect or remain still. This allows for better application in scenarios such as background replacement, making the human body and background more natural after background replacement. It avoids the conflict between a still background where the human body's hair is moving, or a gentle breeze where the human body's hair remains still.
[0101] Figure 9 This is a flowchart illustrating the steps of adjusting a hair area in an embodiment of the present invention;
[0102] Step S61: Identify the hair area within the third person's image range and classify the hair type.
[0103] In this step, the hair is categorized to allow for different adjustment strategies in subsequent steps. For example, hair can be divided into long, medium, and short hair to differentiate the extent of adjustment required. Long hair, which falls more downwards (e.g., past the shoulders), allows for the greatest adjustment. Medium hair, which falls less downwards (e.g., past the lower earlobe but not past the shoulders), also allows for a relatively large adjustment. Short hair, which falls very little (e.g., not past the lower earlobe), is minimally affected by wind and requires only very small adjustments.
[0104] Step S62: Select a motion trajectory for the hair according to the hair type.
[0105] In this step, different hair types correspond to different movement trajectories. For example, for short hair, only the outermost part of the hair undergoes a small positional change; for medium-length hair, most of the hair is adjusted overall; and for long hair, most of the hair is divided into three sections for positional and curvature adjustments to better simulate the natural state of the hair.
[0106] Step S63: Adjust the hair region according to the motion trajectory and the temporal characteristics.
[0107] In this step, the motion trajectory varies depending on the hair type. The motion trajectory is closed, such as a circle, ellipse, or figure-eight. Specific points on the hair move along the motion trajectory. There can be multiple motion trajectories, with different trajectories set for different parts of the hair. If only specific points on a portion of the hair are set with motion trajectories, during adjustment, the specific points are first placed in the preset position, and then other points are adjusted accordingly to ensure the integrity of the hair. The position of the hair in the current frame image is determined by the motion trajectory and temporal features, thus determining the position of the hair after the change.
[0108] This embodiment classifies hair types, selects corresponding motion trajectories, and adjusts the hair region based on temporal features, enabling the hair to move in a manner suitable for its type. Furthermore, this embodiment can control the adjustment of the hair region by specifying motion trajectories, thereby achieving various hair movements. This is particularly suitable for scenarios involving background replacement, ensuring the hair movement adapts to the background and making the video more realistic after background replacement.
[0109] Figure 10This is a schematic diagram of the structure of a depth map-adaptive facial recognition system according to an embodiment of the present invention.
[0110] like Figure 10 As shown, an embodiment of the present invention provides a depth map-adaptive facial recognition system comprising:
[0111] The acquisition module is used to acquire the first frame image in the video; wherein the first frame image includes a first RGB image and a first depth image;
[0112] The preprocessing module is used to obtain the first portrait range on the first depth image using a first preset distance parameter;
[0113] The image matting module is used to identify the first portrait range based on the UNet network structure and Skip-Connection with different dimensional features, combined with the temporal features of the first frame image, to obtain the second portrait range;
[0114] The confidence module is used to calculate the depth value distribution range of the second human image range and obtain the second preset distance parameter and the third preset distance parameter based on the confidence interval a;
[0115] The adjustment module is used to adjust the second portrait range using the second preset distance parameter and the third preset distance parameter to obtain the third portrait range.
[0116] Specifically, the acquisition module obtains the first frame image with depth information, which is then processed on a mask by subsequent modules to obtain the final portrait region. The preprocessing module directly uses the depth information to obtain the first portrait range. The matting module uses a UNet-based network structure within the first portrait range to perform portrait recognition while ensuring features of different dimensions and inter-frame information, obtaining the second portrait range. The confidence module analyzes the depth value distribution range to obtain a second and a third preset distance parameter. The adjustment module then uses the second and third preset distance parameters to fine-tune the second portrait range to ensure the range of the core foreground and background.
[0117] This embodiment uses a first preset distance parameter to obtain the first human image range, and then uses the UNet network and Skip-Connection with different dimensional features, combined with temporal features, to identify the second human image. The second human image range is adaptively adjusted according to the depth distribution, which can completely and effectively identify the human body and interactive objects, remove weakly related object information, and generate high-quality, high-resolution images, enabling real-time video processing.
[0118] This invention also provides a depth map-adaptive facial recognition device, including a processor and a memory storing executable instructions for the processor. The processor is configured to execute steps of a depth map-adaptive facial recognition method by executing the executable instructions.
[0119] As described above, this embodiment uses a first preset distance parameter to obtain the first human image range, and then uses the UNet network and Skip-Connection with different dimensional features, combined with temporal features, to identify the second human image. The second human image range is adaptively adjusted according to the depth distribution, which can completely and effectively identify the human body and interactive objects, remove weakly related object information, and generate high-quality, high-resolution images, enabling real-time video processing.
[0120] Those skilled in the art will understand that various aspects of the present invention can be implemented as systems, methods, or program products. Therefore, various aspects of the present invention can be specifically implemented in the following forms: a completely hardware implementation, a completely software implementation (including firmware, microcode, etc.), or a combination of hardware and software aspects, collectively referred to herein as a "circuit," "module," or "platform."
[0121] Figure 11 This is a schematic diagram of a depth map-adaptive facial recognition device according to an embodiment of the present invention. The following refers to... Figure 11 To describe an electronic device 600 according to this embodiment of the present invention. Figure 11 The electronic device 600 shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of the present invention.
[0122] like Figure 11 As shown, the electronic device 600 is presented in the form of a general-purpose computing device. The components of the electronic device 600 may include, but are not limited to: at least one processing unit 610, at least one storage unit 620, a bus 630 connecting different platform components (including storage unit 620 and processing unit 610), a display unit 640, etc.
[0123] The storage unit stores program code, which can be executed by the processing unit 610 to perform the steps described in the section on a depth map-based adaptive face recognition method of this specification, according to various exemplary embodiments of the present invention. For example, the processing unit 610 can perform actions such as... Figure 1 The steps are shown in the figure.
[0124] Storage unit 620 may include a readable medium in the form of a volatile storage unit, such as random access memory (RAM) 6201 and / or cache memory 6202, and may further include a read-only memory (ROM) 6203.
[0125] Storage unit 620 may also include a program / utility 6204 having a set (at least one) program module 6205, such program module 6205 including but not limited to: operating system, one or more application programs, other program modules and program data, each or some combination of these examples may include an implementation of a network environment.
[0126] Bus 630 can represent one or more of several types of bus structures, including a memory cell bus or memory cell controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the various bus structures.
[0127] Electronic device 600 can also communicate with one or more external devices 700 (e.g., keyboard, pointing device, Bluetooth device, etc.), and with one or more devices that enable a user to interact with electronic device 600, and / or with any device that enables electronic device 600 to communicate with one or more other computing devices (e.g., router, modem, etc.). This communication can be performed via input / output (I / O) interface 650. Furthermore, electronic device 600 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 660. Network adapter 660 can communicate with other modules of electronic device 600 via bus 630. It should be understood that, although... Figure 11 As not shown in the diagram, other hardware and / or software modules may be used in conjunction with the electronic device 600, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage platforms.
[0128] This invention also provides a computer-readable storage medium for storing a program that, when executed, implements the steps of a depth map-based adaptive facial recognition method. In some possible implementations, various aspects of the invention can also be implemented as a program product comprising program code that, when run on a terminal device, causes the terminal device to perform the steps described in the foregoing section on a depth map-based adaptive facial recognition method according to various exemplary embodiments of the invention.
[0129] As shown above, this embodiment uses a first preset distance parameter to obtain the first human image range, and then uses the UNet network and Skip-Connection with different dimensional features, combined with temporal features, to identify the second human image. The second human image range is adaptively adjusted according to the depth distribution, which can completely and effectively identify the human body and interactive objects, remove weakly related object information, and generate high-quality, high-resolution images, enabling real-time video processing.
[0130] Figure 12 This is a schematic diagram of the structure of a computer-readable storage medium according to an embodiment of the present invention. (Reference) Figure 12 As shown, a program product 800 for implementing the above-described method according to an embodiment of the present invention is described. This product may employ a portable compact disc read-only memory (CD-ROM) and include program code, and may run on a terminal device, such as a personal computer. However, the program product of the present invention is not limited thereto. In this document, the readable storage medium may be any tangible medium containing or storing a program that may be used by or in conjunction with an instruction execution system, apparatus, or device.
[0131] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0132] Computer-readable storage media may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable storage medium may also be any readable medium other than a readable storage medium that can transmit, propagate, or transfer a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.
[0133] Program code for performing the operations of this invention can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, and conventional procedural programming languages such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0134] This embodiment uses a first preset distance parameter to obtain the first human image range, and then uses the UNet network and Skip-Connection with different dimensional features, combined with temporal features, to identify the second human image. The second human image range is adaptively adjusted according to the depth distribution, which can completely and effectively identify the human body and interactive objects, remove weakly related object information, and generate high-quality, high-resolution images, enabling real-time video processing.
[0135] The various embodiments described in this specification are presented in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0136] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art can make various modifications or variations within the scope of the claims, which do not affect the essence of the present invention.
Claims
1. A depth map-adaptive facial recognition method, characterized in that, Includes the following steps: Step S1: Acquire the first frame image from the video; wherein the first frame image includes a first RGB image and a first depth image; Step S2: On the first depth image, the first portrait range is obtained using the first preset distance parameter; Step S3: Based on the UNet network structure and Skip-Connection with different dimensional features, and combined with the temporal features of the first frame image, the first portrait range is identified to obtain the second portrait range; Step S4: Calculate the depth value distribution range of the second human image range, and obtain the second preset distance parameter and the third preset distance parameter based on the confidence interval a; Step S5: Adjust the second portrait range using the second preset distance parameter and the third preset distance parameter to obtain the third portrait range; Step S2 includes: Step S21: Filter the first depth image to obtain the second depth image; Step S22: Extract the first portrait range from the second depth image according to the first preset distance parameter; Step S23: Perform an erosion operation on the first portrait area, and mark the eroded area as 1, the eroded area as 0.5, and other areas as 0; Step S3 includes: Step S31: Use an encoder to extract features and condense semantic information for the first portrait area; Step S32: Use a decoder to restore image pixels; Step S33: Supplement information using Skip-Connections with different dimensions of features; Step S34: Obtain the second portrait range by iteratively applying the video temporal correlation loss function; In step S32, the decoder includes a GRU for inter-frame information transmission.
2. The face recognition method based on depth map adaptation according to claim 1, characterized in that, The erosion operation is determined based on the pixel offset between the second depth map and the first RGB image.
3. The face recognition method based on depth map adaptation according to claim 1, characterized in that, Also includes: Step S6: Identify the hair region within the third person's image range and adjust the hair region according to the temporal features.
4. The face recognition method based on depth map adaptation according to claim 3, characterized in that, Step S6 includes: Step S61: Identify the hair region within the third person's image area and classify the hair type; Step S62: Select a motion trajectory for the hair according to the hair type; Step S63: Adjust the hair region according to the motion trajectory and the temporal characteristics.
5. A depth map-adaptive facial recognition system, used to implement the depth map-adaptive facial recognition method according to any one of claims 1 to 4, characterized in that, include: The acquisition module is used to acquire the first frame image in the video; wherein the first frame image includes a first RGB image and a first depth image; The preprocessing module is used to obtain the first portrait range on the first depth image using a first preset distance parameter; The image matting module is used to identify the first portrait range based on the UNet network structure and Skip-Connection with different dimensional features, combined with the temporal features of the first frame image, to obtain the second portrait range; The confidence module is used to calculate the depth value distribution range of the second human image range and obtain the second preset distance parameter and the third preset distance parameter based on the confidence interval a; The adjustment module is used to adjust the second portrait range using the second preset distance parameter and the third preset distance parameter to obtain the third portrait range; The preprocessing module includes the following components during processing: Step S21: Filter the first depth image to obtain the second depth image; Step S22: Extract the first portrait range from the second depth image according to the first preset distance parameter; Step S23: Perform an erosion operation on the first portrait area, and mark the eroded area as 1, the eroded area as 0.5, and other areas as 0; The image matting module includes the following during processing: Step S31: Use an encoder to extract features and condense semantic information for the first portrait area; Step S32: Use a decoder to restore image pixels; Step S33: Supplement information using Skip-Connections with different dimensions of features; Step S34: Obtain the second portrait range by iteratively applying the video temporal correlation loss function; In step S32, the decoder includes a GRU for inter-frame information transmission.
6. A facial recognition device based on depth map adaptation, characterized in that, include: processor; A memory in which executable instructions of the processor are stored; The processor is configured to perform the steps of the depth map-adaptive facial recognition method according to any one of claims 1 to 4 by executing the executable instructions.
7. A computer-readable storage medium for storing a program, characterized in that, When the program is executed, it implements the steps of the depth map-adaptive facial recognition method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Human image recognition method, system and equipment based on depth map double guidance and medium
CN119048946A