Human Body Depth Estimation Method, Device, Electronic Device and Storage Medium
Through three-dimensional human body modeling and compression output from multi-camera perspective, the hollow problem of weak texture area depth estimation is solved, and the accuracy and completeness are improved, while reducing the number of files and transmission space.
Patent Information
- Application Number
- CN202510459227.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-14
AI Technical Summary
The existing human depth estimation task is difficult to calculate accurate depth information when dealing with weak texture areas, resulting in voids on the depth map, and the final output result files are numerous and inconvenient to store and access.
Three-dimensional human body modeling is performed through images from multiple camera perspectives, and the three-dimensional human body model is constructed using the fast parameterized SMPL model, the pose reconstruction model is used for depth estimation, and the depth map is compressed and output.
It effectively solves the problem of depth recovery of weak texture areas, improves the accuracy and completeness of depth estimation, reduces the number of output files, and improves the transmission speed and information integrity.
Smart Images

Figure CN119991764B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and particularly to a human body depth estimation method, device, electronic device and storage medium. Background Art
[0002] In the fields of computer vision and 3D reconstruction, human body depth estimation is a crucial technology, which aims to recover the three-dimensional spatial information of the human body from two-dimensional images. This technology is widely applied in multiple fields such as augmented reality, virtual reality, and human-computer interaction, and is of great significance for improving the user experience and system performance.
[0003] Current human body depth estimation tasks mostly adopt algorithm frameworks such as Patchmatch to obtain dense depth maps from various perspectives. However, due to the lack of accurate initial values in such solutions, it is difficult to calculate accurate depth information for weak texture regions; that is, for human body parts lacking obvious texture features, such as regions wearing white or black clothing, since the image information in these regions is highly consistent or has extremely low contrast, and the initial values are not accurate, it is impossible to effectively calculate depth values in these weak texture regions. This phenomenon is manifested as holes in the depth map, seriously affecting the integrity and accuracy of human body depth estimation. In addition, the current human body depth estimation tasks finally output a large number of files, which are extremely inconvenient for storage and retrieval. Summary of the Invention
[0004] The present invention provides a human body depth estimation method, device, electronic device and storage medium, which are used to solve the deficiencies in the prior art in dealing with weak texture regions and data storage output in human body depth estimation tasks, improve the accuracy, integrity of human body depth estimation, and effectively simplify the output results for subsequent retrieval.
[0005] The present invention provides a human body depth estimation method, including:
[0006] Obtain human body images of the person to be estimated from multiple camera perspectives;
[0007] Perform three-dimensional human body modeling based on the human body images from the multiple camera perspectives to obtain a three-dimensional human body model corresponding to the human body images from the multiple camera perspectives;
[0008] Perform projection based on the three-dimensional human body model, and compress and output the depth maps of the person to be estimated from the multiple camera perspectives obtained by the projection.
[0009] According to the human body depth estimation method provided by the present invention, the performing three-dimensional human body modeling based on the human body images from the multiple camera perspectives to obtain a three-dimensional human body model corresponding to the human body images from the multiple camera perspectives includes:
[0010] Based on the human body images under the multiple camera views, apply a pose reconstruction model to perform 3D human body modeling to obtain a corresponding 3D human body model;
[0011] The pose reconstruction model is trained based on the sample human body images of the sample personnel under the multiple camera views, the sample pose information corresponding to the sample human body images under the multiple camera views, the predicted human body models corresponding to the sample human body images under the multiple camera views, and the projections of the predicted human body models under the multiple camera views.
[0012] According to a human body depth estimation method provided by the present invention, the pose reconstruction model is trained based on the following steps:
[0013] Perform human body detection and key point detection on the sample human body images under the multiple camera views to obtain the sample human body regions and the corresponding sample pose information in the sample human body images under the multiple camera views;
[0014] Perform 3D human body modeling based on an initial reconstruction model to obtain a predicted human body model corresponding to the sample human body regions under the multiple camera views, and determine the projections of the predicted human body model under the multiple camera views;
[0015] Determine a region reconstruction loss based on the consistency between the projections under the multiple camera views and the sample human body regions under the multiple camera views;
[0016] Determine a pose reconstruction loss based on the consistency between the sample pose information corresponding to the sample human body images under the multiple camera views and the predicted pose information corresponding to the projections under the multiple camera views;
[0017] Perform parameter iteration on the initial reconstruction model based on the region reconstruction loss and the pose reconstruction loss to obtain a pose reconstruction model.
[0018] According to a human body depth estimation method provided by the present invention, the sample pose information includes the coordinates and pose parameters of the human body joint points of the sample personnel in the corresponding sample human body images; the predicted pose information includes the coordinates and pose parameters of the human body joint points in the projections under the corresponding camera views;
[0019] The determining the pose reconstruction loss based on the consistency between the sample pose information corresponding to the sample human body images under the multiple camera views and the predicted pose information corresponding to the projections under the multiple camera views includes:
[0020] Determine a position reconstruction loss based on the consistency between the coordinates of the human body joint points in the sample human body images under the multiple camera views and the coordinates of the human body joint points in the projections under the multiple camera views;
[0021] and / or, determining a pose reconstruction loss based on the pose parameters of human body joints in the projections under the multiple camera views, or based on the consistency between the pose parameters of human body joints in the sample human body images under the multiple camera views and the pose parameters of human body joints in the projections under the multiple camera views;
[0022] Determining the pose reconstruction loss based on the position reconstruction loss and / or the pose reconstruction loss.
[0023] According to a human body depth estimation method provided by the present invention, the method of performing projection based on the three-dimensional human body model and compressing and outputting the depth maps of the person to be estimated under the multiple camera views obtained by the projection includes:
[0024] Performing projection based on the three-dimensional human body model to obtain the initial depth map and the initial normal map of the person to be estimated under the multiple camera views;
[0025] Optimizing the initial depth map and the initial normal map under any one camera view based on the human body image under any one camera view and the human body images under other camera views, as well as the camera pose matrix corresponding to any one camera view and the camera pose matrix corresponding to other camera views, to obtain the depth map and the normal map under any one camera view;
[0026] Compressing and outputting the depth maps and the normal maps of the person to be estimated under the multiple camera views.
[0027] According to a human body depth estimation method provided by the present invention, the method of optimizing the initial depth map and the initial normal map under any one camera view based on the human body image under any one camera view and the human body images under other camera views, as well as the camera pose matrix corresponding to any one camera view and the camera pose matrix corresponding to other camera views, to obtain the depth map and the normal map under any one camera view includes:
[0028] Determining the pixel information of a target point in the human body image under any one camera view, the pixel information of the projection point of the target point in the human body images under other camera views, and the included angle between the camera ray corresponding to other camera views and the normal of the projection point of the target point in the human body images under other camera views;
[0029] Based on the pixel information of the target point in the human body image under any one of the camera viewpoints, the pixel information of the projection point of the target point in the human body image under the other camera viewpoint, the included angle, and the camera pose matrix corresponding to any one of the camera viewpoints and the camera pose matrix corresponding to the other camera viewpoint, optimize the initial depth map and the initial normal map under any one of the camera viewpoints to obtain the depth map and the normal map under any one of the camera viewpoints.
[0030] According to a human body depth estimation method provided by the present invention, the compressing and outputting the depth maps and normal maps of the person to be estimated under the multiple camera viewpoints includes:
[0031] Determine the important human body region in the human body image under any one of the camera viewpoints;
[0032] Based on the human body image under any one of the camera viewpoints and the human body images under the other camera viewpoints, and the camera pose matrix corresponding to any one of the camera viewpoints and the camera pose matrix corresponding to the other camera viewpoint, optimize the depth map and the normal map corresponding to the important human body region under any one of the camera viewpoints to obtain the target depth map and the target normal map under any one of the camera viewpoints;
[0033] Store the target depth map, the target normal map, and the human body image under any one of the camera viewpoints by bytes, and compress and output the stored contents under the multiple camera viewpoints.
[0034] The present invention also provides a human body depth estimation device, including:
[0035] An acquisition unit, configured to acquire human body images of a person to be estimated under multiple camera viewpoints;
[0036] A modeling unit, configured to perform three-dimensional human body modeling based on the human body images under the multiple camera viewpoints to obtain a three-dimensional human body model corresponding to the human body images under the multiple camera viewpoints;
[0037] An estimation unit, configured to project based on the three-dimensional human body model and compress and output the depth maps of the person to be estimated under the multiple camera viewpoints obtained by the projection.
[0038] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and running on the processor, where when the processor executes the computer program, the human body depth estimation method as described in any one of the above is implemented.
[0039] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the human body depth estimation method as described in any one of the above is implemented.
[0040] The human body depth estimation method, device, electronic device and storage medium provided by the present invention perform three-dimensional human body modeling through human body images from multiple camera perspectives, so as to obtain a corresponding three-dimensional human body model in a fast parameterization manner. According to the characteristic that the constructed three-dimensional human body model will not have holes, it is projected onto each camera perspective to obtain a depth map, which well solves the problem in the traditional solution that the depth cannot be restored in the weak texture area, resulting in holes in the depth map, and can effectively solve the depth information in the weak texture area, thus ensuring the accuracy and integrity of human body depth estimation. Finally, the depth map is compressed and output, which not only reduces the number of output files, improves the transmission speed, reduces the occupation of transmission space, but also improves the integrity and orderliness of information, realizes the effective simplification of the output result, and is convenient for access. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0042] Figure 1 is a schematic flowchart of the human body depth estimation method provided by the present invention;
[0043] Figure 2 is an example diagram of the three-dimensional human body model provided by the present invention;
[0044] Figure 3 is an example diagram of the bone joint points provided by the present invention;
[0045] Figure 4 is an example diagram of the face feature points provided by the present invention;
[0046] Figure 5 is an example diagram of the sample human body area provided by the present invention;
[0047] Figure 6 is an example diagram of the depth map storage process provided by the present invention;
[0048] Figure 7 is a schematic structural diagram of the human body depth estimation device provided by the present invention;
[0049] Figure 8 is a schematic structural diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0050] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.
[0051] Currently, for the human body depth estimation task, algorithm frameworks such as Patchmatch are mostly used to obtain the dense depth maps from various perspectives. However, such solutions have certain limitations. Firstly, due to the lack of accurate initial values, it is impossible to calculate the depth for weakly textured regions, and there are often holes in the depth map. Secondly, in the final output results, there are a large number of files, the storage method is not concise enough, and it is inconvenient to access.
[0052] In response to this, the present invention provides a human body depth estimation method, aiming to solve the deficiencies in weakly textured regions, data storage and output, etc. during human body depth estimation, improve the accuracy and integrity of human body depth estimation, and effectively simplify the output results, reduce the number of files, improve the transmission speed, and facilitate subsequent access. Figure 1 is a schematic flowchart of the human body depth estimation method provided by the present invention. As Figure 1 shown, the method includes:
[0053] Step 110, obtaining human body images of the person to be estimated from multiple camera perspectives;
[0054] Step 120, performing three-dimensional human body modeling based on the human body images from multiple camera perspectives to obtain three-dimensional human body models corresponding to the human body images from multiple camera perspectives;
[0055] Step 130, performing projection based on the three-dimensional human body model and compressing and outputting the depth maps of the person to be estimated from multiple camera perspectives obtained by the projection.
[0056] Specifically, considering the deficiencies in the traditional human body depth estimation task in dealing with weakly textured regions, as well as in result storage and output; that is, the depth value cannot be effectively calculated for weakly textured regions, resulting in holes in the depth map, and the large number of final output files, which are extremely inconvenient for storage and retrieval. In the embodiments of the present invention, it is proposed that a three-dimensional human body model can be first established based on human body images from multiple camera viewpoints, so as to obtain a corresponding three-dimensional human body model through a fast parameterization method, and then human body depth estimation can be performed based on this three-dimensional human body model to obtain accurate depth maps from multiple camera viewpoints and compress and output them. In this way, not only can the number of output files be reduced, the transmission speed be increased, and the transmission space occupancy be reduced, but also, due to the characteristic that the three-dimensional human body model constructed by the fast parameterization method does not have holes, it can be projected onto each camera viewpoint to obtain the depth map, thus well solving the problem that the depth of weakly textured regions cannot be restored in the traditional solution, and effectively solving the depth information of weakly textured regions, thereby ensuring the accuracy and integrity of human body depth estimation.
[0057] In detail, during the actual application process, before performing human body depth estimation, it is necessary to first determine the two-dimensional images of the person to be estimated, that is, the human body images of the person to be estimated from multiple camera viewpoints, and these human body images can form a complete human body image sequence of the person to be estimated. Here, the human body images can be obtained by camera shooting, that is, the image of the person to be estimated captured by the camera when the camera is at a certain viewpoint is the human body image at that camera viewpoint; the human body images can also be directly searched / found / downloaded, and after obtaining the human body images, their corresponding camera viewpoints can be calibrated by manual, automated, etc. means.
[0058] Among them, the human body image sequence can be one or multiple. In the case of multiple, the persons to be estimated corresponding to different human body image sequences can be the same or different, that is, it can be the human body image sequences corresponding to the same person to be estimated in different states (posture, environment, location, etc.), or it can be the human body image sequences corresponding to different persons in the same or different states.
[0059] After obtaining the human body image sequence, in the embodiments of the present invention, three-dimensional human body modeling can be performed based on this human body image sequence to obtain a three-dimensional human body model corresponding to the human body image sequence. Considering the deficiencies in the traditional human body depth estimation in weak texture regions, in the embodiments of the present invention, when performing three-dimensional human body modeling, a fast parameterization method is adopted to construct a three-dimensional human body model, that is, SMPL (Skinned Multi-Person Linear Model). By virtue of the excellent characteristics of this model (no holes), the defects in the current solution in dealing with weak texture regions can be overcome, thereby significantly improving the integrity and accuracy of the depth map, and further facilitating subsequent use.
[0060] Specifically, here, based on the human body images of the person to be estimated included in the human body image sequence at multiple camera viewpoints, three-dimensional human body modeling can be performed to model a three-dimensional human body model that matches the human body images at each camera viewpoint in the human body image sequence. That is, by aligning the human body images at multiple camera viewpoints through a fast parameterization method, the three-dimensional human body model of the person to be estimated can be obtained. The SMPL constructed in this way can be aligned with all the human body images in the human body image sequence, that is, the projection of the SMPL at each camera viewpoint is basically consistent with the human body region in the human body image of the person to be estimated at each camera viewpoint. In this way, the accuracy of the modeling is ensured, providing an accurate basis for subsequent depth estimation.
[0061] It should be noted that the process of obtaining the three-dimensional human body model SMPL through three-dimensional human body modeling can be realized by a pre-trained network model dedicated to human body modeling, or can also be realized by other algorithms and frameworks, such as the Human Mesh Recovery end-to-end framework, the Multi-View Stereo series, etc. The embodiments of the present invention do not make specific limitations on this.
[0062] Furthermore, after obtaining the three-dimensional human body model corresponding to the human body image sequence through three-dimensional human body modeling, in the embodiments of the present invention, the human body depth can be estimated based on this three-dimensional human body model to obtain the depth maps of the person to be estimated at each camera viewpoint and perform the final output.
[0063] Specifically, unlike the method of randomly initializing the depth in the traditional solution (the initial value is not accurate or even completely wrong), in the embodiment of the present invention, the three-dimensional human body model obtained in the previous step is used for depth estimation to obtain an accurate initial value, based on which an accurate depth map under each camera perspective can be obtained. Here, specifically, the three-dimensional human body model can be projected to project the three-dimensional human body model SMPL to each camera perspective, and its projection under each camera perspective is used as the initial value, so as to ensure the accuracy of the initial value and obtain the initial depth map under each camera perspective; then, based on this initial depth map, multiple iterative optimizations can be performed to update the depth value to ensure precision and accuracy, and finally an accurate depth map of the person to be estimated under each camera perspective can be obtained.
[0064] After that, it can be output. Taking into account the defects in the result output of the current human body depth estimation task, in the embodiment of the present invention, when outputting the result, compression processing can be performed first to reduce the number of files and the transmission space occupied, and ensure the integrity of the information, improve the transmission efficiency, and facilitate subsequent calls.
[0065] Specifically, here the depth maps under multiple camera perspectives may be compressed. Specifically, when performing the compression processing, the information may be stored in units of camera perspectives, and then the overall stored information, including the depth maps, human body images, etc. under all camera perspectives, may be compressed, and the compressed information may be output. In this way, while ensuring the integrity and orderliness of the information, the number of files, transmission time and space occupancy may be greatly reduced, thereby improving the transmission efficiency, and enabling users / downstream tasks to obtain the depth information of the person to be estimated in the three-dimensional space more quickly, thereby achieving complete human depth estimation, and ensuring the integrity, accuracy and effectiveness of human depth estimation.
[0066] The human body depth estimation method provided by the present invention performs three-dimensional human body modeling through human body images under multiple camera perspectives, so as to obtain the corresponding three-dimensional human body model in a fast parameterized manner, and based on the characteristic that the three-dimensional human body model constructed in this way will not have holes, it is projected to each camera perspective to obtain a depth map, which well solves the problem that the depth of weak texture areas cannot be restored in traditional solutions, resulting in holes in the depth map, and can effectively solve the depth information of the weak texture areas, thereby ensuring the accuracy and completeness of the human body depth estimation, and finally compresses the depth map for output, which not only reduces the number of output files, improves the transmission speed, reduces the transmission space occupied, but also improves the integrity and orderliness of the information, and realizes the effective simplification of the output results for easy access.
[0067] Based on the above embodiment, step 120 includes:
[0068] Based on the human body images from multiple camera viewpoints, apply a pose reconstruction model to perform 3D human body modeling and obtain the corresponding 3D human body model;
[0069] The pose reconstruction model is trained based on the sample human body images of the sample personnel from multiple camera viewpoints, the sample pose information corresponding to the sample human body images from multiple camera viewpoints, the predicted human body models corresponding to the sample human body images from multiple camera viewpoints, and the projections of the predicted human body models from multiple camera viewpoints.
[0070] Specifically, the process of performing 3D human body modeling based on the human body images from multiple camera viewpoints can be achieved through the pose reconstruction model. That is, the pose reconstruction model can, in a fast parameterization manner, obtain information such as the pose and shape of the person to be estimated from the human body region of the 2D human body image, and perform 3D modeling based on this, thereby obtaining the corresponding 3D human body model.
[0071] In detail, here it can be to input the human body images from multiple camera viewpoints in the human body image sequence into the pose reconstruction model to perform 3D human body modeling through the model. Specifically, it can be to perform parametric modeling based on the information such as the pose and shape of the person to be estimated obtained from the input 2D human body image. Figure 2 This is an example diagram of the 3D human body model provided by the present invention. As Figure 2 shown, by establishing the human body joint points and surface topologies in the 3D space, an accurate 3D human body model can be obtained.
[0072] However, it is worth noting that, in order to ensure the accuracy of the 3D human body model obtained by modeling, in the embodiments of the present invention, before applying this pose reconstruction model to perform 3D human body modeling, it is also necessary to train it to improve the model performance, so that its performance in the actual 3D human body modeling task is more excellent and the 3D human body model obtained by modeling is more accurate. Here, specifically, it can be to apply the sample human body images with pose information from multiple camera viewpoints for model training to obtain the trained pose reconstruction model.
[0073] Specifically, when performing model training, in order to enable the model to update in the expected direction faster and converge faster, in calculating the loss, the embodiments of the present invention adopt a multi-angle calculation method to measure the loss of the model in the three-dimensional human body modeling task from multiple angles, so as to make it converge faster through multiple constraints, and the modeling is faster and more accurate in actual application. That is, according to the sample human body images of the sample person under multiple camera views and the corresponding sample pose information, as well as the predicted human body model corresponding to the sample human body images under these multiple camera views and the projections of the predicted human body model under multiple camera views, the loss of the model in the three-dimensional human body modeling task is measured from two levels of the human body region and pose. Based on this loss, the model is trained to obtain the finally trained model, that is, the pose reconstruction model. And the trained model can perform better in the actual three-dimensional human body modeling task and the modeling is more accurate.
[0074] Based on the above embodiments, the pose reconstruction model is trained based on the following steps:
[0075] Perform human body detection and key point detection on the sample human body images under multiple camera views to obtain the sample human body regions and the corresponding sample pose information in the sample human body images under multiple camera views;
[0076] Perform three-dimensional human body modeling based on the initial reconstruction model to obtain the predicted human body models corresponding to the sample human body regions under multiple camera views, and determine the projections of the predicted human body models under multiple camera views;
[0077] Based on the consistency between the projections under multiple camera views and the sample human body regions under multiple camera views, determine the region reconstruction loss;
[0078] Based on the consistency between the sample pose information corresponding to the sample human body images under multiple camera views and the predicted pose information corresponding to the projections under multiple camera views, determine the pose reconstruction loss;
[0079] Based on the region reconstruction loss and the pose reconstruction loss, perform parameter iteration on the initial reconstruction model to obtain the pose reconstruction model.
[0080] Specifically, the training process of the pose reconstruction model may include the following steps:
[0081] First, training samples need to be obtained, that is, the sample data required for model training, which may include the sample human body images of the sample person under multiple camera views, and the pose information corresponding to the sample human body images, that is, the sample pose information. The sample pose information here can be obtained by performing key point detection on the sample human body images.
[0082] However, it is worth noting that in the current human body depth estimation task, when reconstructing the human body depth, since the input two-dimensional image contains not only the human body but also the background, and the traditional depth estimation algorithm does not distinguish between them, it will reconstruct the depth of the background that is not needed.
[0083] In view of this, in the embodiments of the present invention, to avoid the problem of reconstructing the background depth during three-dimensional human body modeling, it is proposed that when training the pose reconstruction model used for three-dimensional human body modeling, the background information is directly removed, and the model is trained based on the human body region, so as to avoid the problem of the background environment being reconstructed from the root cause, and the distinction between the background and the human body is realized.
[0084] Specifically, after obtaining the sample human body images of the sample person from multiple camera views, key point detection and human body detection can be performed on them to obtain the sample human body regions in the sample human body images from multiple camera views, and the corresponding sample pose information of each sample human body region. Specifically, key point detection can be performed on each sample human body image here, including face feature point detection and bone joint point detection, to obtain the sample pose information, that is, the coordinates and pose parameters of the human body joint points; the human body joint points here include face feature points and bone joint points; at the same time, human body detection can be performed on each sample human body image to identify the region where the sample person is located and segment it, so as to obtain the sample human body regions in each sample human body image.
[0085] Figure 3 is an example diagram of the bone joint points provided by the present invention, as Figure 3 shown, when performing bone joint point detection on the sample human body image, pose estimation methods such as OpenPose, HRNet, etc. can be used to extract the bone joint points of the human body from the sample human body image, such as shoulders, elbows, knees, ankle bones, etc., to obtain the coordinates and pose parameters of the bone joint points. Figure 4 is an example diagram of the face feature points provided by the present invention, as Figure 4 shown, by performing face feature point detection on the sample human body image, 68 two-dimensional face feature points can be obtained. Figure 5 is an example diagram of the sample human body region provided by the present invention, as Figure 5 shown, the sample human body region can be extracted from each sample human body image by using the human body mask segmentation method.
[0086] Subsequently, the initial model during the training process, i.e., the initial reconstruction model, can be applied to perform 3D human body modeling. Based on the sample human body regions in the sample human body images from multiple camera views in a parametric representation manner, a 3D human body model of the sample person can be reconstructed, i.e., the predicted human body models corresponding to each sample human body region. Here, the initial model can be a model constructed based on a network model with parametric representation, and the specific architecture of the model can be adjusted according to the actual situation, requirements, etc., and the embodiments of the present invention do not make specific limitations on this.
[0087] After obtaining the predicted human body model, in the embodiments of the present invention, it is also necessary to determine the projections of the predicted human body model from multiple camera views, that is, the predicted human body model needs to be projected onto the camera views corresponding to each sample human body image, so as to obtain its projections from multiple camera views.
[0088] Then, based on the sample human body images and the corresponding sample pose information of the sample person from multiple camera views, as well as the predicted human body model and its projections from multiple camera views, the losses of the initial reconstruction model in the 3D human body modeling task can be measured from two levels of human body region and pose, so as to obtain the region reconstruction loss and the pose reconstruction loss.
[0089] Here, specifically, the region reconstruction loss can be determined based on the consistency between the projections of the predicted human body model from multiple camera views and the sample human body regions in the sample human body images from multiple camera views; that is, according to the similarity or difference degree between the projection of the predicted human body model obtained by modeling from the camera view and the sample human body region corresponding to the corresponding camera view, the loss of the model during human body region reconstruction is judged, and this loss is the region reconstruction loss.
[0090] At the same time, the pose reconstruction loss can be determined based on the consistency between the sample pose information corresponding to the sample human body images from multiple camera views and the predicted pose information corresponding to the projections of the predicted human body model from multiple camera views; that is, according to the similarity or difference degree between the predicted pose information (coordinates and pose parameters of human body joint points) corresponding to the projection of the predicted human body model from the camera view and the sample pose information corresponding to the corresponding camera view, the loss of the model during human body pose reconstruction is judged, and this loss is the pose reconstruction loss.
[0091] After that, based on the above region reconstruction loss and pose reconstruction loss, the loss of the model in the entire 3D human body modeling task can be determined, and the parameters of the initial reconstruction model can be iteratively adjusted according to this loss, so that the projections of the predicted human body model output by the model after parameter adjustment can be as close as possible to the real sample human body regions, and the predicted pose information can be as consistent as possible with the real sample pose information, and finally a trained pose reconstruction model can be obtained.
[0092] Based on the above embodiments, the sample pose information includes the coordinates and pose parameters of the human joint points of the sample person in the corresponding sample human body image; the predicted pose information includes the coordinates and pose parameters of the human joint points in the projection under the corresponding camera view;
[0093] Based on the consistency between the sample pose information corresponding to the sample human body images under multiple camera views and the predicted pose information corresponding to the projections under multiple camera views, determine the pose reconstruction loss, including:
[0094] Based on the consistency between the coordinates of the human joint points in the sample human body images under multiple camera views and the coordinates of the human joint points in the projections under multiple camera views, determine the position reconstruction loss;
[0095] And / or, based on the pose parameters of the human joint points in the projections under multiple camera views, or the consistency between the pose parameters of the human joint points in the sample human body images under multiple camera views and the pose parameters of the human joint points in the projections under multiple camera views, determine the pose reconstruction loss;
[0096] Based on the position reconstruction loss and / or the pose reconstruction loss, determine the pose reconstruction loss.
[0097] Specifically, the process of determining the pose reconstruction loss according to the consistency between the sample pose information corresponding to the sample human body images under multiple camera views and the predicted pose information corresponding to the projections under multiple camera views may specifically include:
[0098] First, the loss can be measured from the position level, that is, according to the consistency between the coordinates of the human joint points in the sample human body images under multiple camera views and the coordinates of the human joint points in the projections of the predicted human model under multiple camera views, determine the loss of the position coordinates during the reconstruction of the human joint points by the initial reconstruction model, that is, determine whether the position modeling of the human joint points is accurate and whether there are differences, so as to obtain the position reconstruction loss.
[0099] At the same time, the loss can also be measured from the pose level, that is, according to the consistency between the pose parameters of the human joint points in the sample human body images under multiple camera views and the pose parameters of the human joint points in the projections under multiple camera views, determine the loss of the human pose during the reconstruction of the human joint points by the initial reconstruction model, that is, determine whether the pose modeling of the human joint points is accurate and whether there are differences, so as to obtain the pose reconstruction loss.
[0100] Alternatively, the pose reconstruction loss can be determined only based on the pose parameters of the human body joints in the projections of the predicted human body model from multiple camera viewpoints. Specifically, here, it can be determined by evaluating the rationality, normality, coherence, etc. of the poses of the human body joints in the projections corresponding to the predicted human body model, and judging the loss of the human body pose in the reconstruction of the initial reconstruction model for the human body joints, that is, judging whether the pose modeling of the human body joints is reasonable, so as to obtain the pose reconstruction loss.
[0101] Then, based on the above position reconstruction loss and / or pose reconstruction loss, the pose and position reconstruction loss can be determined. Specifically, here, the position reconstruction loss can be directly used as the pose and position reconstruction loss, or the pose reconstruction loss can be used as the pose and position reconstruction loss, or the position reconstruction loss and the pose reconstruction loss can be combined to determine the pose and position reconstruction loss in a weighted manner.
[0102] Based on the above embodiments, the loss of the initial reconstruction model in the entire three-dimensional human body modeling task can be expressed by the following formula:
[0103]
[0104] In the formula, 、 and respectively represent the position reconstruction loss, the region reconstruction loss, and the pose reconstruction loss, 、 and represent the weights corresponding to the position reconstruction loss, the region reconstruction loss, and the pose reconstruction loss respectively.
[0105] Based on the above embodiments, step 130 includes:
[0106] Performing projection based on the three-dimensional human body model to obtain the initial depth map and the initial normal map of the person to be estimated from multiple camera viewpoints;
[0107] Based on the human body image from any camera viewpoint and the human body images from other camera viewpoints, as well as the camera pose matrix corresponding to this camera viewpoint and the camera pose matrices corresponding to other camera viewpoints, optimizing the initial depth map and the initial normal map from this camera viewpoint to obtain the depth map and the normal map from this camera viewpoint;
[0108] Compressing and outputting the depth maps and normal maps of the person to be estimated from multiple camera viewpoints.
[0109] Specifically, the process of performing projection based on the three-dimensional human body model and compressing and outputting the depth maps of the person to be estimated obtained from the projection from multiple camera viewpoints specifically includes:
[0110] In traditional depth estimation schemes, when performing depth estimation, random depths and normals in space are used as initial values to facilitate subsequent depth value optimization and calculation. However, this leads to slow depth estimation speed and low efficiency.
[0111] Based on this, in the embodiments of the present invention, an initial depth map is first obtained through a three-dimensional human body model, and then the initial depth map is updated and optimized by optimizing each pixel point one by one. Finally, a more accurate depth map can be obtained. Moreover, in this process, by effectively using depth and normal information for update and optimization, the iterative process can be made more efficient, converge faster, and the accuracy of the finally obtained depth map is higher.
[0112] Specifically, here, the three-dimensional human body model can be projected first to obtain the initial depth map and initial normal map of the person to be estimated under multiple camera views. Subsequently, this initial depth map and initial normal map can be optimized to obtain the depth map and normal map under multiple camera views; that is, the depth information and color information under the corresponding camera view are combined to jointly optimize the initial depth map and initial normal map to obtain the depth map and normal map under the corresponding camera view.
[0113] Specifically, for any camera view, when optimizing the initial depth map under it, the characteristic that the color information of the triangular patches on the three-dimensional human body model should be the same under different camera views is utilized, that is, the pixel values (color information) projected by the same triangular patch under different camera views should be consistent. The initial depth map and initial normal map are updated specifically by using the human body image under this camera view and the human body images under other camera views, as well as the camera pose matrix corresponding to this camera view and the camera pose matrices corresponding to other camera views, to update and optimize the depth map and normal map, and obtain the optimized depth map and normal map under this camera view. Through the above method, the optimized depth maps and normal maps under all camera views can be obtained.
[0114] After that, the optimized depth maps and normal maps of the person to be estimated under multiple camera views can be compressed and output. That is, the depth maps and normal maps can be stored in units of camera views, and then the overall stored information is compressed, and the compressed information is output. In this way, while ensuring the integrity and orderliness of the information, the number of files, transmission time, and space occupancy can be significantly reduced, thereby improving the transmission efficiency, and further enabling the user / downstream tasks to obtain the depth information of the person to be estimated in the three-dimensional space faster. In this way, a complete human body depth estimation is realized, ensuring the integrity, accuracy, and effectiveness of the human body depth estimation.
[0115] Based on the above embodiments, optimize the initial depth map and the initial normal map at this camera view based on the human body image at any camera view and the human body images at other camera views, as well as the camera pose matrix corresponding to this camera view and the camera pose matrices corresponding to other camera views, to obtain the depth map and the normal map at this camera view, including:
[0116] Determine the pixel information of the target point in the human body image at this camera view, the pixel information of the projection point of the target point in the human body images at other camera views, and the angle between the camera ray corresponding to the other camera view and the normal of the projection point of the target point in the human body image at the other camera view;
[0117] Based on the pixel information of the target point in the human body image at this camera view, the pixel information of the projection point of the target point in the human body images at other camera views, the angle, as well as the camera pose matrix corresponding to this camera view and the camera pose matrices corresponding to other camera views, optimize the initial depth map and the initial normal map at this camera view to obtain the depth map and the normal map at this camera view.
[0118] Specifically, when optimizing the initial depth map and the initial normal map at any camera view, according to the characteristics of the above three-dimensional human body model, that is, the pixels of the projection points of the same triangular patch at different camera views should be the same, the pixel information of the target point in the human body image at this camera view and the pixel information of the projection point of this target point in the human body images at other camera views can be obtained first. The target point here can be all points or some points in the human body area of the human body image. For different human body areas, such as important human body areas (such as the face) and regular human body areas (such as the limbs), the distribution and quantity of the target points can be the same or different.
[0119] At the same time, it is also necessary to determine the camera pose matrix corresponding to this camera view and the camera pose matrices corresponding to other camera views, as well as the angle between the camera ray corresponding to the other camera view and the normal of the projection point of the target point in the human body image at the other camera view.
[0120] After that, based on this angle, the camera pose matrix, and the pixel information of the target point and its corresponding projection point respectively, the initial depth map and the initial normal map at this camera view can be optimized, so as to obtain the optimized depth map and normal map at this camera view. Here, specifically, the camera pose matrix corresponding to this camera view and the camera pose matrices corresponding to other camera views can be used to project the target point at this camera view to other camera views. On this basis, combined with the pixel information of the target point at this camera view, the angle, and the pixel information of the projection point of the target point in the human body image at the other camera view, the initial depth map and the initial normal map at this camera view are optimized to obtain the high-precision depth map and normal map at this camera view.
[0121] Specifically, when the triangular patches of the 3D human body model SMPL are projected onto multiple camera views, the corresponding pixel regions of the triangular patches in the human body images under multiple camera views can be obtained. By subdividing the triangular patches into individual points, the initial depth values of each point in the initial depth map can be obtained. , denotes the point on this triangular patch at the th camera view. Since the pixel values projected from the same triangular patch under different camera views should be consistent, based on this, when optimizing and estimating the depth values of each camera view of the triangular patch, in order to ensure the accuracy of the projection, a weighting can be done using the direction of the light projection and the normal information of the point in the initial normal map. Assume the optimization of the initial depth map of the camera, and the total number of cameras is , then the optimization process can be expressed by the following formula:
[0122]
[0123] In the formula, denotes projecting the point at the camera view to the camera view. and respectively denote the camera pose matrix corresponding to the camera view and the camera pose matrix corresponding to the camera view. denotes the initial depth value of the point at the camera view. denotes the pixel information of the point at the camera view. denotes the angle between the camera ray corresponding to the camera view and the normal of the projection point of the point at the camera view. denotes the pixel information of the projection point of the point
[0124] Different from the traditional way of randomly initializing depth, in the embodiments of the present invention, a three-dimensional human body model is used to determine the initial depth map and the initial normal map under each camera view. On this basis, a depth and normal estimation method for each pixel is proposed for optimization, which has higher accuracy than the traditional method that only uses the path. Moreover, in the embodiments of the present invention, the normal information is reasonably utilized, making the iterative optimization process faster and finally obtaining a depth map with higher accuracy.
[0125] Based on the above embodiments, the depth map and the normal map of the person to be estimated under multiple camera views are compressed and output, including:
[0126] Determine the important human body regions in the human body image under any camera view;
[0127] Based on the human body image under this camera view and the human body images under other camera views, as well as the camera pose matrix corresponding to this camera view and the camera pose matrices corresponding to other camera views, optimize the depth map and the normal map corresponding to the important human body regions under this camera view to obtain the target depth map and the target normal map under this camera view;
[0128] Store the target depth map, the target normal map, and the human body image under this camera view by bytes, and compress and output the stored contents under multiple camera views.
[0129] Specifically, considering that in the traditional solution, the final output result is only a depth map, and in the result presentation, for important human body regions, such as the face, it is the same as conventional human body regions and no clear distinction is made, resulting in the result not being clear and concise enough and being inconvenient for subsequent use. In the embodiments of the present invention, it is proposed that regions with high detail requirements, such as the face, can be optimized multiple times to make them more accurate and more in line with the usage requirements.
[0130] In detail, for any camera view, here it is possible to first determine the important human body regions in the human body image under this camera view; the important human body regions here can be regions with relatively high accuracy requirements in the human body depth estimation task, for example, the head and face regions; it can also be regions specified by the user or regions of interest to the user, for example, the trunk, limb joints, etc., and the embodiments of the present invention do not make specific limitations on this. Then, use the human body image under this camera view and the human body images under other camera views, as well as the camera pose matrix corresponding to this camera view and the camera pose matrices corresponding to other camera views, to optimize the depth map and the normal map corresponding to the important human body regions under this camera view to obtain the target depth map and the target normal map under this camera view; the specific optimization process is basically the same as the optimization of the initial depth map described above and will not be elaborated here.
[0131] After that, the target depth map, target normal map, and human body image from the perspective of this camera can be stored byte by byte. Figure 6 is an example diagram of the depth map storage process provided by the present invention. As Figure 6 shown, information such as the target depth map, target normal map, displacement matrix, rotation matrix, and human body image can be stored according to the camera perspective, and the stored content from multiple camera perspectives can be compressed and output.
[0132] See Figure 6 In the embodiment of the present invention, the above information is stored and output in the ".bin format". Specifically, in the ".bin format" file, the rotation matrix R (3 × 3) and displacement matrix t (3 × 1) corresponding to the first camera perspective are saved first. Each value is 2 bytes, and a total of 24 (3 × 3 × 2 + 3 × 1 × 2) bytes. The subsequent bytes save the target depth map, target normal map, the number information of the affiliated triangular patches, color information (human body image), etc. By analogy, after storing the above information from all camera perspectives, the entire ".bin format" file is compressed, and the compressed file is output. In this way, not only can the integrity and orderliness of the information be ensured, but also the number of files and the transmission time are significantly reduced, the transmission efficiency is improved, and subsequent access is made more convenient.
[0133] Compared with the traditional depth estimation scheme based on image patch matching, in the embodiment of the present invention, through the per-pixel depth value optimization scheme, a higher-precision estimation of the depth of each pixel can be performed; and, according to actual needs, rapid re-optimization can be performed on important human body regions (which can be marked by the number information of triangular patches) to make them more accurate. This is the advantage of the present invention, and this characteristic cannot be achieved by traditional methods. In addition, a matching storage format is proposed in the embodiment of the present invention, which not only improves the storage efficiency of the depth map, but also enriches the stored content, saves all effective information, and is convenient for subsequent access.
[0134] Next, the human body depth estimation device provided by the present invention will be described. The human body depth estimation device described below can be mutually corresponded and referred to the human body depth estimation method described above.
[0135] Figure 7 is a schematic structural diagram of the human body depth estimation device provided by the present invention. As Figure 7 shown, the device includes:
[0136] An acquisition unit 710, configured to acquire human body images of the person to be estimated from multiple camera perspectives;
[0137] A modeling unit 720 is configured to perform three-dimensional human body modeling based on the human body images under the multiple camera perspectives to obtain three-dimensional human body models corresponding to the human body images under the multiple camera perspectives;
[0138] The estimation unit 730 is used to perform projection based on the three-dimensional human body model, and compress and output the depth map of the person to be estimated under the multiple camera perspectives obtained by the projection.
[0139] The human body depth estimation device provided by the present invention performs three-dimensional human body modeling through human body images under multiple camera perspectives to obtain the corresponding three-dimensional human body model in a fast parameterized manner. Based on the characteristic that the three-dimensional human body model constructed in this way will not have holes, it is projected to each camera perspective to obtain a depth map, which well solves the problem that the depth of weak texture areas cannot be restored in traditional solutions, resulting in holes in the depth map. The depth information of the weak texture areas can be effectively solved, thereby ensuring the accuracy and completeness of the human body depth estimation. Finally, the depth map is compressed and output, which not only reduces the number of output files, improves the transmission speed, reduces the transmission space occupied, but also improves the integrity and orderliness of the information, and realizes the effective simplification of the output results for easy access.
[0140] Based on the above embodiment, the modeling unit 720 is used to:
[0141] Based on the human body images under the multiple camera perspectives, applying the posture reconstruction model to perform three-dimensional human body modeling to obtain a corresponding three-dimensional human body model;
[0142] The posture reconstruction model is obtained based on sample body images of sample persons under multiple camera perspectives, sample pose information corresponding to the sample body images under the multiple camera perspectives, predicted body models corresponding to the sample body images under the multiple camera perspectives, and projection training of the predicted body model under the multiple camera perspectives.
[0143] Based on the above embodiment, the device further includes a training unit, which is used to:
[0144] Performing human body detection and key point detection on the sample human body images under the multiple camera perspectives to obtain sample human body regions and corresponding sample pose information in the sample human body images under the multiple camera perspectives;
[0145] Performing three-dimensional human body modeling based on the initial reconstruction model, obtaining a predicted human body model corresponding to the sample human body region under the multiple camera perspectives, and determining the projection of the predicted human body model under the multiple camera perspectives;
[0146] Determine a region reconstruction loss based on the consistency between the projections under the multiple camera views and the sample human body regions under the multiple camera views;
[0147] Determine a pose reconstruction loss based on the consistency between the sample pose information corresponding to the sample human body images under the multiple camera views and the predicted pose information corresponding to the projections under the multiple camera views;
[0148] Perform parameter iteration on the initial reconstruction model based on the region reconstruction loss and the pose reconstruction loss to obtain a pose reconstruction model.
[0149] Based on the above embodiments, the sample pose information includes the coordinates and pose parameters of the human body joints of the sample person in the corresponding sample human body image; the predicted pose information includes the coordinates and pose parameters of the human body joints in the projection under the corresponding camera view;
[0150] The training unit is configured to:
[0151] Determine a position reconstruction loss based on the consistency between the coordinates of the human body joints in the sample human body images under the multiple camera views and the coordinates of the human body joints in the projections under the multiple camera views;
[0152] And / or, determine a pose reconstruction loss based on the pose parameters of the human body joints in the projections under the multiple camera views, or the consistency between the pose parameters of the human body joints in the sample human body images under the multiple camera views and the pose parameters of the human body joints in the projections under the multiple camera views;
[0153] Determine the pose reconstruction loss based on the position reconstruction loss and / or the pose reconstruction loss.
[0154] Based on the above embodiments, the estimation unit 730 is configured to:
[0155] Perform projection based on the three-dimensional human body model to obtain the initial depth map and the initial normal map of the person to be estimated under the multiple camera views;
[0156] Optimize the initial depth map and the initial normal map under a camera view based on the human body image under any camera view and the human body images under other camera views, as well as the camera pose matrix corresponding to this camera view and the camera pose matrices corresponding to the other camera views, to obtain the depth map and the normal map under this camera view;
[0157] Compress and output the depth maps and normal maps of the person to be estimated under the multiple camera views.
[0158] Based on the above embodiments, the estimation unit 730 is configured to:
[0159] Determine the pixel information of the target point in the human body image from the perspective of this camera, the pixel information of the projection point of the target point in the human body image from the perspective of the other camera, and the angle between the camera light corresponding to the other camera perspective and the normal of the projection point of the target point in the human body image from the perspective of the other camera;
[0160] Based on the pixel information of the target point in the human body image from the perspective of this camera, the pixel information of the projection point of the target point in the human body image from the perspective of the other camera, the angle, as well as the camera pose matrix corresponding to this camera perspective and the camera pose matrix corresponding to the other camera perspective, optimize the initial depth map and the initial normal map from the perspective of this camera to obtain the depth map and the normal map from the perspective of this camera.
[0161] Based on the above embodiments, the estimation unit 730 is used for:
[0162] Determine the important human body regions in the human body image from any camera perspective;
[0163] Based on the human body image from the perspective of this camera and the human body images from the perspectives of other cameras, as well as the camera pose matrix corresponding to this camera perspective and the camera pose matrix corresponding to the other camera perspective, optimize the depth map and the normal map corresponding to the important human body regions from the perspective of this camera to obtain the target depth map and the target normal map from the perspective of this camera;
[0164] Store the target depth map, the target normal map, and the human body image from the perspective of this camera by bytes, and compress and output the stored contents from the perspectives of the multiple cameras.
[0165] Figure 8 Illustrates a schematic physical structure diagram of an electronic device, as Figure 8 shown. The electronic device may include: a processor 810, a communication interface 820, a memory 830, and a communication bus 840. Among them, the processor 810, the communication interface 820, and the memory 830 communicate with each other through the communication bus 840. The processor 810 can call the logical instructions in the memory 830 to execute the human body depth estimation method, which includes: obtaining the human body images of the person to be estimated from multiple camera perspectives; performing three-dimensional human body modeling based on the human body images from the multiple camera perspectives to obtain the three-dimensional human body model corresponding to the human body images from the multiple camera perspectives; performing projection based on the three-dimensional human body model, and compressing and outputting the depth maps of the person to be estimated from the multiple camera perspectives obtained by the projection.
[0166] In addition, when the logical instructions in the above-mentioned memory 830 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0167] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the human body depth estimation method provided by the above-mentioned various methods. The method includes: obtaining human body images of a person to be estimated from multiple camera perspectives; performing three-dimensional human body modeling based on the human body images from the multiple camera perspectives to obtain a three-dimensional human body model corresponding to the human body images from the multiple camera perspectives; performing projection based on the three-dimensional human body model, and compressing and outputting the depth maps of the person to be estimated from the multiple camera perspectives obtained by the projection.
[0168] In yet another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the human body depth estimation method provided by the above-mentioned various methods. The method includes: obtaining human body images of a person to be estimated from multiple camera perspectives; performing three-dimensional human body modeling based on the human body images from the multiple camera perspectives to obtain a three-dimensional human body model corresponding to the human body images from the multiple camera perspectives; performing projection based on the three-dimensional human body model, and compressing and outputting the depth maps of the person to be estimated from the multiple camera perspectives obtained by the projection.
[0169] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.
[0170] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0171] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.
Claims
1. A human body depth estimation method, characterized in that, Including: Obtain human body images of the person to be estimated from multiple camera viewpoints; Perform 3D human body modeling based on the human body images from the multiple camera viewpoints to obtain a 3D human body model corresponding to the human body images from the multiple camera viewpoints; Perform projection based on the 3D human body model, and compress and output the depth maps of the person to be estimated from the multiple camera viewpoints obtained by the projection; The performing 3D human body modeling based on the human body images from the multiple camera viewpoints to obtain a 3D human body model corresponding to the human body images from the multiple camera viewpoints includes: Perform 3D human body modeling based on the human body images from the multiple camera viewpoints by applying a pose reconstruction model to obtain a corresponding 3D human body model; The pose reconstruction model is trained based on sample human body images of sample persons from multiple camera viewpoints, sample pose information corresponding to the sample human body images from the multiple camera viewpoints, predicted human body models corresponding to the sample human body images from the multiple camera viewpoints, and projections of the predicted human body models from the multiple camera viewpoints; The pose reconstruction model is trained based on the following steps: Perform human body detection and key point detection on the sample human body images from the multiple camera viewpoints to obtain sample human body regions and corresponding sample pose information in the sample human body images from the multiple camera viewpoints; Perform 3D human body modeling based on an initial reconstruction model to obtain a predicted human body model corresponding to the sample human body regions from the multiple camera viewpoints, and determine the projections of the predicted human body model from the multiple camera viewpoints; Determine a region reconstruction loss based on the consistency between the projections from the multiple camera viewpoints and the sample human body regions from the multiple camera viewpoints; Determine a pose reconstruction loss based on the consistency between the sample pose information corresponding to the sample human body images from the multiple camera viewpoints and the predicted pose information corresponding to the projections from the multiple camera viewpoints; Perform parameter iteration on the initial reconstruction model based on the region reconstruction loss and the pose reconstruction loss to obtain a pose reconstruction model.
2. The human body depth estimation method according to claim 1, characterized in that The sample pose information includes the coordinates and pose parameters of the human body joint points of the sample person in the corresponding sample human body image; the predicted pose information includes the coordinates and pose parameters of the human body joint points in the projection from the corresponding camera viewpoint; The determining a pose reconstruction loss based on the consistency between the sample pose information corresponding to the sample human body images from the multiple camera viewpoints and the predicted pose information corresponding to the projections from the multiple camera viewpoints includes: Determine a position reconstruction loss based on the consistency between the coordinates of the human body joint points in the sample human body images from the multiple camera viewpoints and the coordinates of the human body joint points in the projections from the multiple camera viewpoints; And / or, determine a pose reconstruction loss based on the pose parameters of the human body joint points in the projections from the multiple camera viewpoints, or the consistency between the pose parameters of the human body joint points in the sample human body images from the multiple camera viewpoints and the pose parameters of the human body joint points in the projections from the multiple camera viewpoints; Determine the pose reconstruction loss based on the position reconstruction loss and / or the pose reconstruction loss.
3. The human body depth estimation method according to claim 1 or 2, characterized in that Performing projection based on the three-dimensional human body model, and compressing and outputting the depth maps of the person to be estimated obtained by the projection under the multiple camera views, including: Performing projection based on the three-dimensional human body model to obtain the initial depth map and initial normal map of the person to be estimated under the multiple camera views; Based on the human body image under any one camera view and the human body images under other camera views, as well as the camera pose matrix corresponding to any one camera view and the camera pose matrix corresponding to the other camera views, optimizing the initial depth map and initial normal map under any one camera view to obtain the depth map and normal map under any one camera view; Compressing and outputting the depth maps and normal maps of the person to be estimated under the multiple camera views.
4. The human body depth estimation method according to claim 3, characterized in that The optimizing the initial depth map and initial normal map under any one camera view based on the human body image under any one camera view and the human body images under other camera views, as well as the camera pose matrix corresponding to any one camera view and the camera pose matrix corresponding to the other camera views to obtain the depth map and normal map under any one camera view, including: Determining the pixel information of a target point in the human body image under any one camera view, the pixel information of the projected point of the target point in the human body image under the other camera views, and the included angle between the camera ray corresponding to the other camera view and the normal of the projected point of the target point in the human body image under the other camera view; Based on the pixel information of the target point in the human body image under any one camera view, the pixel information of the projected point of the target point in the human body image under the other camera views, the included angle, as well as the camera pose matrix corresponding to any one camera view and the camera pose matrix corresponding to the other camera views, optimizing the initial depth map and initial normal map under any one camera view to obtain the depth map and normal map under any one camera view.
5. The human body depth estimation method according to claim 3, characterized in that The compressing and outputting the depth maps and normal maps of the person to be estimated under the multiple camera views, including: Determining the important human body regions in the human body image under any one camera view; Based on the human body image under any one camera view and the human body images under other camera views, as well as the camera pose matrix corresponding to any one camera view and the camera pose matrix corresponding to the other camera views, optimizing the depth maps and normal maps corresponding to the important human body regions under any one camera view to obtain the target depth map and target normal map under any one camera view; Storing the target depth map, target normal map and human body image under any one camera view by bytes, and compressing and outputting the stored contents under the multiple camera views.
6. A human body depth estimation device, characterized in that, Including: An acquisition unit, configured to acquire the human body images of the person to be estimated under multiple camera views; A modeling unit, configured to perform three-dimensional human body modeling based on the human body images under the multiple camera views to obtain the three-dimensional human body model corresponding to the human body images under the multiple camera views; An estimation unit for performing projection based on the three-dimensional human body model and compressing and outputting the depth maps of the person to be estimated obtained by the projection under the multiple camera viewpoints; The three-dimensional human body modeling based on the human body images under the multiple camera viewpoints to obtain the three-dimensional human body model corresponding to the human body images under the multiple camera viewpoints includes: Performing three-dimensional human body modeling on the human body images under the multiple camera viewpoints by applying a pose reconstruction model to obtain the corresponding three-dimensional human body model; The pose reconstruction model is trained based on the sample human body images of the sample person under the multiple camera viewpoints, the sample pose information corresponding to the sample human body images under the multiple camera viewpoints, the predicted human body model corresponding to the sample human body images under the multiple camera viewpoints, and the projections of the predicted human body model under the multiple camera viewpoints; The pose reconstruction model is trained based on the following steps: Performing human body detection and key point detection on the sample human body images under the multiple camera viewpoints to obtain the sample human body regions and the corresponding sample pose information in the sample human body images under the multiple camera viewpoints; Performing three-dimensional human body modeling based on an initial reconstruction model to obtain the predicted human body model corresponding to the sample human body regions under the multiple camera viewpoints, and determining the projections of the predicted human body model under the multiple camera viewpoints; Determining a region reconstruction loss based on the consistency between the projections under the multiple camera viewpoints and the sample human body regions under the multiple camera viewpoints; Determining a pose reconstruction loss based on the consistency between the sample pose information corresponding to the sample human body images under the multiple camera viewpoints and the predicted pose information corresponding to the projections under the multiple camera viewpoints; Based on the region reconstruction loss and the pose reconstruction loss, performing parameter iteration on the initial reconstruction model to obtain a pose reconstruction model.
7. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the human body depth estimation method according to any one of claims 1 to 5.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the human body depth estimation method according to any one of claims 1 to 5.
Citation Information
Patent Citations
High-precision face reconstruction system and method based on light polarization reflectivity, terminal and medium
CN115482329A