Human body depth estimation method and device, electronic equipment and storage medium

Three-dimensional human body modeling and projection through human body images from multiple camera perspectives, the hollow problem of weak texture area depth calculation in human body depth estimation is solved, and the storage and use of results is simplified by compressing the output, improving the accuracy and efficiency of the estimation.

CN119991764AActive Publication Date: 2025-05-13IFLYTEK CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510459227.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-05-13
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

Existing human body depth estimation technology is difficult to calculate accurate depth information when dealing with weak texture areas, resulting in a large number of output files, which is inconvenient to store and access.

Method used

By acquiring human images from multiple camera perspectives for three-dimensional human modeling, a pose reconstruction model is used to build a three-dimensional human model and projecting it to each camera perspective to generate a depth map. The depth map is then compressed and outputted, reducing the number of files and improving the transmission efficiency.

Benefits of technology

It effectively solves the problem of hollowness in the depth calculation of weak texture areas, improves the accuracy and completeness of human depth estimation, and simplifies the storage and use of results through compressed output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991764A_ABST
    Figure CN119991764A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer vision, and provides a human body depth estimation method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining human body images of a to-be-estimated person at a plurality of camera visual angles, carrying out the three-dimensional human body modeling according to the human body images, and obtaining a three-dimensional human body model corresponding to the human body images at the plurality of camera visual angles; projection is carried out based on the three-dimensional human body model, and the depth map of the to-be-estimated person obtained through projection under multiple camera visual angles is compressed and output, so that the problem that the depth of the weak texture area cannot be recovered at present, resulting in cavities in the depth map is solved, and the depth information of the weak texture area can be effectively solved. Therefore, the accuracy and integrity of human body depth estimation are guaranteed, the depth map is compressed and output, the number of output files is reduced, the transmission speed is increased, the occupied transmission space is reduced, the integrity and orderliness of information are improved, and the output result is effectively simplified and is convenient to take and use.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a method, device, electronic device and storage medium for estimating human depth. Background Art

[0002] In the field of computer vision and 3D reconstruction, human depth estimation is a crucial technology that aims to restore the 3D spatial information of the human body from 2D images. This technology is widely used in many fields such as augmented reality, virtual reality, and human-computer interaction, and is of great significance for improving user experience and system performance.

[0003] Current human depth estimation tasks mostly use algorithm frameworks such as Patchmatch to obtain dense depth maps from various perspectives. However, due to the lack of accurate initial values, such schemes are difficult to calculate accurate depth information for weak texture areas; that is, for human body parts that lack obvious texture features, such as areas wearing white or black clothes, the image information in these areas is highly consistent or the contrast is extremely low, and the initial values ​​are not accurate, resulting in the inability to effectively calculate depth values ​​in these weak texture areas. This phenomenon manifests itself as holes in the depth map, which seriously affects the integrity and accuracy of human depth estimation. In addition, the current human depth estimation task ultimately outputs a large number of files, which are extremely inconvenient to store and access. Summary of the invention

[0004] The present invention provides a human body depth estimation method, device, electronic device and storage medium, which are used to solve the deficiencies of the human body depth estimation task in the prior art in processing weak texture areas and data storage output, improve the accuracy and completeness of human body depth estimation, and effectively simplify the output results for subsequent use.

[0005] The present invention provides a method for estimating human body depth, comprising: Obtaining human body images of the person to be estimated under multiple camera perspectives; Performing three-dimensional human body modeling based on the human body images under the multiple camera perspectives to obtain three-dimensional human body models corresponding to the human body images under the multiple camera perspectives; Projection is performed based on the three-dimensional human body model, and depth maps of the person to be estimated obtained by the projection under the multiple camera viewing angles are compressed and output.

[0006] According to a method for estimating human depth provided by the present invention, performing three-dimensional human body modeling based on human body images under the perspectives of multiple cameras to obtain three-dimensional human body models corresponding to the human body images under the perspectives of multiple cameras includes: Based on the human body images under the multiple camera perspectives, applying the posture reconstruction model to perform three-dimensional human body modeling to obtain a corresponding three-dimensional human body model; The posture reconstruction model is obtained based on sample body images of sample persons under multiple camera perspectives, sample pose information corresponding to the sample body images under the multiple camera perspectives, predicted body models corresponding to the sample body images under the multiple camera perspectives, and projection training of the predicted body model under the multiple camera perspectives.

[0007] According to a human body depth estimation method provided by the present invention, the posture reconstruction model is trained based on the following steps: Performing human body detection and key point detection on the sample human body images under the multiple camera perspectives to obtain sample human body regions and corresponding sample pose information in the sample human body images under the multiple camera perspectives; Performing three-dimensional human body modeling based on the initial reconstruction model, obtaining a predicted human body model corresponding to the sample human body region under the multiple camera perspectives, and determining the projection of the predicted human body model under the multiple camera perspectives; Determining a regional reconstruction loss based on consistency between the projections under the multiple camera perspectives and the sample human body regions under the multiple camera perspectives; Determining a pose reconstruction loss based on consistency between sample pose information corresponding to the sample human body images under the multiple camera perspectives and predicted pose information corresponding to the projections under the multiple camera perspectives; Based on the region reconstruction loss and the posture reconstruction loss, the initial reconstruction model is iterated on parameters to obtain a posture reconstruction model.

[0008] According to a method for estimating human depth provided by the present invention, the sample pose information includes coordinates and pose parameters of human joint points of a sample person in a sample human image; the predicted pose information includes coordinates and pose parameters of human joint points in a projection under a corresponding camera perspective; The determining of the pose reconstruction loss based on the consistency between the sample pose information corresponding to the sample human body images under the multiple camera perspectives and the predicted pose information corresponding to the projections under the multiple camera perspectives includes: Determine the position reconstruction loss based on the consistency between the coordinates of the human joint points in the sample human body images under the multiple camera perspectives and the coordinates of the human joint points in the projections under the multiple camera perspectives; And / or, determining the posture reconstruction loss based on the posture parameters of the human joint points in the projections under the multiple camera perspectives, or the consistency between the posture parameters of the human joint points in the sample human body images under the multiple camera perspectives and the posture parameters of the human joint points in the projections under the multiple camera perspectives; The pose reconstruction loss is determined based on the position reconstruction loss and / or the pose reconstruction loss.

[0009] According to a method for estimating human depth provided by the present invention, the method includes: performing projection based on the three-dimensional human body model, and compressing and outputting the depth map of the person to be estimated under the multiple camera perspectives obtained by the projection, including: Performing projection based on the three-dimensional human body model to obtain an initial depth map and an initial normal map of the person to be estimated under the perspectives of the multiple cameras; Based on the human body image under any camera perspective and the human body image under other camera perspectives, as well as the camera pose matrix corresponding to the any camera perspective and the camera pose matrix corresponding to the other camera perspectives, optimizing the initial depth map and the initial normal map under any camera perspective to obtain the depth map and the normal map under any camera perspective; The depth map and the normal map of the person to be estimated under the multiple camera viewing angles are compressed and output.

[0010] According to a method for estimating human depth provided by the present invention, based on a human image under any camera perspective and a human image under other camera perspectives, as well as a camera pose matrix corresponding to any camera perspective and a camera pose matrix corresponding to other camera perspectives, an initial depth map and an initial normal map under any camera perspective are optimized to obtain a depth map and a normal map under any camera perspective, including: Determine pixel information of a target point in the human body image under any camera perspective, pixel information of a projection point of the target point in the human body image under the other camera perspectives, and an angle between a camera ray corresponding to the other camera perspectives and a normal of the projection point of the target point under the other camera perspectives; Based on the pixel information of the target point in the human body image under any camera perspective, the pixel information of the projection point of the target point in the human body image under the other camera perspectives, the angle, and the camera pose matrix corresponding to any camera perspective and the camera pose matrix corresponding to the other camera perspectives, the initial depth map and the initial normal map under any camera perspective are optimized to obtain the depth map and the normal map under any camera perspective.

[0011] According to a method for estimating human depth provided by the present invention, compressing and outputting the depth map and the normal map of the person to be estimated under the multiple camera viewing angles comprises: Determine the important human body regions in the human body image under any camera view; Based on the human body image under any camera perspective and the human body image under other camera perspectives, as well as the camera pose matrix corresponding to any camera perspective and the camera pose matrix corresponding to the other camera perspectives, the depth map and normal map corresponding to the important human body area under any camera perspective are optimized to obtain the target depth map and target normal map under any camera perspective; The target depth map, target normal map and human body image under any camera viewing angle are stored in bytes, and the stored contents under the multiple camera viewing angles are compressed and output.

[0012] The present invention also provides a human body depth estimation device, comprising: An acquisition unit, used for acquiring human body images of a person to be estimated under multiple camera viewing angles; A modeling unit, configured to perform three-dimensional human body modeling based on the human body images under the multiple camera perspectives, and obtain a three-dimensional human body model corresponding to the human body images under the multiple camera perspectives; The estimation unit is used to perform projection based on the three-dimensional human body model, and compress and output the depth map of the person to be estimated under the perspectives of the multiple cameras obtained by the projection.

[0013] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein when the processor executes the computer program, a human body depth estimation method as described in any one of the above is implemented.

[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the human body depth estimation method as described in any one of the above is implemented.

[0015] The human body depth estimation method, device, electronic device and storage medium provided by the present invention perform three-dimensional human body modeling through human body images under multiple camera perspectives, so as to obtain the corresponding three-dimensional human body model in a fast parameterized manner, and project the three-dimensional human body model thus constructed to each camera perspective to obtain a depth map based on the characteristic that there are no holes. This well solves the problem that the depth of weak texture areas cannot be restored in traditional solutions, resulting in holes in the depth map, and can effectively solve the depth information of weak texture areas, thereby ensuring the accuracy and completeness of human body depth estimation. Finally, the depth map is compressed and output, which not only reduces the number of output files, improves the transmission speed, reduces the transmission space occupied, but also improves the integrity and orderliness of the information, and realizes the effective simplification of the output results for easy access. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0017] Figure 1 It is a flow chart of a method for estimating human body depth provided by the present invention; Figure 2 is an example diagram of the three-dimensional human body model provided by the present invention; Figure 3 is an example diagram of the skeletal joint points provided by the present invention; Figure 4 is an example diagram of facial feature points provided by the present invention; Figure 5 is an example diagram of a sample human body region provided by the present invention; Figure 6 is an example diagram of the depth map storage process provided by the present invention; Figure 7 is a schematic structural diagram of a human body depth estimation device provided by the present invention; Figure 8 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0018] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings in the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0019] At present, human depth estimation tasks mostly use algorithm frameworks such as Patchmatch to obtain dense depth maps under various viewing angles. However, such solutions have certain limitations. First, due to the lack of accurate initial values, the depth cannot be calculated for weak texture areas, and holes often appear in the depth map. Second, the final output results have a large number of files, the storage method is not simple enough, and it is inconvenient to access.

[0020] In this regard, the present invention provides a human body depth estimation method, which aims to solve the deficiencies in weak texture areas and data storage output during human body depth estimation, improve the accuracy and completeness of human body depth estimation, and effectively simplify the output results, reduce the number of files, and improve the transmission speed for subsequent retrieval. Figure 1is a flow chart of the method for estimating the depth of a human body provided by the present invention, such as Figure 1 As shown, the method includes: Step 110, obtaining human body images of the person to be estimated under multiple camera viewing angles; Step 120, performing three-dimensional human body modeling based on the human body images under multiple camera perspectives to obtain three-dimensional human body models corresponding to the human body images under multiple camera perspectives; Step 130 , performing projection based on the three-dimensional human body model, and compressing and outputting depth maps of the person to be estimated under multiple camera viewing angles obtained by the projection.

[0021] Specifically, considering the deficiencies of traditional human depth estimation tasks in processing weak texture areas, as well as in storing and outputting results; namely, the depth value cannot be effectively calculated for weak texture areas, resulting in holes in the depth map, and the final output file number is large, which is extremely inconvenient to store and access. In an embodiment of the present invention, it is proposed that three-dimensional human body modeling can be first performed based on human body images under multiple camera perspectives, so as to obtain a corresponding three-dimensional human body model through a fast parameterization method, and then human body depth estimation is performed based on this three-dimensional human body model to obtain accurate depth maps under multiple camera perspectives, and the depth maps are compressed and output. This can not only reduce the number of output files, improve the transmission speed, and reduce the transmission space occupied, but also can be based on the fast parameterization method. The three-dimensional human body model constructed will not have the characteristics of holes, and it can be projected to each camera perspective to obtain a depth map, thereby well solving the problem that the depth of weak texture areas cannot be restored in traditional solutions, and the depth information of weak texture areas can be effectively solved, thereby ensuring the accuracy and completeness of human body depth estimation.

[0022] In detail, in actual application, before performing human depth estimation, it is necessary to first determine the two-dimensional image of the person to be estimated, that is, the human body images of the person to be estimated under multiple camera perspectives, and these human body images can constitute a complete human body image sequence of the person to be estimated. Here, the human body image can be obtained by taking pictures with a camera, that is, when the camera is at a certain perspective, the image of the person to be estimated taken by it is the human body image under the perspective of the camera; the human body image can also be directly searched / searched / downloaded. After obtaining the human body image, the corresponding camera perspective can be calibrated by manual or automatic methods.

[0023] Among them, the human body image sequence can be one or more. In the case of multiple human body image sequences, the persons to be estimated corresponding to different human body image sequences can be the same or different, that is, it can be a human body image sequence corresponding to the human body images of the same person to be estimated in different states (postures, environments, places, etc.), or it can be a human body image sequence corresponding to the human body images of different persons in the same or different states.

[0024] After obtaining the human image sequence, in the embodiment of the present invention, three-dimensional human modeling can be performed based on the human image sequence to obtain a three-dimensional human model corresponding to the human image sequence. In consideration of the deficiencies in the traditional scheme in the weak texture area when estimating the human depth, in the embodiment of the present invention, it is proposed to use a fast parameterization method to construct a three-dimensional human model, namely SMPL (Skinned Multi-Person Linear Model) when performing three-dimensional human modeling. With the excellent characteristics of the model itself (no holes), the defects of the current scheme in processing weak texture areas can be overcome, thereby significantly improving the integrity and accuracy of the depth map, which is more conducive to subsequent use.

[0025] Specifically, here, the human body images of the person to be estimated under multiple camera perspectives contained in the human body image sequence can be used as the basis for three-dimensional human body modeling, so as to model a three-dimensional human body model that matches the human body images under each camera perspective in the human body image sequence. That is, the human body images under multiple camera perspectives are aligned in a fast parameterized manner to obtain a three-dimensional human body model of the person to be estimated, and the SMPL thus constructed can be aligned with all human body images in the human body image sequence, that is, the projection of the SMPL under each camera perspective is basically consistent with the human body area in the human body image of the person to be estimated under each camera perspective, so that the accuracy of the modeling is guaranteed, and an accurate basis is provided for subsequent depth estimation.

[0026] It should be noted that the process of obtaining the three-dimensional human body model SMPL by three-dimensional human body modeling here can be implemented through a pre-trained network model dedicated to human body modeling, or through other algorithms and frameworks, such as the Human MeshRecovery end-to-end framework, the Multi-View Stereo series, etc., and the embodiments of the present invention do not make specific limitations on this.

[0027] Furthermore, after obtaining a three-dimensional human body model corresponding to the human body image sequence through three-dimensional human body modeling, in an embodiment of the present invention, human body depth estimation can be performed based on this three-dimensional human body model to obtain a depth map of the person to be estimated under each camera perspective, and finally output it.

[0028] Specifically, unlike the method of randomly initializing the depth in the traditional solution (the initial value is not accurate or even completely wrong), in the embodiment of the present invention, the three-dimensional human body model obtained in the previous step is used for depth estimation to obtain an accurate initial value, based on which an accurate depth map under each camera perspective can be obtained. Here, specifically, the three-dimensional human body model can be projected to project the three-dimensional human body model SMPL to each camera perspective, and its projection under each camera perspective is used as the initial value, so as to ensure the accuracy of the initial value and obtain the initial depth map under each camera perspective; then, based on this initial depth map, multiple iterative optimizations can be performed to update the depth value to ensure precision and accuracy, and finally an accurate depth map of the person to be estimated under each camera perspective can be obtained.

[0029] After that, it can be output. Taking into account the defects in the result output of the current human body depth estimation task, in the embodiment of the present invention, when outputting the result, compression processing can be performed first to reduce the number of files and the transmission space occupied, and ensure the integrity of the information, improve the transmission efficiency, and facilitate subsequent calls.

[0030] Specifically, here the depth maps under multiple camera perspectives may be compressed. Specifically, when performing the compression processing, the information may be stored in units of camera perspectives, and then the overall stored information, including the depth maps, human body images, etc. under all camera perspectives, may be compressed, and the compressed information may be output. In this way, while ensuring the integrity and orderliness of the information, the number of files, transmission time and space occupancy may be greatly reduced, thereby improving the transmission efficiency, and enabling users / downstream tasks to obtain the depth information of the person to be estimated in the three-dimensional space more quickly, thereby achieving complete human depth estimation, and ensuring the integrity, accuracy and effectiveness of human depth estimation.

[0031] The human body depth estimation method provided by the present invention performs three-dimensional human body modeling through human body images under multiple camera perspectives, so as to obtain the corresponding three-dimensional human body model in a fast parameterized manner, and based on the characteristic that the three-dimensional human body model constructed in this way will not have holes, it is projected to each camera perspective to obtain a depth map, which well solves the problem that the depth of weak texture areas cannot be restored in traditional solutions, resulting in holes in the depth map, and can effectively solve the depth information of the weak texture areas, thereby ensuring the accuracy and completeness of the human body depth estimation, and finally compresses the depth map for output, which not only reduces the number of output files, improves the transmission speed, reduces the transmission space occupied, but also improves the integrity and orderliness of the information, and realizes the effective simplification of the output results for easy access.

[0032] Based on the above embodiment, step 120 includes: Based on human body images from multiple camera perspectives, a posture reconstruction model is used to perform 3D human body modeling to obtain a corresponding 3D human body model; The posture reconstruction model is obtained based on sample body images of sample persons under multiple camera perspectives, sample pose information corresponding to the sample body images under multiple camera perspectives, predicted body models corresponding to the sample body images under multiple camera perspectives, and projection training of the predicted body model under multiple camera perspectives.

[0033] Specifically, the above process of performing three-dimensional human body modeling based on human body images under multiple camera perspectives can be implemented through a posture reconstruction model. That is, the posture reconstruction model can obtain information such as the posture and shape of the person to be estimated from the human body area of ​​the two-dimensional human body image in a fast parameterized manner, and perform three-dimensional modeling based on the information, thereby obtaining a corresponding three-dimensional human body model.

[0034] Specifically, human body images from multiple camera perspectives in a human body image sequence may be input into a posture reconstruction model to perform three-dimensional human body modeling through the model. Specifically, parameterized modeling may be performed by obtaining information such as the posture and shape of the person to be estimated from the input two-dimensional human body image. Figure 2 is an example diagram of the three-dimensional human body model provided by the present invention, such as Figure 2 As shown, by establishing the human body joint points and surface topology in three-dimensional space, an accurate three-dimensional human body model can be obtained.

[0035] However, it is worth noting that in order to ensure the accuracy of the three-dimensional human body model obtained by modeling, in the embodiment of the present invention, before applying this posture reconstruction model for three-dimensional human body modeling, it is necessary to train it to improve the model performance, so that it can perform better in the actual three-dimensional human body modeling task, and the three-dimensional human body model obtained by modeling is more accurate. Here, specifically, the model training can be performed by applying sample human body images with posture information under multiple camera perspectives to obtain a trained posture reconstruction model.

[0036] Specifically, when training the model, in order to enable the model to update faster in the expected direction and converge faster, when calculating the loss, a multi-angle calculation method is used in the embodiment of the present invention to measure the loss of the model in the three-dimensional human body modeling task from multiple angles, so that it can converge faster through multiple constraints, and the modeling is faster and more accurate in actual application. That is, the loss of the model in the three-dimensional human body modeling task can be measured from the two levels of human body area and posture based on the sample human body images and corresponding sample posture information of the sample personnel under multiple camera perspectives, as well as the predicted human body models corresponding to the sample human body images under these multiple camera perspectives and the projections of the predicted human body models under multiple camera perspectives. The model training is carried out based on this loss, and the final trained model, that is, the posture reconstruction model, can be obtained. And the trained model can perform better in the actual three-dimensional human body modeling task and model more accurately.

[0037] Based on the above embodiment, the posture reconstruction model is trained based on the following steps: Performing human body detection and key point detection on sample human body images under multiple camera perspectives to obtain sample human body regions and corresponding sample pose information in sample human body images under multiple camera perspectives; Performing 3D human body modeling based on the initial reconstruction model, obtaining a predicted human body model corresponding to the sample human body region under multiple camera perspectives, and determining the projection of the predicted human body model under multiple camera perspectives; Determine the regional reconstruction loss based on the consistency between the projections under multiple camera views and the sample human body regions under multiple camera views; Determine the pose reconstruction loss based on the consistency between the sample pose information corresponding to the sample human body images under multiple camera perspectives and the predicted pose information corresponding to the projections under multiple camera perspectives; Based on the regional reconstruction loss and the pose reconstruction loss, the parameters of the initial reconstruction model are iterated to obtain the pose reconstruction model.

[0038] Specifically, the training process of the posture reconstruction model may include the following steps: First, it is necessary to obtain training samples, that is, sample data required for model training, which may include sample human body images of sample persons under multiple camera perspectives, and the pose information corresponding to the sample human body images, that is, sample pose information. The sample pose information here can be obtained by performing key point detection on the sample human body images.

[0039] However, it is worth noting that in the current human depth estimation task, when reconstructing the human body depth, since the input two-dimensional image contains not only the human body but also the background, the traditional depth estimation algorithm does not distinguish between them and will reconstruct the depth of the unnecessary background.

[0040] In this regard, in an embodiment of the present invention, in order to avoid the problem of reconstructing background depth when modeling a three-dimensional human body, it is proposed to directly eliminate background information when training a posture reconstruction model used for three-dimensional human body modeling, and perform model training based on the human body area, thereby avoiding the problem of background environment being reconstructed from the root and achieving the distinction between the background and the human body.

[0041] Here, specifically, after obtaining sample human body images of sample persons under multiple camera perspectives, key point detection and human body detection can be performed on them to obtain sample human body regions in the sample human body images under multiple camera perspectives, as well as sample pose information corresponding to each sample human body region. Specifically, key point detection can be performed on each sample human body image, including facial feature point detection and skeletal joint point detection, to obtain sample pose information, i.e., coordinates and pose parameters of human body joint points; human body joint points here include facial feature points and skeletal joint points; at the same time, human body detection can be performed on each sample human body image to identify the area where the sample person is located, and segment it, thereby obtaining the sample human body area in each sample human body image.

[0042] Figure 3 is an example diagram of the skeletal joint points provided by the present invention, such as Figure 3 As shown in the figure, when detecting the skeletal joints of a sample human image, a posture estimation method such as OpenPose, HRNet, etc. can be used to extract the skeletal joints of the human body, such as shoulders, elbows, knees, ankles, etc., from the sample human image to obtain the coordinates and posture parameters of the skeletal joints. Figure 4 is an example diagram of facial feature points provided by the present invention, such as Figure 4 As shown, by detecting facial feature points on a sample human image, 68 two-dimensional facial feature points can be obtained. Figure 5 is an example diagram of a sample human body region provided by the present invention, such as Figure 5 As shown, the human body mask segmentation method can be used to extract the sample human body area from each sample human body image.

[0043] Subsequently, the initial model in the training process, i.e., the initial reconstruction model, can be applied to perform three-dimensional human body modeling, so as to reconstruct the three-dimensional human body model of the sample person based on the sample human body regions in the sample human body images under multiple camera perspectives in a parametric representation manner, i.e., the predicted human body model corresponding to each sample human body region. Here, the initial model can be a model constructed on the basis of a network model based on a parametric representation, and the specific architecture of the model can be adjusted according to actual conditions, requirements, etc., and the embodiments of the present invention do not specifically limit this.

[0044] After obtaining the predicted human body model, in the embodiment of the present invention, it is also necessary to determine the projection of the predicted human body model under multiple camera perspectives, that is, the predicted human body model needs to be projected to the camera perspective corresponding to each sample human body image, so as to obtain its projection under multiple camera perspectives.

[0045] Then, based on the sample body images and corresponding sample pose information of the sample persons under multiple camera perspectives, as well as the predicted body model and its projection under multiple camera perspectives, the loss of the initial reconstructed model in the three-dimensional body modeling task can be measured from the two levels of body area and pose, thereby obtaining the area reconstruction loss and pose reconstruction loss.

[0046] Here, the regional reconstruction loss can be determined specifically based on the consistency between the projection of the predicted human body model under multiple camera perspectives and the sample human body region in the sample human body images under multiple camera perspectives; that is, based on the similarity or difference between the projection of the predicted human body model under the camera perspective obtained by modeling and the sample human body region under the corresponding camera perspective, the loss of the model in reconstructing the human body region is judged, and the loss is the regional reconstruction loss.

[0047] At the same time, the pose reconstruction loss can be determined based on the consistency between the sample pose information corresponding to the sample human body images under multiple camera perspectives and the predicted pose information corresponding to the projections of the predicted human body model under multiple camera perspectives; that is, the loss of the model in human body pose reconstruction is judged based on the similarity or difference between the predicted pose information (coordinates and posture parameters of human body joints) corresponding to the projections of the predicted human body model under the camera perspective and the sample pose information under the corresponding camera perspective. This loss is the pose reconstruction loss.

[0048] Afterwards, the loss of the model in the entire 3D human body modeling task can be determined based on the above-mentioned regional reconstruction loss and pose reconstruction loss, and the parameters of the initial reconstruction model can be iterated based on this loss, so that the projection of the predicted human body model output by the model after parameter adjustment under the camera perspective can be as close as possible to the real sample human body area, and the predicted pose information can be as consistent as possible with the real sample pose information, and finally a trained pose reconstruction model can be obtained.

[0049] Based on the above embodiment, the sample pose information includes coordinates and pose parameters of the human joint points of the sample person in the sample human image; the predicted pose information includes coordinates and pose parameters of the human joint points in the projection under the corresponding camera perspective; Based on the consistency between the sample pose information corresponding to the sample human body images under multiple camera perspectives and the predicted pose information corresponding to the projections under multiple camera perspectives, the pose reconstruction loss is determined, including: Determine the position reconstruction loss based on the consistency between the coordinates of the human joint points in the sample human body images under multiple camera perspectives and the coordinates of the human joint points in the projections under multiple camera perspectives; and / or, determining a posture reconstruction loss based on the posture parameters of the human joint points in the projections under multiple camera perspectives, or the consistency between the posture parameters of the human joint points in the sample human body images under multiple camera perspectives and the posture parameters of the human joint points in the projections under multiple camera perspectives; Based on the position reconstruction loss and / or the pose reconstruction loss, a pose reconstruction loss is determined.

[0050] Specifically, the process of determining the pose reconstruction loss according to the consistency between the sample pose information corresponding to the sample human body images under multiple camera perspectives and the predicted pose information corresponding to the projections under multiple camera perspectives may specifically include: First, the loss can be measured from the position level, that is, based on the consistency between the coordinates of the human joint points in the sample human body images under multiple camera perspectives and the coordinates of the human joint points in the projections of the predicted human body model under multiple camera perspectives, the loss of the position coordinates of the initial reconstructed model when reconstructing the human joint points can be determined, that is, whether the position modeling of the human joint points is accurate and whether there are any differences, thereby obtaining the position reconstruction loss.

[0051] At the same time, loss measurement can also be performed at the posture level, that is, the loss of human posture in the initial reconstruction model when reconstructing the human joints can be determined based on the consistency between the posture parameters of the human joints in the sample human images under multiple camera perspectives and the posture parameters of the human joints in the projections under multiple camera perspectives. That is, whether the posture modeling of the human joints is accurate and whether there are any differences, thereby obtaining the posture reconstruction loss.

[0052] Alternatively, the posture reconstruction loss can be determined only based on the posture parameters of the human joints in the projections of the predicted human body model under multiple camera perspectives; here, specifically, the loss of the human body posture when the initial reconstruction model is reconstructing the human joints can be determined by judging the rationality, normality, coherence, etc. of the posture of the human joints in the projections corresponding to the predicted human body model, that is, whether the posture modeling of the human joints is reasonable, thereby obtaining the posture reconstruction loss.

[0053] Then, the pose reconstruction loss can be determined according to the above-mentioned position reconstruction loss and / or pose reconstruction loss. Specifically, the position reconstruction loss can be directly used as the pose reconstruction loss, or the pose reconstruction loss can be used as the pose reconstruction loss, or the position reconstruction loss and the pose reconstruction loss can be combined to determine the pose reconstruction loss in a weighted manner.

[0054] Based on the above embodiment, the loss of the initial reconstruction model in the entire 3D human body modeling task It can be expressed by the following formula:

[0055] In the formula, , and Represents position reconstruction loss, region reconstruction loss and posture reconstruction loss respectively. , and It represents the weights corresponding to the position reconstruction loss, area reconstruction loss and posture reconstruction loss respectively.

[0056] Based on the above embodiment, step 130 includes: Projection is performed based on the three-dimensional human body model to obtain an initial depth map and an initial normal map of the person to be estimated under multiple camera perspectives; Based on the human body image under any camera perspective and the human body image under other camera perspectives, as well as the camera pose matrix corresponding to the camera perspective and the camera pose matrix corresponding to the other camera perspectives, the initial depth map and the initial normal map under the camera perspective are optimized to obtain the depth map and the normal map under the camera perspective; The depth map and normal map of the person to be estimated under multiple camera perspectives are compressed and output.

[0057] Specifically, the process of performing projection based on the three-dimensional human body model and compressing and outputting the depth map of the person to be estimated under multiple camera perspectives obtained by the projection specifically includes: Considering that in traditional depth estimation schemes, random depth and normal are used in space as initial values ​​during depth estimation to facilitate subsequent depth value optimization and calculation, but this will result in slow depth estimation and low efficiency.

[0058] Based on this, in an embodiment of the present invention, an initial depth map is first obtained through a three-dimensional human body model, and then the initial depth map is updated and optimized by optimizing each pixel point one by one, and finally a more accurate depth map can be obtained. Moreover, in this process, by effectively utilizing the depth and normal information for updating and optimization, the iterative process can be made more efficient, converge faster, and the final depth map can have higher accuracy.

[0059] Specifically, the three-dimensional human body model may be projected first to obtain an initial depth map and an initial normal map of the person to be estimated under multiple camera perspectives. Then, the initial depth map and the initial normal map may be optimized to obtain depth maps and normal maps under multiple camera perspectives; that is, the depth information and color information under the corresponding camera perspectives are combined, and the initial depth map and the initial normal map are jointly optimized to obtain the depth map and the normal map under the corresponding camera perspectives.

[0060] Specifically, for any camera perspective, when optimizing the initial depth map under it, the characteristic that the color information of the triangular facets on the three-dimensional human body model should be the same under different camera perspectives, that is, the pixel values ​​(color information) obtained by projecting the same triangular facet under different camera perspectives should be consistent, is used to update the initial depth map and the initial normal map. Specifically, the human body image under the camera perspective and the human body image under other camera perspectives, as well as the camera pose matrix corresponding to the camera perspective and the camera pose matrix corresponding to other camera perspectives are used to update and optimize the depth map and normal map to obtain the optimized depth map and normal map under the camera perspective. In the above manner, the optimized depth map and normal map under all camera perspectives can be obtained.

[0061] After that, the optimized depth map and normal map of the person to be estimated under multiple camera perspectives can be compressed and output, that is, the depth map and normal map can be stored according to the camera perspective, and then the stored overall information is compressed and output. In this way, while ensuring the integrity and orderliness of the information, the number of files, transmission time and space occupancy can be greatly reduced, thereby improving the transmission efficiency, and enabling users / downstream tasks to obtain the depth information of the person to be estimated in the three-dimensional space more quickly, thus realizing complete human depth estimation, and ensuring the integrity, accuracy and effectiveness of human depth estimation.

[0062] Based on the above embodiment, based on the human body image under any camera perspective and the human body image under other camera perspectives, as well as the camera pose matrix corresponding to the camera perspective and the camera pose matrix corresponding to the other camera perspectives, the initial depth map and the initial normal map under the camera perspective are optimized to obtain the depth map and the normal map under the camera perspective, including: Determine the pixel information of the target point in the human body image under the camera's perspective, the pixel information of the projection point of the target point in the human body image under other camera's perspectives, and the angle between the camera light corresponding to other camera's perspectives and the normal of the projection point of the target point under other camera's perspectives; Based on the pixel information of the target point in the human body image under the camera's perspective, the pixel information and the angle of the projection point of the target point in the human body image under other camera's perspectives, as well as the camera pose matrix corresponding to the camera's perspective and the camera pose matrix corresponding to other camera's perspectives, the initial depth map and the initial normal map under the camera's perspective are optimized to obtain the depth map and the normal map under the camera's perspective.

[0063] Specifically, when optimizing the initial depth map and the initial normal map under any camera perspective, according to the characteristics of the above-mentioned three-dimensional human body model, that is, the pixels of the projection points of the same triangle patch under different camera perspectives should be consistent, the pixel information of the target point in the human body image under the camera perspective and the pixel information of the projection point of this target point in the human body image under other camera perspectives can be obtained first. The target point here can be all the points in the human body area in the human body image or some of the points. For different human body areas, such as important human body areas (such as face) and conventional human body areas (such as limbs), the distribution and number of target points can be the same or different.

[0064] At the same time, it is also necessary to determine the camera pose matrix corresponding to the camera perspective and the camera pose matrix corresponding to other camera perspectives, as well as the angle between the camera light corresponding to other camera perspectives and the normal of the projection point of the target point under other camera perspectives.

[0065] After that, the initial depth map and initial normal map under the camera's perspective can be optimized according to the angle, the camera pose matrix, and the pixel information of the target point and its corresponding projection point, so as to obtain the optimized depth map and normal map under the camera's perspective. Here, the camera pose matrix corresponding to the camera's perspective and the camera pose matrix corresponding to other camera perspectives can be used to project the target point under the camera's perspective to other camera perspectives. On this basis, the pixel information and angle of the target point under the camera's perspective, as well as the pixel information of the projection point of the target point in the human body image under other camera perspectives, are combined to optimize the initial depth map and initial normal map under the camera's perspective, so as to obtain a high-precision depth map and normal map under the camera's perspective.

[0066] Specifically, if the triangular patch of the 3D human body model SMPL is projected to multiple camera perspectives, the pixel area corresponding to the triangular patch in the human body image under multiple camera perspectives can be calculated. By subdividing the triangular patch into points one by one, the initial depth value of each point in the initial depth map can be obtained. , Indicates the point on this triangle In the The initial depth value under the camera's perspective. Since the pixel values ​​obtained by projecting the same triangle under different camera perspectives should be consistent, based on this, when optimizing the depth value of the triangle under each camera perspective, in order to ensure the accuracy of the projection, the direction of the light projection and the normal information of the points in the initial normal map can be used as a weighted method. Assuming the optimization The initial depth map of the camera, the number of cameras is The optimization process can be expressed by the following formula:

[0067] In the formula, Indicates that Points from the camera's perspective Project to From the camera's perspective, and Respectively The camera pose matrix corresponding to the camera view and The camera pose matrix corresponding to the camera view, express Camera view point The initial depth value of express Camera view point Pixel information, express Camera rays and points corresponding to the camera view exist The angle of the normal of the projection point from the camera's perspective, Indicate point exist Projection point from the camera's perspective Pixel information.

[0068] Different from the traditional method of randomly initializing the depth, the embodiment of the present invention uses a three-dimensional human body model to determine the initial depth map and initial normal map under each camera perspective. On this basis, a pixel-by-pixel depth and normal estimation method is proposed for optimization, which is more accurate than the traditional path-only method. In addition, the embodiment of the present invention also reasonably utilizes the normal information, which makes the iterative optimization process faster and the final depth map has higher accuracy.

[0069] Based on the above embodiment, the depth map and normal map of the person to be estimated under multiple camera perspectives are compressed and output, including: Determine the important human body regions in the human body image under any camera view; Based on the human body image under the camera's perspective and the human body images under other camera's perspectives, as well as the camera pose matrix corresponding to the camera's perspective and the camera pose matrix corresponding to other camera's perspectives, the depth map and normal map corresponding to the important human body area under the camera's perspective are optimized to obtain the target depth map and target normal map under the camera's perspective; The target depth map, target normal map and human body image under the camera's perspective are stored in bytes, and the stored contents under multiple camera perspectives are compressed and output.

[0070] Specifically, considering that the final output result in the traditional solution is only a depth map, and in the result presentation, important human body areas, such as faces, are consistent with conventional human body areas and are not clearly distinguished, resulting in the result being unclear and concise, and inconvenient for subsequent use, in an embodiment of the present invention, it is proposed that multiple optimization processes can be performed on areas with high detail requirements such as faces, so that they are more accurate and better meet usage requirements.

[0071] In detail, for any camera perspective, here we can first determine the important human body areas in the human body image under the camera perspective; the important human body areas here can be areas with relatively high accuracy requirements in the human depth estimation task, such as the head and face areas; they can also be areas specified by the user, or areas of interest to the user, such as the torso, limb joints, etc., which are not specifically limited in the embodiments of the present invention. Then, using the human body image under the camera perspective and the human body images under other camera perspectives, as well as the camera pose matrix corresponding to the camera perspective and the camera pose matrix corresponding to other camera perspectives, the depth map and normal map corresponding to the important human body area under the camera perspective are optimized to obtain the target depth map and target normal map under the camera perspective; the specific optimization process is basically the same as the optimization of the initial depth map described above, and will not be repeated here.

[0072] After that, the target depth map, target normal map and human body image under the camera's perspective can be stored in bytes. Figure 6 is an example diagram of the depth map storage process provided by the present invention, such as Figure 6 As shown, information such as a target depth map, a target normal map, a displacement matrix, a rotation matrix, and a human body image can be stored according to the camera perspective, and the stored content under multiple camera perspectives can be compressed and output.

[0073] See also Figure 6 In the embodiment of the present invention, the above information is stored in a ".bin format" for output. Specifically, the rotation matrix R (3 3) and the displacement matrix t(3 1), a value of 2 bytes, a total of 24 (3 3 2+3 1 2) bytes, and the subsequent bytes save the target depth map, target normal map, number information of the triangle facets, color information (human body image), etc., and so on. After storing the above information of all camera perspectives, the entire ".bin format" file is compressed and output. This not only ensures the integrity and order of information, but also greatly reduces the number of files and transmission time, improves transmission efficiency, and makes subsequent access more convenient.

[0074] Compared with the traditional depth estimation scheme based on image patch matching, in the embodiment of the present invention, the depth of each pixel can be estimated with higher accuracy through the pixel-by-pixel depth value optimization scheme; and, according to actual needs, important human body areas (which can be marked by the numbering information of the triangular facets) can be quickly re-optimized to make them more accurate. This is the advantage of the present invention, and this feature is not available in traditional methods. In addition, the embodiment of the present invention also proposes a corresponding storage format, which not only improves the storage efficiency of the depth map, but also enriches the storage content, saves all the effective information, and facilitates subsequent access.

[0075] The human body depth estimation device provided by the present invention is described below. The human body depth estimation device described below and the human body depth estimation method described above can be referred to each other.

[0076] Figure 7 Schematic diagram of the structure of the human body depth estimation device provided by the present invention. Figure 7 As shown, the device comprises: An acquisition unit 710 is used to acquire human body images of a person to be estimated under multiple camera viewing angles; A modeling unit 720 is configured to perform three-dimensional human body modeling based on the human body images under the multiple camera perspectives to obtain three-dimensional human body models corresponding to the human body images under the multiple camera perspectives; The estimation unit 730 is used to perform projection based on the three-dimensional human body model, and compress and output the depth map of the person to be estimated under the multiple camera perspectives obtained by the projection.

[0077] The human body depth estimation device provided by the present invention performs three-dimensional human body modeling through human body images under multiple camera perspectives to obtain the corresponding three-dimensional human body model in a fast parameterized manner. Based on the characteristic that the three-dimensional human body model constructed in this way will not have holes, it is projected to each camera perspective to obtain a depth map, which well solves the problem that the depth of weak texture areas cannot be restored in traditional solutions, resulting in holes in the depth map. The depth information of the weak texture areas can be effectively solved, thereby ensuring the accuracy and completeness of the human body depth estimation. Finally, the depth map is compressed and output, which not only reduces the number of output files, improves the transmission speed, reduces the transmission space occupied, but also improves the integrity and orderliness of the information, and realizes the effective simplification of the output results for easy access.

[0078] Based on the above embodiment, the modeling unit 720 is used to: Based on the human body images under the multiple camera perspectives, applying the posture reconstruction model to perform three-dimensional human body modeling to obtain a corresponding three-dimensional human body model; The posture reconstruction model is obtained based on sample body images of sample persons under multiple camera perspectives, sample pose information corresponding to the sample body images under the multiple camera perspectives, predicted body models corresponding to the sample body images under the multiple camera perspectives, and projection training of the predicted body model under the multiple camera perspectives.

[0079] Based on the above embodiment, the device further includes a training unit, which is used to: Performing human body detection and key point detection on the sample human body images under the multiple camera perspectives to obtain sample human body regions and corresponding sample pose information in the sample human body images under the multiple camera perspectives; Performing three-dimensional human body modeling based on the initial reconstruction model, obtaining a predicted human body model corresponding to the sample human body region under the multiple camera perspectives, and determining the projection of the predicted human body model under the multiple camera perspectives; Determining a regional reconstruction loss based on consistency between the projections under the multiple camera perspectives and the sample human body regions under the multiple camera perspectives; Determining a pose reconstruction loss based on consistency between sample pose information corresponding to the sample human body images under the multiple camera perspectives and predicted pose information corresponding to the projections under the multiple camera perspectives; Based on the region reconstruction loss and the posture reconstruction loss, the initial reconstruction model is iterated on parameters to obtain a posture reconstruction model.

[0080] Based on the above embodiment, the sample pose information includes coordinates and pose parameters of human joints of the sample person in the sample human image; the predicted pose information includes coordinates and pose parameters of human joints in the projection under the corresponding camera perspective; The training unit is used to: Determine the position reconstruction loss based on the consistency between the coordinates of the human joint points in the sample human body images under the multiple camera perspectives and the coordinates of the human joint points in the projections under the multiple camera perspectives; And / or, determining the posture reconstruction loss based on the posture parameters of the human joint points in the projections under the multiple camera perspectives, or the consistency between the posture parameters of the human joint points in the sample human body images under the multiple camera perspectives and the posture parameters of the human joint points in the projections under the multiple camera perspectives; The pose reconstruction loss is determined based on the position reconstruction loss and / or the pose reconstruction loss.

[0081] Based on the above embodiment, the estimation unit 730 is used for: Performing projection based on the three-dimensional human body model to obtain an initial depth map and an initial normal map of the person to be estimated under the perspectives of the multiple cameras; Based on the human body image under any camera perspective and the human body image under other camera perspectives, as well as the camera pose matrix corresponding to the camera perspective and the camera pose matrix corresponding to the other camera perspectives, optimizing the initial depth map and the initial normal map under the camera perspective to obtain the depth map and the normal map under the camera perspective; The depth map and the normal map of the person to be estimated under the multiple camera viewing angles are compressed and output.

[0082] Based on the above embodiment, the estimation unit 730 is used for: Determine pixel information of a target point in the human body image under the camera's perspective, pixel information of a projection point of the target point in the human body image under the other camera's perspective, and an angle between a camera ray corresponding to the other camera's perspective and a normal of the projection point of the target point under the other camera's perspective; Based on the pixel information of the target point in the human body image under the camera perspective, the pixel information of the projection point of the target point in the human body image under the other camera perspectives, the angle, and the camera pose matrix corresponding to the camera perspective and the camera pose matrix corresponding to the other camera perspectives, the initial depth map and the initial normal map under the camera perspective are optimized to obtain the depth map and the normal map under the camera perspective.

[0083] Based on the above embodiment, the estimation unit 730 is used for: Determine the important human body regions in the human body image under any camera view; Based on the human body image under the camera's perspective and the human body images under other camera's perspectives, as well as the camera pose matrix corresponding to the camera's perspective and the camera pose matrix corresponding to the other camera's perspectives, the depth map and normal map corresponding to the important human body area under the camera's perspective are optimized to obtain a target depth map and a target normal map under the camera's perspective; The target depth map, target normal map and human body image under the camera viewing angle are stored in bytes, and the stored contents under the multiple camera viewing angles are compressed and output.

[0084] Figure 8 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 8As shown, the electronic device may include: a processor 810, a communication interface 820, a memory 830 and a communication bus 840, wherein the processor 810, the communication interface 820 and the memory 830 communicate with each other through the communication bus 840. The processor 810 may call the logic instructions in the memory 830 to execute a method for estimating the depth of a human body, the method comprising: obtaining human body images of a person to be estimated under multiple camera perspectives; performing three-dimensional human body modeling based on the human body images under the multiple camera perspectives to obtain a three-dimensional human body model corresponding to the human body images under the multiple camera perspectives; performing projection based on the three-dimensional human body model, and compressing and outputting the depth map of the person to be estimated under the multiple camera perspectives obtained by projection.

[0085] In addition, the logic instructions in the above-mentioned memory 830 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.

[0086] On the other hand, the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the human depth estimation method provided by the above-mentioned methods, and the method includes: obtaining human body images of the person to be estimated under multiple camera perspectives; performing three-dimensional human body modeling based on the human body images under the multiple camera perspectives to obtain a three-dimensional human body model corresponding to the human body images under the multiple camera perspectives; performing projection based on the three-dimensional human body model, and compressing and outputting the depth map of the person to be estimated under the multiple camera perspectives obtained by the projection.

[0087] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the human body depth estimation method provided by the above-mentioned methods, the method comprising: obtaining human body images of a person to be estimated under multiple camera perspectives; performing three-dimensional human body modeling based on the human body images under the multiple camera perspectives to obtain a three-dimensional human body model corresponding to the human body images under the multiple camera perspectives; performing projection based on the three-dimensional human body model, and compressing and outputting the depth map of the person to be estimated under the multiple camera perspectives obtained by the projection.

[0088] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0089] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0090] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for estimating human body depth, characterized in that: include: Obtaining human body images of the person to be estimated under multiple camera perspectives; Performing three-dimensional human body modeling based on the human body images under the multiple camera perspectives to obtain three-dimensional human body models corresponding to the human body images under the multiple camera perspectives; Projection is performed based on the three-dimensional human body model, and depth maps of the person to be estimated obtained by the projection under the multiple camera viewing angles are compressed and output.

2. The method for estimating human body depth according to claim 1, characterized in that: The performing three-dimensional human body modeling based on the human body images under the multiple camera perspectives to obtain a three-dimensional human body model corresponding to the human body images under the multiple camera perspectives includes: Based on the human body images under the multiple camera perspectives, applying the posture reconstruction model to perform three-dimensional human body modeling to obtain a corresponding three-dimensional human body model; The posture reconstruction model is obtained based on sample body images of sample persons under multiple camera perspectives, sample pose information corresponding to the sample body images under the multiple camera perspectives, predicted body models corresponding to the sample body images under the multiple camera perspectives, and projection training of the predicted body model under the multiple camera perspectives.

3. The method for estimating human body depth according to claim 2, characterized in that: The posture reconstruction model is trained based on the following steps: Performing human body detection and key point detection on the sample human body images under the multiple camera perspectives to obtain sample human body regions and corresponding sample pose information in the sample human body images under the multiple camera perspectives; Performing three-dimensional human body modeling based on the initial reconstruction model, obtaining a predicted human body model corresponding to the sample human body region under the multiple camera perspectives, and determining the projection of the predicted human body model under the multiple camera perspectives; Determining a regional reconstruction loss based on consistency between the projections under the multiple camera perspectives and the sample human body regions under the multiple camera perspectives; Determining a pose reconstruction loss based on consistency between sample pose information corresponding to the sample human body images under the multiple camera perspectives and predicted pose information corresponding to the projections under the multiple camera perspectives; Based on the region reconstruction loss and the posture reconstruction loss, the initial reconstruction model is iterated on parameters to obtain a posture reconstruction model.

4. The method for estimating human body depth according to claim 3, characterized in that: The sample pose information includes coordinates and pose parameters of the human joints of the sample person in the sample human image; the predicted pose information includes coordinates and pose parameters of the human joints in the projection under the corresponding camera perspective; The determining of the pose reconstruction loss based on the consistency between the sample pose information corresponding to the sample human body images under the multiple camera perspectives and the predicted pose information corresponding to the projections under the multiple camera perspectives includes: Determine the position reconstruction loss based on the consistency between the coordinates of the human joint points in the sample human body images under the multiple camera perspectives and the coordinates of the human joint points in the projections under the multiple camera perspectives; And / or, determining the posture reconstruction loss based on the posture parameters of the human joint points in the projections under the multiple camera perspectives, or the consistency between the posture parameters of the human joint points in the sample human body images under the multiple camera perspectives and the posture parameters of the human joint points in the projections under the multiple camera perspectives; The pose reconstruction loss is determined based on the position reconstruction loss and / or the pose reconstruction loss.

5. The method for estimating human body depth according to any one of claims 1 to 4, characterized in that: The projecting based on the three-dimensional human body model and compressing and outputting the depth map of the person to be estimated under the multiple camera perspectives obtained by the projection includes: Performing projection based on the three-dimensional human body model to obtain an initial depth map and an initial normal map of the person to be estimated under the perspectives of the multiple cameras; Based on the human body image under any camera perspective and the human body image under other camera perspectives, as well as the camera pose matrix corresponding to the any camera perspective and the camera pose matrix corresponding to the other camera perspectives, optimizing the initial depth map and the initial normal map under any camera perspective to obtain the depth map and the normal map under any camera perspective; The depth map and the normal map of the person to be estimated under the multiple camera viewing angles are compressed and output.

6. The method for estimating human body depth according to claim 5, characterized in that: The method optimizes the initial depth map and the initial normal map under any camera perspective based on the human body image under any camera perspective and the human body image under other camera perspectives, and the camera pose matrix corresponding to any camera perspective and the camera pose matrix corresponding to the other camera perspectives, to obtain the depth map and the normal map under any camera perspective, including: Determine pixel information of a target point in the human body image under any camera perspective, pixel information of a projection point of the target point in the human body image under the other camera perspectives, and an angle between a camera ray corresponding to the other camera perspectives and a normal of the projection point of the target point under the other camera perspectives; Based on the pixel information of the target point in the human body image under any camera perspective, the pixel information of the projection point of the target point in the human body image under the other camera perspectives, the angle, and the camera pose matrix corresponding to any camera perspective and the camera pose matrix corresponding to the other camera perspectives, the initial depth map and the initial normal map under any camera perspective are optimized to obtain the depth map and the normal map under any camera perspective.

7. The method for estimating human body depth according to claim 5, characterized in that: The compressing and outputting the depth map and the normal map of the person to be estimated under the multiple camera perspectives includes: Determine the important human body regions in the human body image under any camera view; Based on the human body image under any camera perspective and the human body image under other camera perspectives, as well as the camera pose matrix corresponding to any camera perspective and the camera pose matrix corresponding to the other camera perspectives, the depth map and normal map corresponding to the important human body area under any camera perspective are optimized to obtain the target depth map and target normal map under any camera perspective; The target depth map, target normal map and human body image under any camera viewing angle are stored in bytes, and the stored contents under the multiple camera viewing angles are compressed and output.

8. A human body depth estimation device, characterized in that: include: An acquisition unit, used for acquiring human body images of a person to be estimated under multiple camera viewing angles; A modeling unit, configured to perform three-dimensional human body modeling based on the human body images under the multiple camera perspectives, and obtain a three-dimensional human body model corresponding to the human body images under the multiple camera perspectives; The estimation unit is used to perform projection based on the three-dimensional human body model, and compress and output the depth map of the person to be estimated under the perspectives of the multiple cameras obtained by the projection.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the human body depth estimation method according to any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for estimating human depth as claimed in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • High-precision face reconstruction system and method based on light polarization reflectivity, terminal and medium

    CN115482329A

  • Three-dimensional human body shape reconstruction method and system based on multi-view projection contour consistency constraint

    CN116152432A

  • Three-dimensional object classification method and system based on descriptor and AdaBoost algorithm

    CN116434220A

  • Monocular three-dimensional human body reconstruction method and device and electronic equipment

    CN116486009A

  • Multi-view three-dimensional human body posture estimation method and device and storage medium

    CN117911494A