Single-frame image 3D human pose estimation and reconstruction method based on projection loss constraint
Patent Information
- Application Number
- CN202211024246.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-24
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2042-08-24
AI Technical Summary
[0074]1、利用SMPL模型在重建3D人体过程中,考虑了关键点投影损失、人体轮廓投影损失、多角度顶点投影损失、关节弯曲的先验惩罚,减少了重建的3D人体的扭曲现象;
Smart Images

Figure CN115393512B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of three-dimensional human body reconstruction technology, specifically to a method for 3D human body pose estimation and reconstruction based on single-frame images with projection loss constraints. Background Technology
[0002] Human 3D pose estimation and reconstruction technology can obtain 3D parameter information from single or multiple frames of 2D images and further reconstruct a 3D human body, with wide applications in sports, film and television, security, and many other fields. Since recovering 3D information from a 2D image is an underdetermined problem, existing technologies typically utilize various additional information for the solution, such as 3D human body reconstruction using consecutive frame images or multi-view images. Currently, 3D human body reconstruction from a single frame image captured by a monocular camera remains a challenging task, usually requiring iterative optimization of the underdetermined problem. Constraints used include human contour constraints and 2D keypoint position constraints, but these optimizations are prone to local optima, leading to distortion and interlacing problems in the reconstructed 3D human body.
[0003] The Skinned Multi-Person Linear Model (SMPL) is a vertex-based parametric 3D human body model. It generates a 3D human body mesh driven by shape and pose parameters. Using this model, the reconstruction optimization problem can be transformed into solving the SMPL model parameters, thus obtaining relatively more accurate human body reconstruction results more quickly. However, existing methods for generating 3D human bodies using the SMPL model still suffer from problems such as human body distortion and positional offset, resulting in a lack of spatial realism in the reconstructed 3D human body within the scene. Summary of the Invention
[0004] This invention aims to further reduce the distortion and positional shift of the 3D human body reconstructed by the SMPL model and improve the spatial realism of the reconstructed 3D human body in the scene. It provides a single-frame image 3D human body pose estimation and reconstruction method based on projection loss constraints.
[0005] To achieve this objective, the present invention adopts the following technical solution:
[0006] A method for 3D human pose estimation and reconstruction based on single-frame image with projection loss constraint is provided, the steps of which include:
[0007] S1, using a 2D pose estimation model to perform 2D human key points and 2D human contour recognition on a single frame image captured by a monocular camera, to obtain 2D human key points K and 2D human contour map S.
[0008] S2, using the 2D human body key points K and the 2D human body contour map S as inputs to the SMPL parameter estimation model, the shape parameter β, pose parameter θ and camera parameter T are inferred.
[0009] S3, the SMPL model is driven by β and θ to reconstruct a 3D human body, and the reconstructed 3D human body is projected using the camera parameters T.
[0010] Preferably, the SMPL parameter estimation model includes a shape estimation module, a pose estimation module, and a camera estimation module that are parallel to each other. The shape estimation module includes a first shape estimation module and a second shape estimation module connected in parallel with the first shape estimation module.
[0011] The first body size shape estimation module takes the 2D human body contour image S as input, predicts and outputs an intermediate feature vector f, and inputs f into the second body size shape estimation module to predict and output the body size shape parameter β;
[0012] The pose estimation module takes the 2D human key points K as input and predicts and outputs the pose parameters θ.
[0013] The camera estimation module takes the coordinates K of the 2D human key points as input and predicts and outputs the camera parameters T.
[0014] Preferably, the SMPL parameter estimation model includes a shape estimation module, a pose estimation module, and a camera estimation module that are sequentially connected to each other. The shape estimation module includes a first shape estimation module and a third shape estimation module that is sequentially connected to the first shape estimation module.
[0015] The first shape estimation module takes the 2D human body contour image S as input and predicts the intermediate feature vector f.
[0016] The pose estimation module takes the coordinates K of the 2D human key points as input and predicts and outputs the pose parameters θ.
[0017] The camera estimation module takes the 2D human key points K as input and predicts and outputs the camera parameters T.
[0018] The third shape estimation module uses the intermediate feature vector f, the pose parameter θ, and the camera parameter T as inputs to predict and output the shape parameter β.
[0019] Preferably, the method by which the pose estimation module estimates the parameter θ includes the following steps:
[0020] A1: The 2D human key points K with a dimension of 2×M, where M represents the number of human key points, are used as the input of the pose estimation module. After feature extraction by the fully connected layer FC1 in the pose estimation module, the first feature with a dimension of 512 is output.
[0021] A2: The output of the fully connected layer FC1 is used as the input of the convolutional layer Conv1 in the pose estimation module. The pose estimation module first transforms the first feature of 512 dimensions into a feature map of 1×1×512 dimensions, and then inputs it into the convolutional layer Conv1 for further feature extraction, and outputs a second feature map of 1×1×512 dimensions. The convolutional kernel size of the convolutional layer Conv1 is 1×1, the number of channels is 1, and the stride is 1.
[0022] A3: The pose estimation module transforms the output of the convolutional layer Conv1 into 512 dimensions, and then inputs it into the fully connected layer FC2. The fully connected layer FC2 outputs a 256-dimensional third feature. The fully connected layer FC1 has 256 nodes.
[0023] A4: The output of the fully connected layer FC2 is used as the input of the fully connected layer FC3 in the pose estimation module. The fully connected layer FC3 outputs a 72-dimensional fourth feature. The fully connected layer FC3 has 72 nodes, and the final output of the 72 dimensions is the predicted parameter θ.
[0024] The method steps for the camera estimation module to estimate the camera parameter T include:
[0025] B1: The 2D human keypoints K with dimensions of 2×M are used as the input of the camera estimation module. After feature extraction by the fully connected layer FC4 in the camera estimation module, the fifth feature with 512 dimensions is output. The fully connected layer FC4 has 512 nodes, and M represents the number of human keypoints.
[0026] B2: The output of the fully connected layer FC4 is used as the input of the fully connected layer FC5 in the camera estimation module. The fully connected layer FC5 outputs a 128-dimensional sixth feature and has 128 nodes.
[0027] B3: The output of the fully connected layer FC4 is used as the input of the fully connected layer FC6 in the camera estimation module. The fully connected layer FC6 outputs a 3-dimensional seventh feature. The fully connected layer FC6 has 3 nodes, and the final output 3-dimensional feature is the predicted camera parameter T.
[0028] The method steps for the first shape size estimation module to estimate the intermediate feature vector f include:
[0029] C1: The 2D human body contour map S with a size of 1×256×256 is used as the input of the convolutional layer Conv2 in the first body size estimation module. The convolutional layer Conv2 has a kernel size of 7×7, 32 channels, a stride of 2, and padding of 3. The convolutional layer Conv2 outputs an eighth feature map with a dimension of 32×128×128.
[0030] C2: The output of the convolutional layer Conv2 is the input of the convolutional layer Conv3 in the first shape estimation module. The convolutional layer Conv3 has a kernel size of 3×3, 128 channels, a stride of 2, and padding of 1. The convolutional layer Conv3 outputs a ninth feature map with dimensions of 128×64×64.
[0031] C3: The output of the convolutional layer Conv3 is the input of the convolutional layer Conv4 in the first shape estimation module. The convolutional layer Conv4 has a kernel size of 3×3, 512 channels, a stride of 2, and padding of 1. The output of the convolutional layer Conv4 is a tenth feature map with a dimension of 512×32×32.
[0032] C4: The output of the convolutional layer Conv4 is the input of the max pooling layer MaxPool in the second shape size estimation module. The max pooling layer MaxPool has a region size of 32×32 and an output of 512×1. After transformation, the intermediate feature vector f of 512 dimensions is generated.
[0033] The method steps for estimating parameter β in the second shape size estimation module, which is connected in parallel with the first shape size estimation module, include:
[0034] M1: Input the 512-dimensional intermediate feature vector f output in step C4 into the fully connected layer FC7 in the second shape size estimation module 2, and output the eleventh feature with a dimension of 256. The fully connected layer FC7 has 256 nodes.
[0035] M2: The output of the fully connected layer FC7 is the input of the fully connected layer FC8 in the second shape estimation module. The fully connected layer FC8 outputs the twelfth feature with a dimension of 10. The fully connected layer FC8 has 10 nodes, and the final output of the 10 dimensions is the predicted parameter β.
[0036] The method steps for estimating parameter β in the third shape size estimation module, which is serially connected to the first shape size estimation module, include:
[0037] N1: The 512-dimensional intermediate feature vector f output in step C4, the parameter θ output by the fully connected layer FC3 in the pose estimation module, and the parameter T output by the fully connected layer FC6 in the camera estimation module are horizontally concatenated in the order of (f, θ, T) to generate a 587-dimensional feature. This feature is then input into the fully connected layer FC9 in the serial third shape estimation module, and the output is a thirteenth feature with a dimension of 256. The fully connected layer FC9 has 256 nodes.
[0038] N2: The output of the fully connected layer FC9 is the input of the fully connected layer FC10 in the serial third shape size estimation module. The fully connected layer FC10 outputs the fourteenth feature with a dimension of 10. The fully connected layer FC10 has 10 nodes, and the final output of the 10 dimensions is the predicted parameter β.
[0039] Preferably, the objective function of the SMPL parameter estimation model is expressed by the following formula (1):
[0040] L = L J +λ S L S +λ V L V +λ a L a Formula (1)
[0041] In formula (1), L J This indicates the keypoint projection loss after 3D human body reconstruction.
[0042] L S This indicates the loss of human body contour projection after 3D human body reconstruction.
[0043] L V This indicates that for the reconstructed 3D human body, the vertex projection loss of each triangular mesh used to characterize the 3D human body is calculated.
[0044] L a Indicates the prior penalty for joint bending;
[0045] λ S , λ V , λ a L respectively S L V L a The loss weighting coefficient.
[0046] As a preferred option, L J It is calculated using the following formula (2):
[0047]
[0048] In formula (2), K i This represents the coordinates of the i-th key point in the 2D human body key point coordinate K;
[0049] J proj-i This indicates that the 3D human body reconstructed by the SMPL model is related to K. i The projected coordinates of key points that have a one-to-one correspondence;
[0050] M represents the number of key points in the 2D human body keypoint K.
[0051] As a preferred option, L S It is calculated using the following formula (3):
[0052]
[0053] In formula (3), S represents the 2D human body contour map S;
[0054] S proj This represents the projection of the human body outline after 3D human reconstruction.
[0055] As a preferred option, L V It is calculated using the following formula (4):
[0056]
[0057] In formula (4), c represents the projection camera set at a specified angle c;
[0058] “3” represents three projection cameras at different angles;
[0059] μ c This indicates that for the reconstructed 3D human body, the vertex projection loss L of each triangular mesh used to characterize the 3D human body is calculated. V hour, The weight it occupies;
[0060] V proj-cjThis represents the projection coordinates of the j-th vertex calculated by projecting the reconstructed 3D human body onto a projection camera set at the specified angle c.
[0061] V gt-cj This represents the projected coordinates of the j-th vertex obtained by projecting the 3D reconstructed human body ground truth onto a projection camera set at the specified angle c.
[0062] Preferably, the positional angle between each pair of the three different projection cameras set at the specified angle c is 120°.
[0063] As a preferred option, L a It is calculated using the following formula (5):
[0064] L a =∑ i∈(eblows,knees) exp(θ i ) Formula (5)
[0065] In formula (5), θ i This represents the i-th value of the pose parameter;
[0066] eblows represents the set of elbow joint positions in the 72-dimensional pose parameter θ;
[0067] knees represents the set of knee joint positions in the 72-dimensional pose parameter θ;
[0068] Preferably, when reconstructing the 3D human body, the world coordinate system is converted to the camera coordinate system using the camera parameters T, and the rotation matrix when converting the world coordinate system to the camera coordinate system is set as the identity matrix, which is expressed by the following formula (8):
[0069]
[0070] v x v y v z These represent the coordinates of the vertex v of the 3D human body reconstructed in the world coordinate system on the x-axis, y-axis, and z-axis, respectively.
[0071] t x t y t z These represent the values of the camera parameter T, and respectively represent the translation coefficients of the vertex v in the x-axis, y-axis, and z-axis directions;
[0072] v′ x v′ y v′ z These represent the coordinates of vertex v in the camera coordinate system on the x-axis, y-axis, and z-axis, respectively.
[0073] The present invention has the following beneficial effects:
[0074] 1. In the process of reconstructing the 3D human body using the SMPL model, key point projection loss, human body contour projection loss, multi-angle vertex projection loss, and prior penalty for joint bending are considered, which reduces the distortion phenomenon of the reconstructed 3D human body.
[0075] 2. In the process of reconstructing a 3D human body using the SMPL model, the connection relationship between the pose estimation module, the shape module, and the camera estimation module was considered. Two model structures, parallel and serial, were constructed. The results of the two models were compared, and the better structure was selected.
[0076] 3. Using 2D human keypoints K and 2D human contour map S as inputs to the SMPL parameter estimation model, the SMPL parameter estimation model avoids the influence of features such as image background in the original 2D human RGB image on the estimation results when estimating the shape parameter β, pose parameter θ and camera parameter T. This makes the parameter estimation accuracy of the SMPL parameter estimation model higher, thereby improving the stability of the SMPL model in reconstructing 3D human bodies. Attached Figure Description
[0077] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments of the present invention will be briefly described below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.
[0078] Figure 1 This is a diagram illustrating the implementation steps of a single-frame image 3D human pose estimation and reconstruction method based on projection loss constraints provided in an embodiment of the present invention.
[0079] Figure 2 This is a network structure diagram for three-dimensional human pose estimation according to an embodiment of the present invention;
[0080] Figure 3 This is a schematic diagram illustrating the angular enhancement of sample data;
[0081] Figure 4 This is a network model structure diagram for 2D pose estimation and semantic segmentation tasks;
[0082] Figure 5 This is a model structure diagram of the pose estimation module and the camera estimation module in the SMPL parameter estimation model;
[0083] Figure 6This is a model structure diagram of the shape estimation module, which includes both parallel and serial structures, in the SMPL parameter estimation model.
[0084] Figure 7 This is a schematic diagram of the parallel connection of modules in the SMPL parameter estimation model;
[0085] Figure 8 This is a schematic diagram of the sequential connection of modules in the SMPL parameter estimation model;
[0086] Figure 9 It is a layout diagram of three simulated cameras that project onto the vertices of each triangular mesh representing the same 3D human body at different specified angles c.
[0087] Figure 10 This is an experimental comparison diagram of the 3D human body reconstructed from β and θ estimated by the SMPL parameter estimation model with a parallel structure based on a single frame image captured by a camera, and the 3D human body reconstructed from β and θ estimated by the SMPL parameter estimation model with a serial structure based on the same single frame image captured by the camera. Detailed Implementation
[0088] The technical solution of the present invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0089] The accompanying drawings are for illustrative purposes only and are schematic diagrams, not actual images. They should not be construed as limiting the scope of this patent. To better illustrate the embodiments of the present invention, some parts in the drawings may be omitted, enlarged, or reduced, and do not represent the actual dimensions of the product. It is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.
[0090] In the accompanying drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components. In the description of the present invention, it should be understood that if terms such as "upper," "lower," "left," "right," "inner," and "outer" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, they are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, the terms used to describe positional relationships in the drawings are only for illustrative purposes and should not be construed as limiting the present patent. For those skilled in the art, the specific meaning of the above terms can be understood according to the specific circumstances.
[0091] In the description of this invention, unless otherwise explicitly specified and limited, the term "connection" or similar designation indicating a connection between components should be interpreted broadly. For example, it can refer to a fixed connection, a detachable connection, or an integral part; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium; it can refer to the internal communication between two components or the interaction between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0092] The present invention provides a method for 3D human pose estimation and reconstruction based on projection loss constraints for single-frame images, such as... Figure 2 As shown, it mainly consists of two parts: 1. Generating 2D human keypoints and 2D human contours; 2. SMPL model parameter estimation. In application, a single-frame human image captured by a monocular camera is first input into the 2D pose estimation model to infer the 2D human keypoints and 2D human contours. Then, using the SMPL parameter estimation model with the 2D human keypoint coordinates K and the 2D human contour image S as input, the model infers the 10-dimensional (the 10-dimensional parameters jointly control the human's weight, height, and head-to-body ratio; for example, the first dimension represents the overall weight and size of the human body) shape parameters β, 72-dimensional (each of the 72 dimensions represents the rotation angle of a human body or joint; for example, the first 3 dimensions represent the angles of rotation around the x, y, and z axes, respectively, and the 4th to 6th and 7th to 9th dimensions represent the angles of rotation of the left and right hip joints around the x, y, and z axes, respectively) pose parameters θ and 3-dimensional (t) parameters required for reconstructing the 3D human body. x t y t z , where β and θ represent the camera parameters T (the translation coefficients of vertices v of the triangular mesh representing the 3D human body in the x, y, and z axes, respectively). Finally, the SMPL model is used to reconstruct the 3D human body by driving β and θ, and a 2D image of the 3D human body is reconstructed by projecting T.
[0093] In this embodiment, the training of the model is divided into two independent sub-processes: 2D pose estimation model training and SMPL parameter estimation model training. HRNet (High-Resolution Net) is a neural network for pose estimation of a single human body. It consists of several modules, including convolution, BatchNorm, and ReLU, forming layers 1, stage 2, stage 3, and stage 4. UDP (Unbiased Data Processing) optimizes the model by adding unbiased data processing to HRNet, exhibiting superior performance in 2D keypoint detection. However, due to its lack of semantic segmentation capabilities, it cannot be directly applied to semantic segmentation of a single-frame human image in this application to obtain a 2D human contour map S. To address this issue, this application adds a semantic segmentation feature selection branch after stage 4 of the UDP model structure, resulting in... Figure 4 The UDP-S model shown fixes all network parameters except for the semantic segmentation branch, and updates the semantic segmentation branch with a small learning rate. Figure 4 The training is performed on the solid line box portion of the image.
[0094] The second sub-training process involves training the SMPL parameter estimation network PENet (Parameters Estimation Network). PENet includes a pose estimation module, a shape estimation module, and a camera estimation module. The pose estimation module consists of one convolutional layer and three fully connected layers. It takes the 2D human keypoint coordinates K as input and predicts the output pose parameters θ. The network structure of the pose estimation module is shown below. Figure 5 ;
[0095] The camera estimation module consists of three fully connected layers. It takes the coordinates K of 2D human keypoints as input and predicts the camera parameters T as output. The network structure of the camera estimation module is shown below. Figure 5 ;
[0096] The shape estimation module (including a first shape estimation module and a second shape estimation module connected in parallel, or a first shape estimation module and a third shape estimation module connected in serial) consists of three convolutional layers, one max pooling layer, and two fully connected layers. It takes a 2D human contour image S as input and predicts the output shape parameter β. The network structure of the shape estimation module is shown below. Figure 6 The method for estimating parameter θ in the pose estimation module includes the following steps:
[0097] A1: Take the 2D human keypoints K with dimension 2×M, where M represents the number of human keypoints, as the input to the pose estimation module. After feature extraction by the fully connected layer FC1 in the pose estimation module, the first feature with dimension 512 is output.
[0098] A2: The output of the fully connected layer FC1 is used as the input of the convolutional layer Conv1 in the pose estimation module. The pose estimation module first transforms the 512-dimensional first feature into a 1×1×512-dimensional feature map, and then inputs it into the convolutional layer Conv1 for further feature extraction, outputting a 1×1×512-dimensional second feature map. The convolutional kernel size of the convolutional layer Conv1 is 1×1, the number of channels is 1, and the stride is 1.
[0099] A3: The pose estimation module transforms the output of the convolutional layer Conv1 into 512 dimensions, and then inputs it into the fully connected layer FC2. The fully connected layer FC2 outputs a 256-dimensional third feature. The fully connected layer FC1 has 256 nodes.
[0100] A4: The output of the fully connected layer FC2 is used as the input of the fully connected layer FC3 in the pose estimation module. The fully connected layer FC3 outputs a 72-dimensional fourth feature. The fully connected layer FC3 has 72 nodes, and the final output of the 72 dimensions is the predicted parameter θ.
[0101] The steps involved in estimating camera parameters T using the camera estimation module include:
[0102] B1: The 2D human keypoints K with dimensions of 2×M are used as the input of the camera estimation module. After feature extraction by the fully connected layer FC4 in the camera estimation module, the fifth feature with dimensions of 512 is output. The fully connected layer FC4 has 512 nodes, and M represents the number of human keypoints.
[0103] B2: The output of the fully connected layer FC4 is used as the input of the fully connected layer FC5 in the camera estimation module. The fully connected layer FC5 outputs a 128-dimensional sixth feature and has 128 nodes.
[0104] B3: The output of the fully connected layer FC4 serves as the input to the fully connected layer FC6 in the camera estimation module. The fully connected layer FC6 outputs a 3-dimensional seventh feature. The fully connected layer FC6 has 3 nodes, and the final output 3-dimensional feature is the predicted camera parameter T.
[0105] The method steps for the first shape estimation module to estimate the intermediate feature vector f include:
[0106] C1: The 2D human contour map S with a size of 1×256×256 is used as the input of the convolutional layer Conv2 in the first body size estimation module. The convolutional layer Conv2 has a kernel size of 7×7, 32 channels, a stride of 2, and padding of 3. The output of the convolutional layer Conv2 is the eighth feature map with a dimension of 32×128×128.
[0107] C2: The output of convolutional layer Conv2 is the input of convolutional layer Conv3 in the first shape estimation module. Convolutional layer Conv3 has a kernel size of 3×3, 128 channels, a stride of 2, and padding of 1. Convolutional layer Conv3 outputs a ninth feature map with dimensions of 128×64×64.
[0108] C3: The output of convolutional layer Conv3 is the input of convolutional layer Conv4 in the first shape estimation module. Convolutional layer Conv4 has a kernel size of 3×3, 512 channels, a stride of 2, and padding of 1. Convolutional layer Conv4 outputs a tenth feature map with dimensions of 512×32×32.
[0109] C4: The output of the convolutional layer Conv4 is the input to the max pooling layer MaxPool in the second shape estimation module. The max pooling layer MaxPool has a region size of 32×32 and an output of 512×1. After transformation, it generates a 512-dimensional intermediate feature vector f.
[0110] The method steps for the second shape estimation module, which is connected in parallel with the first shape estimation module, to estimate the parameter β include:
[0111] M1: Input the 512-dimensional intermediate feature vector f output in step C4 into the fully connected layer FC7 in the second shape size estimation module, and output the eleventh feature with a dimension of 256. The fully connected layer FC7 has 256 nodes.
[0112] M2: The output of the fully connected layer FC7 is the input of the fully connected layer FC8 in the second shape estimation module. The fully connected layer FC8 outputs the twelfth feature with a dimension of 10. The fully connected layer FC8 has 10 nodes, and the final output of the 10 dimensions is the predicted parameter β.
[0113] The method steps for the third shape estimation module, which is serially connected to the first shape estimation module, to estimate the parameter β include:
[0114] N1: The 512-dimensional intermediate feature vector f output in step C4, the parameter θ output by the fully connected layer FC3 in the pose estimation module, and the parameter T output by the fully connected layer FC6 in the camera estimation module are horizontally concatenated in the order of (f, θ, T) to generate a 587-dimensional feature. This feature is then input into the fully connected layer FC9 in the serial third shape estimation module, and the output is a thirteenth feature with a dimension of 256. The fully connected layer FC9 has 256 nodes.
[0115] N2: The output of the fully connected layer FC9 is the input of the fully connected layer FC10 in the serial third shape estimation module. The fully connected layer FC10 outputs the fourteenth feature with a dimension of 10. The fully connected layer FC10 has 10 nodes, and the final output of the 10 dimensions is the predicted parameter β.
[0116] After obtaining the parameters θ, β, and T, the SMPL model is used to reconstruct the 3D human body using β and θ, generating a triangular mesh M with N = 6890 vertices. Then, M is projected onto T to obtain the keypoint projection coordinates J. proj ∈R 2×M (M two-dimensional coordinates in real space), vertex projection coordinates V proj ∈R 2×N (N two-dimensional coordinates in real space) and the projection of the human body contour S proj .
[0117] The core technology of this application lies in improving the correlation of parameters β, θ, and T by constructing both parallel and serial structures, and improving the loss function of the SMPL model to reconstruct the 3D human body, making the SMPL model training process more stable, convergence faster, and the reconstructed 3D human body more spatially realistic.
[0118] To ensure training effectiveness, the following technical measures are adopted in this application:
[0119] 1. Enhance human body key points and human body contour data.
[0120] In this application, Figure 2 The input to the PENet network shown is not the RGB human image itself captured by a monocular camera, but rather the 2D human keypoint coordinates K and 2D human contour map S predicted by the UDP-S model. The effective effect of this is that when the 2D human keypoint coordinates K and 2D human contour map S are used as inputs to the SMPL parameter estimation model, the influence of features such as the image background in the original 2D human RGB image on the estimation results is avoided, resulting in higher parameter estimation accuracy of the SMPL parameter estimation model and thus improving the stability of the SMPL model in reconstructing the 3D human body.
[0121] This embodiment uses publicly available 3D human body datasets UP3D (Unite the People-3D body), 3DPW (3D POSES IN THE WILD DATASET), and MTP (Mimic-The-Pose) as samples to train PENet. These datasets record the β and θ parameters of different body shapes, but the overall data volume is limited and the pose variation is relatively small. To solve this problem, this application adopts a data augmentation method to improve data diversity and robustness. The specific method is as follows: First, a 256×256 frame simulated camera is set, and the camera intrinsic parameters are fixed. The world coordinates of the simulated camera coincide with the camera coordinates, that is, the rotation matrix is an identity matrix. In the SMPL parameters, the first three dimensions of the pose parameter θ are global orient, representing the global rotation coefficient of the human body. The orientation and angle of the human body relative to the camera can be rotated by changing the global orient. The β value of different body shapes can be randomly changed. In this way, by transforming and combining the β and θ of each sample, a new human body is generated to obtain V. gt Using the human body's vertex ground truth as the rendering projection achieves data augmentation. A new human body rotated 180° is obtained using global rotation coefficients. Figure 3 As shown.
[0122] 2. Correlation between Pose and Shape parameters in prediction
[0123] The PENet model (i.e., the SMPL parameter estimation model) consists of three modules. This application designs two different structures for the connection between the modules: parallel structure and serial structure.
[0124] Among them, parallel structures are as follows Figure 6 and Figure 7 As shown, the three modules are independent of each other, with each module predicting its corresponding parameters independently. Specifically, the 2D human contour map S serves as the input to the shape estimation module, used to predict β; the 2D human keypoint coordinates K serve as the input to the pose estimation module and the camera estimation module, predicting the outputs θ and T, respectively.
[0125] The serial structure is as follows Figure 6 and Figure 8As shown, the three modules are connected in series and share parameters. The area and shape of the 2D human contour map S are affected by different poses and camera distances. When the modules are parallel, the shape estimation module is sensitive to S, and S contains less information, making it difficult to handle complex human figures in scenes with occlusion or side profiles. The serial structure links the modules together, first determining the pose and human position, and then predicting the shape dimensions. Specifically, the serial connection is as follows: the shape estimation module outputs an intermediate feature vector f, which is horizontally concatenated with θ and T in the order (f, θ, T) and then input into the fully connected layer to predict the output shape dimension parameters β. The serial structure is as follows. Figure 6 As shown.
[0126] 3. Enhance loss function constraints
[0127] The loss function calculation for the PENet model includes: calculating the projection loss using the 2D human keypoint coordinates K and the 2D human contour map S as self-supervised information, and penalizing abnormal poses based on joint priors. In this application, the objective function minimized during training is:
[0128] L = L J +λ S L S +λ V L V +λ a L a Formula (1)
[0129] In formula (1), L J This indicates the keypoint projection loss after 3D human body reconstruction.
[0130] L S This indicates the loss of human body contour projection after 3D human body reconstruction.
[0131] L V This indicates that for the reconstructed 3D human body, the vertex projection loss of each triangular mesh used to characterize the 3D human body is calculated.
[0132] L a Indicates the prior penalty for joint bending;
[0133] λ S , λ V , λ a L S L V L a The loss weighting coefficient.
[0134] L J It is calculated using the following formula (2):
[0135]
[0136] In formula (2), K i This represents the coordinates of the i-th keypoint in the 2D human body keypoint coordinate system K.
[0137] J proj-i This indicates that on the 3D human body reconstructed from the SMPL model, there is K i The projected coordinates of key points that have a one-to-one correspondence;
[0138] M represents the number of key points in the 2D human body key point coordinates K.
[0139] L S It is calculated using the following formula (3):
[0140]
[0141] In formula (3), S represents the 2D human body outline diagram S;
[0142] S proj This represents the projection of the human body outline after 3D human reconstruction.
[0143] Because there is no one-to-one correspondence between two-dimensional pose and three-dimensional pose, L J The constraint on the body pose remains ambiguous, affecting model convergence. To improve convergence speed and reconstruction accuracy, this application introduces a vertex projection coordinate loss constraint L into the loss function. V This constraint sets one main projection camera and two auxiliary projection cameras. Figure 6 In this configuration, one camera is designated as the primary projection camera, and the other two are designated as auxiliary cameras. The three cameras are angled together at 120°, and each camera is assigned a weight μ corresponding to the calculation of the vertex projection loss. c This application reduces the error caused by missing depth information by calculating the vertex projection coordinate loss from the camera c at different angles.
[0144] L V It is calculated using the following formula (4):
[0145]
[0146] In formula (4), c represents the projection camera at a specified angle, and the vertex projection loss of each triangular mesh used to characterize the 3D human body is calculated.
[0147] "3" represents three different projection cameras c at specified angles; as a preferred option, such as Figure 9 As shown, the angle between any two of the three different specified angles c is 120°.
[0148] μ c This represents the weight of the vertex projection loss calculated by projecting the reconstructed 3D human body onto camera c at a specified angle.
[0149] V proj-cj This represents the projected coordinates of the j-th vertex calculated by projecting the reconstructed 3D human body onto camera c at a specified angle.
[0150] V gt-cj This represents the projected coordinates of the j-th vertex calculated by projecting the 3D reconstructed human body ground truth onto camera c at a specified angle.
[0151] L a Representing simple a priori assumptions about joint flexion, constraining the unnatural rotational directions of the elbow and knee, L a It is calculated using the following formula (5):
[0152] L a =∑ i∈(eblows,knees) exp(θ i ) Formula (5)
[0153] In formula (5), θ i This represents the i-th value of the Pose parameter;
[0154] eblows represents the set of elbow joint positions in the 72-dimensional pose parameter θ;
[0155] knees represents the set of knee joint positions in the 72-dimensional pose parameter θ.
[0156] 4. Projecting predicted camera parameters
[0157] The reconstructed human body, in the world coordinate system, needs to be converted to the camera coordinate system. The 3D human body uses the camera estimation parameters T to convert the world coordinate system to the camera coordinate system. The rotation matrix for the conversion is set as the identity matrix, and is expressed by the following formula (6):
[0158]
[0159] v x u y v z These represent the coordinates of vertex v of the reconstructed 3D human body in the world coordinate system on the x-axis, y-axis, and z-axis, respectively.
[0160] t x t y t and t represent the values of the camera parameter T, which are the translation coefficients of vertex v in the x-axis, y-axis, and z-axis directions, respectively.
[0161] v′ x v′ y v′ z These represent the coordinates of vertex v on the x-axis, y-axis, and z-axis after transformation to the camera coordinate system.
[0162] The following are experimental verification and comparison results of the technical effects achieved by the single-frame image 3D human pose estimation and reconstruction method based on camera parameter constraints provided in this embodiment:
[0163] 1. Comparison of the beneficial effects of parallel and serial structures within the SMPL parameter estimation model.
[0164] Figure 10 The image shows a comparison of the results of 3D human reconstruction using a set of keyframes of a moving human body captured by a camera from a distance to a close distance, and then using SMPL parameter estimation models with parallel and serial structures respectively. Figure 10 The solid box in the middle selects the 3D human body reconstructed from two different angles by the SMPL model, which is driven by the SMPL parameter estimation model with a parallel structure, based on β and θ estimates. Figure 10 The dashed box highlights the 3D human figures reconstructed from the same two different angles using the SMPL parameter estimation model driven by β and θ estimates from the serial structure SMPL model. The results show that the serial structure reconstruction exhibits better consistency in body size, while the parallel structure reconstruction results in inconsistent body proportions, with abrupt size changes during actions such as bending or turning. Regarding pose reconstruction, the range of joint flexion and body tilt angles during human movement significantly impact the parallel structure reconstruction, while the serial structure reconstruction is smoother and more consistent with the pose of the 2D human image. This is because during model training, the parameters output by each module in PENet collectively influence the projection loss, and parameter updates are mutually constrained. In contrast, the parallel structure has weaker parameter constraints, and the shape estimation module relies on the size and shape of the 2D human contour map S, leading to significant variations in body size and proportions with different poses and distances. The serial structure enhances the mutual constraints between parameters, resulting in greater stability in pose and shape estimation.
[0165] In summary, the single-frame image 3D human pose estimation and reconstruction method based on projection loss constraints provided by the embodiments of the present invention, such as... Figure 1 As shown, the steps include:
[0166] S1, using a 2D pose estimation model to perform 2D human key points and 2D human contour recognition on a single frame image captured by a monocular camera, to obtain 2D human key points K and 2D human contour map S.
[0167] S2, using 2D human keypoints K and 2D human contour map S as inputs to the SMPL parameter estimation model, infers the shape parameter β, pose parameter θ and camera parameter T;
[0168] S3 reconstructs a 3D human body by driving the SMPL model through β and θ, and projects the reconstructed 3D human body onto the camera parameter T.
[0169] It should be stated that the above-described specific embodiments are merely preferred embodiments of the present invention and the technical principles employed. Those skilled in the art should understand that various modifications, equivalent substitutions, and variations can be made to the present invention. However, such variations, as long as they do not depart from the spirit of the present invention, should be within the scope of protection of the present invention. Furthermore, some terminology used in this specification and claims is not limiting, but merely for ease of description.
Claims
1. A method for 3D human pose estimation and reconstruction based on single-frame image with projection loss constraints, characterized in that the steps are as follows: include: S1, using a 2D pose estimation model to perform 2D human keypoint and 2D human contour recognition on a single frame image acquired by a monocular camera, obtaining 2D human keypoints. and 2D human body outline ; S2, using the 2D human key points and the 2D human body outline diagram As input to the SMPL parameter estimation model, the shape parameters are inferred. Pose parameters and camera parameters ; S3, via , Drive the SMPL model to reconstruct a 3D human body and utilize the camera parameters. Project the reconstructed 3D human body; The SMPL parameter estimation model includes a shape estimation module, a pose estimation module, and a camera estimation module that operate in parallel. The shape estimation module includes a first shape estimation module and a second shape estimation module connected in parallel with the first shape estimation module. The first shape estimation module uses the 2D human body contour map. As input, predict the intermediate feature vector as output. f is input to the second shape estimation module, which predicts and outputs the shape parameters. ; The pose estimation module uses the 2D human key points As input, predict the output of the pose parameters. The camera estimation module uses the coordinates of the 2D human key points. As input, predict the output camera parameters. ; Alternatively, the SMPL parameter estimation model includes a shape estimation module, a pose estimation module, and a camera estimation module that are sequentially connected to each other. The shape estimation module includes a first shape estimation module and a third shape estimation module that is sequentially connected to the first shape estimation module. The first shape estimation module uses the 2D human body contour map. As input, predict the intermediate feature vector as output. ; The pose estimation module uses the coordinates of the 2D human key points. As input, predict the output of the pose parameters. ; The camera estimation module uses the 2D human key points As input, predict the output camera parameters. ; The third shape estimation module uses the intermediate feature vector f and the pose parameters. The camera parameter T is the input to predict the output shape size parameter. ; The pose estimation module estimates the parameters. The method includes the following steps: A1: The 2D human key points K with a dimension of 2×M, where M represents the number of human key points, are used as the input of the pose estimation module. After feature extraction by the fully connected layer FC1 in the pose estimation module, the first feature with a dimension of 512 is output. A2: The output of the fully connected layer FC1 is used as the input of the convolutional layer Conv1 in the pose estimation module. The pose estimation module first transforms the first feature of 512 dimensions into a feature map of 1×1×512 dimensions, and then inputs it into the convolutional layer Conv1 for further feature extraction, and outputs a second feature map of 1×1×512 dimensions. The convolutional kernel size of the convolutional layer Conv1 is 1×1, the number of channels is 1, and the stride is 1. A3: The pose estimation module transforms the output of the convolutional layer Conv1 into 512 dimensions, and then inputs it into the fully connected layer FC2. The fully connected layer FC2 outputs a 256-dimensional third feature, and the fully connected layer FC2 has 256 nodes. A4: The output of the fully connected layer FC2 serves as the input to the fully connected layer FC3 in the pose estimation module. The fully connected layer FC3 outputs a 72-dimensional fourth feature. The fully connected layer FC3 has 72 nodes, and the final 72-dimensional output is the predicted parameter. ; The camera estimation module estimates the camera parameters. The method steps include: B1: The 2D human keypoints K with dimensions of 2×M are used as the input of the camera estimation module. After feature extraction by the fully connected layer FC4 in the camera estimation module, the fifth feature with 512 dimensions is output. The fully connected layer FC4 has 512 nodes, and M represents the number of human keypoints. B2: The output of the fully connected layer FC4 is used as the input of the fully connected layer FC5 in the camera estimation module. The fully connected layer FC5 outputs a 128-dimensional sixth feature and has 128 nodes. B3: The output of the fully connected layer FC5 is used as the input of the fully connected layer FC6 in the camera estimation module. The fully connected layer FC6 outputs a 3-dimensional seventh feature. The fully connected layer FC6 has 3 nodes, and the final output 3-dimensional feature is the predicted camera parameter T. The method steps for the first shape size estimation module to estimate the intermediate feature vector f include: C1: The 2D human body contour map S with a size of 1×256×256 is used as the input of the convolutional layer Conv2 in the first body size estimation module. The convolutional layer Conv2 has a kernel size of 7×7, 32 channels, a stride of 2, and padding of 3. The convolutional layer Conv2 outputs an eighth feature map with a dimension of 32×128×128. C2: The output of the convolutional layer Conv2 is the input of the convolutional layer Conv3 in the first shape estimation module. The convolutional layer Conv3 has a kernel size of 3×3, 128 channels, a stride of 2, and padding of 1. The convolutional layer Conv3 outputs a ninth feature map with dimensions of 128×64×64. C3: The output of the convolutional layer Conv3 is the input of the convolutional layer Conv4 in the first shape estimation module. The convolutional layer Conv4 has a kernel size of 3×3, 512 channels, a stride of 2, and padding of 1. The output of the convolutional layer Conv4 is a tenth feature map with a dimension of 512×32×32. C4: The output of the convolutional layer Conv4 is the input of the max pooling layer MaxPool in the first shape size estimation module. The max pooling layer MaxPool has a region size of 32×32 and an output of 512×1. After transformation, the intermediate feature vector f of 512 dimensions is generated. The second shape estimation module, which is connected in parallel with the first shape estimation module, estimates parameters in the shape estimation module. The method steps include: M1: Input the 512-dimensional intermediate feature vector f output in step C4 into the fully connected layer FC7 in the second shape size estimation module 2, and output the eleventh feature with a dimension of 256. The fully connected layer FC7 has 256 nodes. M2: The output of the fully connected layer FC7 is the input of the fully connected layer FC8 in the second shape estimation module. The fully connected layer FC8 outputs a 10-dimensional twelfth feature. The fully connected layer FC8 has 10 nodes, and the final 10-dimensional output is the predicted parameter. ; The third shape estimation module, which is serially connected to the first shape estimation module, estimates the parameters. The method steps include: N1: The 512-dimensional intermediate feature vector f output in step C4, and the parameters output by the fully connected layer FC3 in the pose estimation module. The parameter T output by the fully connected layer FC6 in the camera estimation module is calculated according to (f, The 587-dimensional feature is generated by sequentially and horizontally concatenating the T) features. It is then input into the fully connected layer FC9 in the serial third shape size estimation module and outputs the thirteenth feature with a dimension of 256. The fully connected layer FC9 has 256 nodes. N2: The output of the fully connected layer FC9 is the input of the fully connected layer FC10 in the serial third shape estimation module. The fully connected layer FC10 outputs a fourteenth feature with a dimension of 10. The fully connected layer FC10 has 10 nodes, and the final 10-dimensional output is the predicted parameter. .
2. The method for 3D human pose estimation and reconstruction based on projection loss constraints in a single frame image according to claim 1, characterized in that, The objective function of the SMPL parameter estimation model is expressed by the following formula (1): In formula (1), This indicates the keypoint projection loss after 3D human body reconstruction. This indicates the loss of human body contour projection after 3D human body reconstruction. This indicates that for the reconstructed 3D human body, the vertex projection loss of each triangular mesh used to characterize the 3D human body is calculated. Indicates the prior penalty for joint bending; They represent The loss weighting coefficient.
3. The method for single-frame image 3D human pose estimation and reconstruction based on projection loss constraints according to claim 2, characterized in that, It is calculated using the following formula (2): In formula (2), Represents the coordinates of the key points of the 2D human body. The first in The coordinates of the key points; This indicates that the 3D human body reconstructed by the SMPL model is related to... The projected coordinates of key points that have a one-to-one correspondence; Representing the key points of the 2D human body The number of key points in the text.
4. The method for 3D human pose estimation and reconstruction based on projection loss constraints for single-frame images according to claim 2, characterized in that, It is calculated using the following formula (3): In formula (3), Represents the 2D human body outline. ; This represents the projection of the human body outline after 3D human reconstruction.
5. The method for single-frame image 3D human pose estimation and reconstruction based on projection loss constraints according to claim 2, characterized in that, It is calculated using the following formula (4): In formula (4), Indicates the specified angle The projector camera is set up; " "This indicates three projection cameras at different angles; This involves calculating the vertex projection loss of each triangular mesh used to characterize the reconstructed 3D human body. hour, The weight it occupies; This indicates that the reconstructed 3D human body is viewed at the specified angle. The projection calculation obtained by setting the projection camera The projected coordinates of each vertex; This indicates that the 3D reconstructed human body ground truth is at the specified angle. The first step in setting up the projection camera for projection calculation The projected coordinates of each vertex.
6. The method for 3D human pose estimation and reconstruction based on projection loss constraints for single-frame images according to claim 5, characterized in that, 3 different at the specified angle The angle between each pair of projection cameras is set to 120°.
7. The method for 3D human pose estimation and reconstruction based on projection loss constraints for single-frame images according to claim 2, characterized in that, It is calculated using the following formula (5): In formula (5), The first parameter represents the pose parameter. One value; Represents 72-dimensional pose parameters A collection of locations of the mid-elbow joint; Represents 72-dimensional pose parameters A set of locations for the knee joint.
8. The method for 3D human pose estimation and reconstruction based on projection loss constraints for single-frame images according to claim 1, characterized in that, When reconstructing the 3D human body, the world coordinate system is transformed into the camera coordinate system using the camera parameters T. The rotation matrix when transforming the world coordinate system into the camera coordinate system is set as the identity matrix, and is expressed by the following formula (6): These represent the coordinates of the vertex v of the 3D human body reconstructed in the world coordinate system on the x-axis, y-axis, and z-axis, respectively. These represent the values of the camera parameter T, and respectively represent the translation coefficients of the vertex v in the x-axis, y-axis, and z-axis directions; These represent the coordinates of vertex v in the camera coordinate system on the x-axis, y-axis, and z-axis, respectively.
Citation Information
Patent Citations
Picture-based SMPL parameter prediction and human body model generation method
CN111968217A
Method and apparatus with human body estimation
US20220189123A1