Single-frame image 3D human pose estimation and reconstruction method based on camera parameter constraint
By calibrating camera parameters and constraining the SMPL model using 2D human key points and contour information, the problems of distortion and positional offset in 3D human reconstruction of single-frame images from a monocular camera were solved, achieving more accurate and realistic 3D human reconstruction.
Patent Information
- Application Number
- CN202211021869.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-24
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2042-08-24
AI Technical Summary
In existing technologies, 3D human reconstruction using single-frame images acquired by a monocular camera suffers from problems such as reconstruction distortion, positional shift, and insufficient spatial realism. In particular, it is difficult to accurately recover the 3D human pose when using the SMPL model.
A camera parameter constraint-based approach is adopted. By calibrating the parameters of a monocular camera, a 2D pose estimation model is used to identify key points and contours of the 2D human body. Combined with the SMPL parameter estimation model, the body size, pose, and camera parameters are inferred. The SMPL model is driven to reconstruct the 3D human body using the real camera parameters as constraints. The projection loss and camera parameter constraints are enhanced to improve the reconstruction accuracy.
It reduces distortion in 3D human reconstruction, improves reconstruction stability and spatial realism, ensures the accurate position of the reconstructed 3D human in the scene, and enhances the physical credibility of the reconstruction results.
Smart Images

Figure CN115393436B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of three-dimensional human body reconstruction technology, specifically to a method for 3D human body pose estimation and reconstruction based on single-frame images with camera parameter constraints. Background Technology
[0002] Human 3D pose estimation and reconstruction technology can obtain 3D parameter information from single or multiple frames of 2D images and further reconstruct a 3D human body, with wide applications in sports, film and television, security, and many other fields. Since recovering 3D information from a 2D image is an underdetermined problem, existing technologies typically utilize various additional information for the solution, such as 3D human body reconstruction using consecutive frame images or multi-view images. Currently, 3D human body reconstruction from a single frame image captured by a monocular camera remains a challenging task, usually requiring iterative optimization of the underdetermined problem. Constraints used include human contour constraints and 2D keypoint position constraints, but these optimizations are prone to local optima, leading to distortion and interlacing problems in the reconstructed 3D human body.
[0003] The Skinned Multi-Person Linear Model (SMPL) is a vertex-based parametric 3D human body model that generates a 3D human body mesh driven by shape and pose parameters. Using this model, the reconstruction optimization problem can be transformed into solving the SMPL model parameters, thus obtaining relatively more accurate human body reconstruction results more quickly. However, existing methods for generating 3D human bodies using the SMPL model also suffer from problems such as human body distortion and positional offset, resulting in a lack of spatial realism in the reconstructed 3D human body within the scene. Summary of the Invention
[0004] This invention aims to further reduce the distortion and positional shift of the 3D human body reconstructed by the SMPL model and improve the spatial realism of the reconstructed 3D human body in the scene. It provides a method for 3D human body pose estimation and reconstruction based on single-frame images with camera parameter constraints.
[0005] To achieve this objective, the present invention adopts the following technical solution:
[0006] A method for 3D human pose estimation and reconstruction based on single-frame images with camera parameter constraints is provided, the steps of which include:
[0007] S1, calibrate the parameters of the monocular camera;
[0008] S2, using a 2D pose estimation model, perform 2D human keypoint and 2D human contour recognition on a single frame image acquired by the monocular camera to obtain 2D human keypoints. and 2D human body outline ;
[0009] S3, using the aforementioned 2D human key points and the 2D human body outline diagram As input to the SMPL parameter estimation model, the shape parameters are inferred. Pose parameters and camera parameters ;
[0010] S4, using the actual camera parameters calibrated in step S1 and the camera parameters inferred in step S3. To address the spatial constraints when reconstructing a 3D human body, through... , Drive the SMPL model to reconstruct a 3D human body.
[0011] Preferably, the SMPL parameter estimation model includes a shape estimation module, a pose estimation module, and a camera estimation module. The shape estimation module includes a first shape estimation module and a second shape estimation module.
[0012] The pose estimation module uses the 2D human key points As input, predict the output of the pose parameters. ;
[0013] The camera estimation module uses the 2D human key points As input, predict the output camera parameters. ;
[0014] The first shape estimation module uses the 2D human body contour map. As input, predict the intermediate feature vector as output. ;
[0015] The second shape estimation module uses the intermediate feature vector The posture Pose parameters The camera parameters The shape parameters are used as input to predict the output. .
[0016] Preferably, the objective function of the SMPL parameter estimation model is expressed by the following formula (1):
[0017]
[0018] In formula (1), This indicates the keypoint projection loss after 3D human body reconstruction.
[0019] This indicates the loss of human body contour projection after 3D human body reconstruction.
[0020] This indicates that for the reconstructed 3D human body, the vertex projection loss of each triangular mesh used to characterize the 3D human body is calculated.
[0021] Indicates the prior penalty for joint bending;
[0022] Indicates a priori penalty for full-body posture;
[0023] They represent The loss weighting coefficient.
[0024] As a preferred option It is calculated using the following formula (2):
[0025]
[0026] In formula (2), Representing the key points of the 2D human body The first in The coordinates of the key points;
[0027] This indicates that the 3D human body reconstructed by the SMPL model is related to... The projected coordinates of key points that have a one-to-one correspondence;
[0028] Representing the key points of the 2D human body The number of key points in the text.
[0029] As a preferred option It is calculated using the following formula (3):
[0030]
[0031] In formula (3), Represents the 2D human body outline. ;
[0032] This represents the projection of the human body outline after 3D human reconstruction.
[0033] As a preferred option It is calculated using the following formula (4):
[0034]
[0035] In formula (4), Indicates the specified angle The projector camera is set up;
[0036] “ "Indicates 3 different angles" The projector camera is set up;
[0037] This involves calculating the vertex projection loss of each triangular mesh used to characterize the reconstructed 3D human body. hour, The weight it occupies;
[0038] This indicates that the reconstructed 3D human body is viewed at the specified angle. The projection calculation obtained by setting the projection camera The projected coordinates of each vertex;
[0039] This indicates that the 3D reconstructed human body ground truth is at the specified angle. The projection calculation obtained by setting the projection camera The projected coordinates of each vertex.
[0040] Preferably, the three different specified angles The angle between each pair is 120°.
[0041] It is calculated using the following formula (5):
[0042]
[0043] In formula (5), The pose parameters are represented. The first in One value;
[0044] Represents 72-dimensional pose parameters A collection of locations of the mid-elbow joint;
[0045] Represents 72-dimensional pose parameters The set of locations of the knee joint;
[0046] It is calculated using the following formula (6):
[0047]
[0048] In formula (6), Indicates the component subscripts of the Gaussian mixture model;
[0049] The Gaussian mixture model represents the first... The weight of each component;
[0050] The first Gaussian mixture model represents the... One component;
[0051] Indicates the first Attitude parameters in each component The mean vector;
[0052] Indicates the first The covariance matrix of each component;
[0053] It is a constant.
[0054] Preferably, in step S4, the spatial constraints when reconstructing the 3D human body are expressed by the following formula (7):
[0055]
[0056] In formula (7), This represents the scaling factor used to scale the SMPL model when converting its coordinate system to the world coordinate system.
[0057] This represents the rotation matrix used to rotate the SMPL model when converting its coordinate system to the world coordinate system.
[0058] These represent the vertices in the coordinate system of the SMPL model. exist axis, axis, The coordinates of the axis;
[0059] Each represents a vertex exist axis, axis, Translation coefficient in the axial direction;
[0060] Each represents a vertex In the transformed world coordinate system axis, axis, The coordinates of the axis;
[0061] It is calculated using the following formula (8):
[0062]
[0063] In formula (8), This indicates the tilt or elevation angle of the monocular camera that has been calibrated in step S1;
[0064] In step S3, the camera parameters obtained through reasoning are... Including the scaling factor and the translation coefficient .
[0065] The present invention has the following beneficial effects:
[0066] 1. By using the SMPL model to reconstruct the 3D human body, key point projection loss, human body contour projection loss and multi-angle vertex projection loss are considered, which reduces the distortion of the reconstructed 3D human body.
[0067] 2. Using 2D human body key points and 2D human body outline As input to the SMPL parameter estimation model, the SMPL parameter estimation model estimates the shape parameters. Pose parameters and camera parameters In this way, the influence of features such as the image background in the original 2D human RGB image on the estimation results is avoided, which makes the parameter estimation accuracy of the SMPL parameter estimation model higher, thereby improving the stability of the SMPL model in reconstructing 3D human body.
[0068] 3. By using the projection constraints of camera parameters on the human body, the coordinate system of the SMPL model is associated with the world coordinate system of the physical world, which reduces the positional drift of the reconstructed 3D human body and improves the spatial realism of the reconstructed 3D human body in the scene. Attached Figure Description
[0069] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments of the present invention will be briefly described below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.
[0070] Figure 1 This is a diagram illustrating the implementation steps of a single-frame image 3D human pose estimation and reconstruction method based on camera parameter constraints provided in an embodiment of the present invention.
[0071] Figure 2 This is a network structure diagram for three-dimensional human pose estimation according to an embodiment of the present invention;
[0072] Figure 3 This is a schematic diagram illustrating the angular enhancement of sample data;
[0073] Figure 4 This is a network model structure diagram for 2D pose estimation and semantic segmentation tasks;
[0074] Figure 5 This is a schematic diagram showing the connection between the Shape estimation module, Pose estimation module, and Camera estimation module in the SMPL parameter estimation model.
[0075] Figure 6 This is a diagram of the overall model structure of the SMPL parameter estimation model;
[0076] Figure 7 It involves projecting cameras at different specified angles onto the vertices of the triangular mesh representing the same 3D human body. A diagram showing the layout of the three simulated cameras used for projection;
[0077] Figure 8 The SMPL parameter estimation model is estimated based on the single-frame image captured by the camera. , , Experimental results of reconstructed 3D human body;
[0078] Figure 9 The camera parameters are based on calibrated real camera parameters and inferred camera parameters. The reconstructed image after fine-tuning the position of the 3D human body. Detailed Implementation
[0079] The technical solution of the present invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0080] The accompanying drawings are for illustrative purposes only and are schematic diagrams, not actual images. They should not be construed as limiting the scope of this patent. To better illustrate the embodiments of the present invention, some parts in the drawings may be omitted, enlarged, or reduced, and do not represent the actual dimensions of the product. It is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.
[0081] In the accompanying drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components. In the description of the present invention, it should be understood that if terms such as "upper," "lower," "left," "right," "inner," and "outer" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, they are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, the terms used to describe positional relationships in the drawings are only for illustrative purposes and should not be construed as limiting the present patent. For those skilled in the art, the specific meaning of the above terms can be understood according to the specific circumstances.
[0082] In the description of this invention, unless otherwise explicitly specified and limited, the term "connection" or similar designation indicating a connection between components should be interpreted broadly. For example, it can refer to a fixed connection, a detachable connection, or an integral part; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium; it can refer to the internal communication between two components or the interaction between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0083] The present invention provides a method for 3D human pose estimation and reconstruction based on camera parameter constraints in a single frame image, such as... Figure 2 As shown, it mainly consists of three parts: 1. Generating 2D human keypoints and 2D human contours; 2. SMPL model parameter estimation; 3. Real camera parameter space transformation. In application, the monocular camera first needs to be calibrated to obtain its intrinsic and extrinsic parameters for calculating the projection of the 3D human model onto the real scene. After camera parameter calibration, a single frame of human image captured by the monocular camera is first input into the 2D pose estimation model to infer the 2D human keypoints and 2D human contours. Then, the SMPL parameter estimation model is used to transform the 2D human keypoints... (Coordinates) and 2D human body outline Using this as input, inference obtains the 10-dimensional shape parameters required to reconstruct the 3D human body. (10 parameters together control the human body's size, height, and head-to-body ratio; for example, the first dimension represents the overall size and weight of the human body.) 72 dimensions (each of the 72 dimensions represents the rotation angle of the entire human body or a joint; for example, the first three dimensions represent the rotation angle of the human body around the joint.) axis, axis, The angle of rotation of the axis, with dimensions 4 to 6 and 7 to 9 representing the rotation of the left and right hip joints, respectively. axis, axis, (Axis rotation angle) Pose parameters and 4D ( , , This represents the scaling factor used to scale the SMPL model when converting its coordinate system to the world coordinate system. These represent the vertices of the triangular mesh that characterizes the 3D human body. exist axis, axis, Camera parameters (translation coefficients in the axial direction) Finally, based on the estimated camera parameters... The 3D spatial transformation of the human body model is constrained by the calibrated real camera parameters.
[0084] In this embodiment, the training of the model is divided into two independent sub-processes: 2D pose estimation model training and SMPL parameter estimation model training. HRNet (High-Resolution Net) is a neural network proposed in 2019 for pose estimation of a single human body. It consists of several modules, including convolutional Conv, BatchNorm, and ReLU, forming layers 1, stage 2, stage 3, and stage 4. UDP (Unbiased Data Processing) adds unbiased data processing to HRNet to optimize the model, and it has superior performance in the field of 2D keypoint detection. However, because it lacks semantic segmentation capabilities, it cannot be directly applied to the semantic segmentation of a single frame of human image in this application to obtain a 2D human contour map. To address this issue, this application adds a semantic segmentation feature selection branch after stage 4 of the UDP model structure, resulting in the following... Figure 4 The UDP-S model shown fixes all network parameters except for the semantic segmentation branch, and updates the semantic segmentation branch with a small learning rate. Figure 4 The training is performed on the solid line box portion of the image.
[0085] The second sub-training process involves training the SMPL parameter estimation network PENet (Parameters Estimation Network). PENet includes a pose estimation module, a shape estimation module, and a camera estimation module. The pose estimation module consists of one convolutional layer and three fully connected layers, using 2D human keypoint coordinates... As input, predict the output pose parameters. For a diagram of the internal network structure of the pose estimation module, please refer to [link / reference]. Figure 6 ;
[0086] The shape estimation module consists of... Figure 6The first shape estimation module in the system (i.e. Figure 6 The shape estimation module 1) and the second shape estimation module (i.e. Figure 6 The shape estimation module (2) consists of three convolutional layers, one max pooling layer, and two fully connected layers, based on the 2D human contour map. As input, predict the output shape parameters (Shape). For a diagram of the internal network structure of the shape estimation module, please refer to [link / reference]. Figure 6 ;
[0087] The camera estimation module consists of three fully connected layers, using 2D human keypoint coordinates. Given the input, predict the output camera parameters. For information on the internal network structure of the camera estimation module, please refer to [link / reference]. Figure 6 .
[0088] The following details the methods used by the shape estimation module (including the first and second shape estimation modules), pose estimation module, and camera estimation module to predict the corresponding parameters:
[0089] The pose estimation module estimates parameters. The process is as follows:
[0090] Step A1: 2D Human Body Key Points The dimension is ("2" represents a two-dimensional coordinate;) (The number of key points) is used as the input to the pose estimation module. After passing through the fully connected layer FC1 of the pose estimation module, the output is 512-dimensional. The fully connected layer FC1 has 512 nodes.
[0091] Step A2: The output of the fully connected layer FC1 is used as the input to the convolutional layer Conv1 of the pose estimation module. First, the 512-dimensional layer is resized (transformed) to 1. 1 A 512-dimensional feature map is input into a convolutional layer Conv1, and the output is 1. 1 The feature map has 512 dimensions, and the kernel size of the Conv1 convolutional layer is 1. 1. Number of channels: 1; Step size: 1;
[0092] Step A3: The output of the convolutional layer Conv1 is first resized to 512 dimensions, and then input into the fully connected layer FC2. The output of FC2 is 256 dimensions, and the fully connected layer FC2 has 256 nodes.
[0093] Step A4: The output of the fully connected layer FC2 is used as the input to the fully connected layer FC3 of the pose estimation module. FC3 outputs 72 dimensions, and the fully connected layer FC3 has 72 nodes. The final 72-dimensional output is the predicted pose parameters. ;
[0094] Camera estimation module estimates parameters The process is as follows:
[0095] Step B1: 2D Human Body Keypoint Coordinates Dimensions ("2" represents a two-dimensional coordinate;) (The number of key points) is used as the input to the Camera estimation module. After passing through the fully connected layer FC4 of the Camera estimation module, the output is 512-dimensional. The fully connected layer FC4 has 512 nodes.
[0096] Step B2: The output of the fully connected layer FC4 is used as the input of the fully connected layer FC5 of the Camera estimation module. FC5 outputs 128 dimensions and has 128 nodes.
[0097] Step B3: The output of the fully connected layer FC4 is used as the input of the fully connected layer FC6 of the camera estimation module. FC6 outputs 4 dimensions. The fully connected layer FC6 has 4 nodes, and the final output 4-dimensional features are the predicted camera parameters. ;
[0098] The process of estimating the intermediate feature vector f in the Shape estimation module 1 includes:
[0099] Step C1: 2D Human Body Contour Drawing Size 1 256 192, as the input to the convolutional layer Conv2 of the shape estimation module 1, the kernel size of convolutional layer Conv2 is 7. 32 channels, 2-step size, 3-padding, Conv2 output 32 128 96;
[0100] Step C2: The output of convolutional layer Conv2 is the input of convolutional layer Conv3 of shape estimation module 1, and the kernel size of convolutional layer Conv3 is 3. 3. Number of channels: 128, Step size: 2, Padding: 1, Conv3 output: 128 64 48;
[0101] Step C3: The output of convolutional layer Conv3 is the input of convolutional layer Conv4 of shape estimation module 1, and the kernel size of convolutional layer Conv4 is 3. 3. Number of channels: 512, step size: 2, padding: 1, Conv3 output: 512. 32 twenty four;
[0102] Step C4: The output of the convolutional layer Conv4 is the input to the max pooling layer MaxPool of the shape estimation module 1. The region size of the max pooling layer MaxPool is 32. 24, output 512 1. After resizing, a 512-dimensional intermediate feature vector f is generated;
[0103] Shape estimation module 2 estimates parameters :
[0104] Step C5: 512-dimensional intermediate feature vector f, FC3 outputs 72-dimensional pose parameters. FC6 output camera parameters according to Sequential horizontal stitching generates 588-dimensional features, which are input into the fully connected layer FC7 of the Shape estimation module 2, and output 256-dimensional features. FC7 has 256 nodes.
[0105] Step C6: The output of FC7 is the input to the fully connected layer FC8 of the shape estimation module 2. FC8 outputs 10-dimensional features and has 10 nodes. The final 10-dimensional output is the predicted shape parameters. .
[0106] Obtain parameters Afterwards, through , Drive the SMPL model to reconstruct a 3D human body and generate A triangular mesh with 10 vertices. Then, using camera parameters Using the pre-calibrated real camera parameters as projection constraints, for Projection is performed to obtain the projected coordinates of the key points. (in real number space) (two-dimensional coordinates), vertex projection coordinates (in real space) (two-dimensional coordinates) and human body contour projection .
[0107] The core technology of this application lies in increasing parameters through inter-module connections. , The correlation of T, the improved loss function for reconstructing 3D human bodies using the calibrated camera parameters as constraints for human body projection, make the SMPL model training process more stable, converge faster, and the reconstructed 3D human body more spatially realistic.
[0108] To ensure training effectiveness, the following technical measures are adopted in this application:
[0109] 1. Enhance human body key points and human body contour data.
[0110] In this application, Figure 2 The input to the PENet network shown is not the RGB human image itself captured by a monocular camera, but rather 2D human keypoints predicted by the UDP-S model. and 2D human body outline The effective effect of doing so is: using 2D human key points and 2D human body outline As input to the SMPL parameter estimation model, the SMPL parameter estimation model estimates the shape parameters. Pose parameters and camera parameters This avoids the influence of features such as the image background in the original 2D human RGB image on the estimation results, making the parameter estimation accuracy of the SMPL parameter estimation model higher, and thus improving the stability of the SMPL model in reconstructing 3D human bodies.
[0111] This embodiment uses publicly available 3D human body datasets UP3D (Unite the People-3D body), 3DPW (3D POSES IN THE WILD DATASET), and MTP (Mimic-The-Pose) as samples to train PENet. These datasets record different body shapes. While the parameters are available, the overall data volume is limited and pose variations are minimal. To address this issue, this application employs a data augmentation method to enhance data diversity and robustness. Specifically, a simulated camera with the same image size as the calibration camera is first established, and its intrinsic parameters are set to those of the calibration camera. The world coordinates of the simulated camera coincide with the camera coordinates, meaning the rotation matrix is an identity matrix. In the SMPL parameters, the pose parameter... The first three dimensions are global orient, representing the global rotation coefficient of the human body. The orientation and angle of the human body relative to the camera can be rotated by changing the global orient. Different shapes... The value can be changed randomly, thus by applying the values to each sample... and By performing transformations and combinations, the SMPL model is driven to reconstruct and generate a new human body. Using the human body's vertex ground truth as the rendering projection achieves data augmentation. A new human body rotated 180° is obtained using global rotation coefficients. Figure 3 As shown.
[0112] 2. Predictive Correlation between Pose and Shape Parameters
[0113] The PENet model (i.e., the SMPL parameter estimation model) consists of three modules.
[0114] The model structure is as follows Figure 5 As shown, the three modules are connected in series and share parameters. (2D human body contour diagram) The area and shape are affected by different poses and camera distances. Figure 5 The model structure links modules together, first determining the pose and human position, then predicting the body size. Specifically, the shape estimation module outputs an intermediate feature vector. ,and according to After sequential horizontal stitching, the fully connected layer is input to predict the output shape parameters. .
[0115] 3. Enhance loss function constraints
[0116] The loss function calculation for the PENet model includes: 2D human keypoints and 2D human body outline The projection loss is calculated using self-supervised information, and abnormal poses are penalized based on joint and whole-body priors. In this application, the objective function minimized during training is:
[0117]
[0118] In formula (1), This indicates the keypoint projection loss after 3D human body reconstruction.
[0119] This indicates the loss of human body contour projection after 3D human body reconstruction.
[0120] This indicates that for the reconstructed 3D human body, the vertex projection loss of each triangular mesh used to represent the 3D human body is calculated.
[0121] Indicates the prior penalty for joint bending;
[0122] Indicates a priori penalty for full-body posture;
[0123] They represent The loss weighting coefficient.
[0124] It is calculated using the following formula (2):
[0125]
[0126] In formula (2), Representing key points of a 2D human body The first in The coordinates of the key points;
[0127] This indicates that on the 3D human body reconstructed from the SMPL model, and... The projected coordinates of key points that have a one-to-one correspondence;
[0128] Representing key points of a 2D human body The number of key points in the text.
[0129] It is calculated using the following formula (3):
[0130]
[0131] In formula (3), Represents a 2D human body outline. ;
[0132] This represents the projection of the human body outline after 3D human reconstruction.
[0133] Because there is no one-to-one correspondence between two-dimensional pose and three-dimensional pose. The constraint on the body pose remains ambiguous, affecting model convergence. To improve convergence speed and reconstruction accuracy, this application introduces vertex projection coordinate loss constraints into the loss function. This constraint uses the calibration camera as the primary projection camera, adds two auxiliary projection cameras, and sets the angle between the three cameras to 120°. Each projection camera is assigned a weight corresponding to the calculation of the vertex projection loss. This application utilizes projection cameras from different angles. Calculate the vertex projection coordinate loss to reduce the error caused by the lack of depth information.
[0134] It is calculated using the following formula (4):
[0135]
[0136] In formula (4), Indicates the specified angle Projector camera;
[0137] “ "Indicates 3 at different specified angles" A projection camera is set up; preferably, such as Figure 7 As shown, three at different specified angles The angle between each pair of projection cameras is set to 120°.
[0138] This indicates that for the reconstructed 3D human body, the vertex projection loss of each triangular mesh used to represent the 3D human body is calculated. hour, The weight it occupies;
[0139] This indicates the reconstructed 3D human body at a specified angle. The projection calculation obtained by setting the projection camera The projected coordinates of each vertex;
[0140] This indicates that the 3D reconstructed human body ground truth is at the specified angle. The projection calculation obtained by setting the projection camera The projected coordinates of each vertex.
[0141] This represents a simple a priori representation of joint flexion, constraining the unnatural rotational directions of the elbow and knee. It is calculated using the following formula (5):
[0142]
[0143] In formula (5), Represents the pose parameters The first in One value;
[0144] Represents 72-dimensional pose parameters A collection of locations of the mid-elbow joint;
[0145] Represents 72-dimensional pose parameters A set of locations for the knee joint.
[0146] The prior penalty for the whole-body posture is represented by the negative logarithm of the weighted sum of Gaussian mixtures, which penalizes abnormal movements. To simplify the calculation, the maximum error term of the Gaussian mixture model is taken and a constant is introduced. As an approximation of the sum. It is calculated using the following formula (6):
[0147]
[0148] In formula (6), Indicates the component subscripts of the Gaussian mixture model;
[0149] The Gaussian mixture model represents the first... The weight of each component;
[0150] The first Gaussian mixture model represents the... One component;
[0151] Indicates the first Attitude parameters in each component The mean vector;
[0152] Indicates the first The covariance matrix of each component; It is a constant.
[0153] It should be noted here that... The calculation process is quite complex. To ensure the accuracy of the model's predictions, we will... Incorporate it into the loss function. One of the loss methods is considered as a preferred option. In practical applications, to simplify the model training process, It can be disregarded.
[0154] 4. Camera parameter projection constraints
[0155] The reconstructed 3D human body model described above was created using a simulated camera. To make the 3D human body model more realistic in the scene, this application uses a fixed camera in a real scene. The intrinsic and extrinsic parameters of the camera are obtained through parameter calibration, and then the reconstructed model is based on the camera parameters predicted during reconstruction. The relationship between the reconstructed 3D human body and the calibrated real camera parameters is used to apply spatial constraints to the reconstructed 3D human body. The specific method is as follows:
[0156] The Zhang Zhengyou calibration method was used to calibrate the parameters of the monocular camera. Through calibration, the camera's intrinsic parameters and its tilt or elevation angles were obtained. The camera intrinsic parameters are used to calculate keypoint and vertex projections, as well as for rendering 2D images of the reconstructed human body. To simplify calculations, the camera coordinate system is set to coincide with the world coordinate system, in which case the camera extrinsic matrix is the identity matrix. The coordinate system of the SMPL model is not consistent with the world coordinate system; therefore, the SMPL model needs to be rotated, scaled, and translated to convert it into a physically meaningful representation in the world coordinate system. The scaling factor... Translation coefficient Predicted by the PENet network. Rotation matrix. Depending on the camera's tilt or elevation angle It is calculated, and the specific calculation formula is as follows:
[0157]
[0158] Vertices of the triangular mesh of the 3D human body reconstructed from the SMPL model Converted under camera parameter constraints The conversion formula is:
[0159]
[0160] In formula (7), This represents the scaling factor used to scale the SMPL model when converting its coordinate system to the world coordinate system.
[0161] This represents the rotation matrix used to rotate the SMPL model when converting its coordinate system to the world coordinate system.
[0162] These represent the vertices in the SMPL model coordinate system. exist axis, axis, The coordinates of the axis;
[0163] Representing vertices respectively exist axis, axis, Translation coefficient in the axial direction;
[0164] Representing vertices respectively In the transformed world coordinate system axis, axis, The coordinates of the axis;
[0165] The following are experimental verification results of the technical effects achieved by the single-frame image 3D human pose estimation and reconstruction method based on camera parameter constraints provided in this embodiment:
[0166] 1. SMPL parameter estimation model for human body reconstruction results
[0167] Figure 8 This paper presents a series of keyframes captured by a camera, showing a moving human body from a distance to a close-up. The PENet network is used to predict the parameters of the SMPL model, driving the human body to change its shape and posture to obtain a 3D reconstruction. Two-dimensional images are then rendered using camera intrinsic and extrinsic parameters. The results show that the reconstructed body dimensions have good consistency, and the reconstructed posture, such as the joint flexion range and body tilt angle, appears natural.
[0168] 2. Comparison Results of Real-World 3D Human Body Reconstruction
[0169] After calibrating a monocular camera to obtain camera parameters, a 3D human body model under the real camera (the world coordinate system of the physical world) can be reconstructed through rotation, translation, and scaling operations. Since there are errors between the camera parameters of the simulated camera used in the training model and the calibrated camera, this application uses keypoint coordinates to adjust the scaling factor to address this issue. Translation coefficient Fine-tuning was performed, and the reconstruction effect was as follows: Figure 9 As shown, the 3D human body model has a more realistic spatial position in the camera image, conforms to the camera's projection relationship, and the reconstructed model has a better fit with the ground in the image, so that the person is walking on the ground rather than floating.
[0170] In summary, the single-frame image 3D human pose estimation and reconstruction method based on camera parameter constraints provided in this invention embodiment is as follows: Figure 1 As shown, the steps include:
[0171] S1, calibrate the parameters of the monocular camera;
[0172] S2, using a 2D pose estimation model, performs 2D human keypoint and 2D human contour recognition on a single frame image captured by a monocular camera, obtaining 2D human keypoints. and 2D human body outline ;
[0173] S3, based on 2D human key points and the 2D human body outline diagram As input to the SMPL parameter estimation model, the shape parameters are inferred. Pose parameters and camera parameters ;
[0174] S4, using the actual camera parameters calibrated in step S1 and the camera parameters inferred in step S3. To address the spatial constraints when reconstructing a 3D human body, through... , Drive the SMPL model to reconstruct a 3D human body.
[0175] It should be stated that the above-described specific embodiments are merely preferred embodiments of the present invention and the technical principles employed. Those skilled in the art should understand that various modifications, equivalent substitutions, and variations can be made to the present invention. However, such variations, as long as they do not depart from the spirit of the present invention, should be within the scope of protection of the present invention. Furthermore, some terminology used in this specification and claims is not limiting, but merely for ease of description.
Claims
1. A method for 3D human pose estimation and reconstruction based on single-frame images with camera parameter constraints, characterized in that the steps include... include: S1, calibrate the parameters of the monocular camera; S2, using a 2D pose estimation model, perform 2D human keypoint and 2D human contour recognition on a single frame image acquired by the monocular camera to obtain 2D human keypoints. and 2D human body outline ; S3, using the aforementioned 2D human key points and the 2D human body outline diagram As input to the SMPL parameter estimation model, the shape parameters are inferred. Pose parameters and camera parameters ; S4, using the actual camera parameters calibrated in step S1 and the camera parameters inferred in step S3. To address the spatial constraints when reconstructing a 3D human body, through... , Drive SMPL model to reconstruct 3D human body; The SMPL parameter estimation model includes a shape estimation module, a pose estimation module, and a camera estimation module. The shape estimation module includes a first shape estimation module and a second shape estimation module. The pose estimation module uses the 2D human key points As input, predict the output of the pose parameters. ; The camera estimation module uses the 2D human key points As input, predict the output camera parameters. ; The first shape estimation module uses the 2D human body contour map. As input, predict the intermediate feature vector as output. ; The second shape estimation module uses the intermediate feature vector The posture Pose parameters The camera parameters The shape parameters are used as input to predict the output. .
2. The method for 3D human pose estimation and reconstruction based on camera parameter constraints in a single frame image according to claim 1, characterized in that, The objective function of the SMPL parameter estimation model is expressed by the following formula (1): In formula (1), This indicates the keypoint projection loss after 3D human body reconstruction. This indicates the loss of human body contour projection after 3D human body reconstruction. This indicates that for the reconstructed 3D human body, the vertex projection loss of each triangular mesh used to characterize the 3D human body is calculated. Indicates the prior penalty for joint bending; Indicates a priori penalty for full-body posture; They represent The loss weighting coefficient.
3. The method for 3D human pose estimation and reconstruction based on camera parameter constraints in a single frame image according to claim 2, characterized in that, It is calculated using the following formula (2): In formula (2), Representing the key points of the 2D human body The first in The coordinates of the key points; This indicates that the 3D human body reconstructed by the SMPL model is related to... The projected coordinates of key points that have a one-to-one correspondence; Representing the key points of the 2D human body The number of key points in the text.
4. The method for 3D human pose estimation and reconstruction based on camera parameter constraints in a single frame image according to claim 2, characterized in that, It is calculated using the following formula (3): In formula (3), Represents the 2D human body outline. ; This represents the projection of the human body outline after 3D human reconstruction.
5. The method for 3D human pose estimation and reconstruction based on camera parameter constraints in a single frame image according to claim 2, characterized in that, It is calculated using the following formula (4): In formula (4), Indicates the specified angle The projector camera is set up; " "Indicates 3 different angles" The projector camera is set up; This involves calculating the vertex projection loss of each triangular mesh used to characterize the reconstructed 3D human body. hour, The weight it occupies; This indicates that the reconstructed 3D human body is viewed at the specified angle. The projection calculation obtained by setting the projection camera The projected coordinates of each vertex; This indicates that the 3D reconstructed human body ground truth is at the specified angle. The projection calculation obtained by setting the projection camera The projected coordinates of each vertex.
6. The method for 3D human pose estimation and reconstruction based on camera parameter constraints in a single frame image according to claim 5, characterized in that, 3 different specified angles The angle between each pair is 120°.
7. The method for 3D human pose estimation and reconstruction based on camera parameter constraints in a single frame image according to claim 2, characterized in that, It is calculated using the following formula (5): In formula (5), The pose parameters are represented. The first in One value; Represents 72-dimensional pose parameters A collection of locations of the mid-elbow joint; Represents 72-dimensional pose parameters The set of locations of the knee joint; It is calculated using the following formula (6): In formula (6), Indicates the component subscripts of the Gaussian mixture model; The Gaussian mixture model represents the first... The weight of each component; The first Gaussian mixture model represents the... One component; Indicates the first Attitude parameters in each component The mean vector; Indicates the first The covariance matrix of each component; It is a constant.
8. The method for 3D human pose estimation and reconstruction based on camera parameter constraints in a single frame image according to claim 1, characterized in that, In step S4, the spatial constraints when reconstructing the 3D human body are expressed by the following formula (7): In formula (7), This represents the scaling factor used to scale the SMPL model when converting its coordinate system to the world coordinate system. This represents the rotation matrix used to rotate the SMPL model when converting its coordinate system to the world coordinate system. These represent the vertices in the coordinate system of the SMPL model. exist axis, axis, The coordinates of the axis; Each represents a vertex exist axis, axis, Translation coefficient in the axial direction; Each represents a vertex In the transformed world coordinate system axis, axis, The coordinates of the axis; It is calculated using the following formula (8): In formula (8), This indicates the tilt or elevation angle of the monocular camera that has been calibrated in step S1; In step S3, the camera parameters obtained through reasoning are... Including the scaling factor and the translation coefficient .
Citation Information
Patent Citations
Picture-based SMPL parameter prediction and human body model generation method
CN111968217A
Method and apparatus with human body estimation
US20220189123A1