Method, apparatus, and system for three-dimensional pose estimation
By combining time 1D convolution and LSTM network architecture, combined with motion chain space regularization and data enhancement technology, the problem of combining global trajectory and relative pose in human body 3D pose estimation is solved, and stable pose estimation under a monocular camera is realized to adapt to outdoor video and dynamic motion.
Patent Information
- Application Number
- CN202180016890.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-08-25
- Filing Date
- 2021-08-18
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2041-08-18
AI Technical Summary
The prior art has failed to effectively solve the combination of global trajectory and relative pose in human body 3D pose estimation, especially in outdoor videos, where depth ambiguity and trajectory drift problems, and common data sets cannot cover real-world scenarios.
Using a network architecture combining time 1D convolution and long-term short-term memory (LSTM), the network architecture is used to regularize through feature extraction and bone length and unit vector estimation, combined with motion chain space (KCS), and uses automatic enhancement and perturbation data to simulate dynamic motion and occlusion, and directly regress to the root position and relative pose.
The root position and relative posture of the human body's 3D pose are achieved under the monocular camera conditions, adapting to various fields of view and dynamic movements, and improving the accuracy and robustness of pose estimation.
Smart Images

Figure CN115151944B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to content estimation. More specifically, the present invention relates to 3D pose estimation. Background Art
[0002] After great success in human 2D pose estimation, in order to expand its applications such as in movies, surveillance, and human-computer interaction, human 3D pose estimation has attracted wide attention. Many methods have been proposed, including multi-view methods, temporal methods, monocular 3D pose methods for skeletons, and monocular 3D pose methods with 3D meshes. Summary of the Invention
[0003] Recent advances in neural networks have shown significant progress in human pose estimation tasks. Pose estimation can be divided into monocular 2D pose estimation, multi-view Figure 3 D pose estimation, and monocular Figure 3 D pose estimation, where 3D pose has recently received increasing attention for applications in AR / VR, gaming, and human-computer interaction applications. However, current academic benchmarks for human 3D pose estimation only consider the performance of its relative pose. Root localization over time, in other words, the "trajectory" of the entire body in 3D space, is not considered well enough. Applications such as motion capture require not only the precise relative pose of the body but also the root position of the entire body in 3D space. Therefore, an effective monocular full 3D pose recovery model from 2D pose input is described herein, which can be applied to the above applications. Described herein are the network architecture that combines temporal 1D convolution and long short-term memory (LSTM) for root position estimation, how to formulate the output, the design of the loss function, and a comparison with existing technology models to demonstrate the effectiveness of the method for application use. As described herein, 3D pose estimation for 15 and 17 key points is performed, but it can be extended to any key point definition.
[0004] In one aspect, a method includes receiving camera information, where the camera information includes a two-dimensional pose and camera parameters including a focal length; applying feature extraction to the camera information, including determination using a one-dimensional convolutional residual; estimating a bone length based on the feature extraction; estimating a bone unit vector based on the feature extraction conditioned on the bone length; and estimating a relative pose from the bone length and the bone unit vector, and deriving a root position based on the feature extraction conditioned on the bone length and the bone unit vector. The method further includes receiving one or more frames as input. It is assumed that each bone length does not exceed a length of 1 meter. A long short-term memory is used to estimate the root position to stabilize the root position. The method further includes applying auto-augmentation to the global position and rotation to simulate dynamic motion. The method further includes randomly changing the field of view of the camera for each batch of samples to estimate any video using different camera parameters. The method further includes performing perturbation of the two-dimensional pose with Gaussian noise and random key-point dropout on the two-dimensional pose input to simulate noise and occlusion in two-dimensional pose prediction.
[0005] In another aspect, a device includes: a non-transitory memory for storing an application; and a processor coupled to the memory, the processor being configured to process the application, the application being for: receiving camera information, where the camera information includes a two-dimensional pose and camera parameters including a focal length; applying feature extraction to the camera information, including determination using a one-dimensional convolutional residual; estimating a bone length based on the feature extraction; estimating a bone unit vector based on the feature extraction conditioned on the bone length; and estimating a relative pose from the bone length and the bone unit vector, and deriving a root position based on the feature extraction conditioned on the bone length and the bone unit vector. In the device, the application is configured to receive one or more frames as input. It is assumed that each bone length does not exceed a length of 1 meter. A long short-term memory is used to estimate the root position to stabilize the root position. The application is configured to apply auto-augmentation to the global position and rotation to simulate dynamic motion. The application is configured to randomly change the field of view of the camera for each batch of samples to estimate any video using different camera parameters. The application is configured to perform perturbation of the two-dimensional pose with Gaussian noise and random key-point dropout on the two-dimensional pose input to simulate noise and occlusion in two-dimensional pose prediction.
[0006] In another aspect, a system includes: a camera configured to acquire content; and a computing device configured to: receive camera information from the camera, where the camera information includes a two-dimensional pose and camera parameters including a focal length; apply feature extraction to the camera information, including determination using a one-dimensional convolutional residual; estimate a bone length based on the feature extraction; estimate a bone unit vector based on the feature extraction conditioned on the bone length; and estimate a relative pose from the bone length and the bone unit vector, and derive a root position based on the feature extraction conditioned on the bone length and the bone unit vector. The second device is also configured to receive one or more frames as input. Assume that each bone length does not exceed a length of 1 meter. Use a long short-term memory for estimating the root position to stabilize the root position. The second device is also configured to apply auto-augmentation to the global position and rotation to simulate dynamic motion. The second device is also configured to randomly change the field of view of the camera for each batch of samples to estimate arbitrary videos using different camera parameters. The second device is also configured to perform perturbation of the two-dimensional pose with Gaussian noise and random keypoint dropout on the two-dimensional pose input to simulate noise and occlusion conditions in two-dimensional pose prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Figure 1 Shows a visualization of the model output for an outdoor video (e.g., with uncontrolled variables / environment) according to some embodiments.
[0008] Figure 2 Represents that the model described herein overcomes depth ambiguity according to some embodiments.
[0009] Figure 3 Represents the same 3D pose reprojected to UV with different FOVs according to some embodiments.
[0010] Figure 4 Represents a change in keypoint definition according to some embodiments.
[0011] Figure 5 Represents the 3D camera coordinates of the design described herein according to some embodiments.
[0012] Figure 6 Represents how values are encoded in the normalized space according to some embodiments.
[0013] Figure 7 Represents the distribution of the root position in the X-Z and Z-Y planes of the camera coordinates according to some embodiments.
[0014] Figure 8 Represents perturbation of the input and keypoint dropout according to some embodiments.
[0015] Figure 9 Represents a simplified block diagram of the network as described herein according to some embodiments.
[0016] Figure 10 Visualization of the predicted root position of the target according to some embodiments.
[0017] Figure 11 Table showing the results of the Human3.6M plus data augmentation scheme described herein according to some embodiments.
[0018] Figure 12 Table showing the comparison of LSTM and 1D convolutional root position estimation according to some embodiments.
[0019] Figure 13 Visualization of the Z-axis root position tracking of a sample sequence according to some embodiments to compare models using LSTM and 1D convolution.
[0020] Figure 14A ~B represents a backflip video from YouTube and applying AlphaPose as a 2D pose detector according to some embodiments, which is then performed by using the method described herein.
[0021] Figure 15 Block diagram of an exemplary computing device configured to implement a full-skeleton 3D pose recovery method according to some embodiments. Detailed Description
[0022] Recent advances in neural networks have shown significant progress in human pose estimation tasks. Pose estimation can be divided into monocular 2D pose estimation, multi-view Figure 3 D pose estimation, and monocular Figure 3 D pose estimation, where 3D pose has recently received increasing attention for applications in AR / VR, gaming, and human-computer interaction. However, current academic benchmarks for human 3D pose estimation only consider the performance of its relative pose. Root localization over time, in other words, the "trajectory" of the whole body in 3D space, is not well considered. Applications such as motion capture require not only the precise relative pose of the body but also the root position of the whole body in 3D space. Therefore, an effective monocular full 3D pose recovery model from 2D pose input is described herein, which can be applied to the above applications. Described herein are the network architecture that combines temporal 1D convolution and long short-term memory (LSTM) for root position estimation, how to formulate the output, the design of the loss function, and the comparison with prior art models to show the effectiveness of the method for application use. As described herein, 3D pose estimation of 15 and 17 key points is performed, but it can be extended to any key point definition.
[0023] Monocular human 3D pose estimation has become a hot topic in the research community because it can be applied to outdoor videos (e.g., uncontrolled environments), which are available as consumer-generated videos over the Internet. Additionally, enabling pose estimation in a monocular setup eliminates the need for the installation of multiple cameras and their alignment, enabling triangulation to be solved. Although recent work on monocular human 3D pose estimation has shown significant improvements over time, combining global trajectories and relative poses is an extremely difficult problem due to the nature of the ambiguity caused by multiple 3D poses mapping to the same 2D pose. Moreover, since the ambiguity of depth and the drift of its trajectory cannot be fully observed if the results are only overlaid on the input image plane, it is quite difficult to qualitatively evaluate these methods from video results and sometimes misleading in terms of performance. Secondly, the methods mentioned in the background section only evaluate their relative 3D poses, where the relative pose is defined as the root bone in a fixed (e.g., zero) position, and the recovery of trajectories in motion is not fully studied. Finally, the main dataset Human3.6M used in the above human 3D pose estimation evaluation lacks a real-world setting and cannot cover the situations that may occur when applied to outdoor videos. This dataset was captured in a laboratory setup with 8 cameras having almost identical camera parameters in a 3×4-meter area. Therefore, additional 2D pose data is usually used in a semi-supervised manner with adversarial loss.
[0024] To address the above problems that are very important for applying monocular human 3D pose estimation to motion capture purposes, the following aspects are described: A unified human 3D relative pose and trajectory recovery network from 2D pose inputs, combined with 1D convolution for relative pose and LSTM for trajectories. Compared with previous state-of-the-art methods, this model is efficient in terms of parameter size, and more stable trajectory recovery can be observed by using LSTM on top of convolution.
[0025] Like the method in VP3D, the model does take multiple frames (if available), but is not limited to using multiple frames by design. VP3D stands for VideoPose3D and is from "3D Human Pose Estimation in Video with Temporal Convolutions and Semi-Supervised Training" at https: / / github.com / facebookresearch / VideoPose3D. The VP3D method reaches its best performance when the input is 243 frames, but the model described in this paper can work even when the input is 1 frame. This is particularly important when one wants to apply any number of frame inputs to the processing. More value lies in usability rather than making the relative pose accuracy differ by a few millimeters.
[0026] To simultaneously regress the root position and relative pose in a unified network, the kinematic chain space (KCS) is used for regularization purposes. As an alternative to using regularization, the bone unit vectors and bone lengths are directly estimated and respective losses are applied to enhance the consistency of bone lengths between input frames. Assuming each bone length is in the range {0,1}, within which it is assumed that the human bone length does not exceed 1 meter. Additionally, tanh is applied to the root position (such as in an encoding / decoding scheme) so that network parameters can be generated within the same dynamic range.
[0027] Since the Human3.6M dataset has a relatively small coverage in 3D space and actions, auto-augmentation is applied to the global position and rotation to simulate dynamic motions such as backflips or handsprings. The camera field of view for each batch of samples is randomly changed so that, given the camera parameters for prediction as a condition, different camera parameters can be used to estimate any video. Perturbation of 2D poses with Gaussian noise and random keypoint dropout is performed on the 2D pose input to simulate the noise and occlusion situations in 2D pose prediction. This allows using only motion capture data where adversarial modules or losses do not occur, thus making the preparation and training time shorter. The Human3.6M data is merely an exemplary dataset used with the methods and systems described in this paper and is not meant to be limiting in any way. Any 3D human motion capture dataset can be used with the methods and systems described in this paper.
[0028] Figure 1 Visualization of the model output for an outdoor video (e.g., with uncontrolled variables / environment) according to some embodiments is shown. In Figure 1In it, column (a) represents video frames with 2D pose estimation, column (b) represents 3D pose in the X-Y plane, and column (c) represents 3D pose in the X-Z plane. The red line on the 3D graph (the line that usually passes through the back and head of a person) indicates the global trajectory. The model can output a trajectory with a stable z position for dynamic motion. For the detailed definition of camera coordinates, see Figure 5 .
[0029] As Figure 2 shown, the model described in this paper overcomes depth ambiguity. Figure 2 The top row represents the camera plane projection of the 3D human pose prediction. The bottom row represents the reconstructed side view. Even if a person only moves parallel to the camera, the whole body, especially in the depth direction, is not well estimated.
[0030] Monocular 3D human pose estimation methods can be roughly divided into two categories: grid-based methods and 2D lifting methods.
[0031] Grid-based method
[0032] Grid-based methods use a prior model such as a human body grid to adapt to the image plane by not only recovering the pose but also the skin. Specifically, if they overlap in the image plane, grid-based methods will show better results, but if you view from a different angle (such as Figure 2 the side view), unstable trajectory tracking is visible. This stems from the nature of the high ambiguity problem that monocular methods suffer from. Even though the problem space is made smaller by using a human body prior model, this problem still remains to be well solved.
[0033] 2D lifting method
[0034] Another type is monocular 3D human bone pose, where the input to the model is the 2D pose predicted by a well-established human 2D pose detector. To be stable over the time dimension, some implementations use the LSTM sequence-to-sequence method. However, their methods include encoding all frames into a fixed length. VP3D utilizes temporal information by performing 1D convolution over the time dimension. They also split the network into two, where the relative pose and trajectory estimation networks are separated and jointly trained. However, the networks for relative pose and trajectory each use 16M parameters, and the network for full pose estimation is 32M parameters. It also uses 243 frames of input to obtain the best performance, and due to the limited camera configuration of Human3.6M, it cannot handle well videos with camera parameters different from the training data.
[0035] Motion chain space
[0036] The Kinematic Chain Space (KCS) can be used to decompose poses into bone vectors and their lengths. The idea of using KCS instead of estimating relative poses in Cartesian coordinates is followed. The models described in this paper differ in how to utilize KCS for optimization. KCS has been used to map relative poses in KCS and make the adversarial loss serve as a regularization term to train the model in a semi-supervised manner. Different from the above, the method described in this paper directly regresses bone vectors and bone lengths located in the normalized space.
[0037] What is described in this paper are the definitions of the input and output, the dataset, and how to perform augmentation, network design, and loss formula.
[0038] Input
[0039] Following a scheme similar to the 2D pose lifting method described in this paper, where 2D poses can be estimated from an arbitrary 2D pose detector. For example, AlphaPose can be used. As Figure 4 shown, the 2D pose detector outputs 17 - 25 various key points. For example, Human3.6M uses 17 key points (17 out of 32 defined movable). To make the model described in this paper work on an arbitrary 2D pose detector, 15 most intersecting key points are defined and can be evaluated using Human3.6M data (or other data). As input, UV-normalized 2D coordinates are used, where u ∈ {0, 1}. Additionally, due to occlusion, the 2D pose detector usually cannot detect some key points, which is a common situation. For this, these values are set to zero. The camera focal length is also used as input. Monocular 3D human pose estimation methods use Human3.6M and Human-Eva, but neither of these datasets has various camera settings, and attempts are made to apply the model to outdoor videos and images by applying semi-supervised training using 2D annotations. The camera parameters can be estimated to calculate the reprojection error, but the camera parameters are still implicitly modeled by the pose generator network. Instead, the network described in this paper is modeled conditioned on 2D pose input and camera focal length. The focal length is a very important queue to support arbitrary cameras. As Figure 3 shown, even with the same relative pose and root position in 3D space, different camera fields of view (FOV) make the 2D pose appearance very different. Otherwise, it is difficult to estimate the correct pose in 3D. As described in this paper, it is assumed that a perspective projection camera with the principal point at the center of the image and lens distortion is not considered.
[0040] Figure 3Represents the same 3D pose reprojected onto UV with different FOVs. The FOV of figure (a) is 60°; the FOV of figure (b) is 90° and the FOV of figure (c) is 120°. In each clip taken outdoors, the camera parameters may be different.
[0041] Figure 4 Represents the change in key point definition. Image (a) is MSCOCO with 17 points. Image (b) is OpenPose with 18 points. Image (c) is OpenPose with 25 points. Image (d) is Human3.6 with 17 points (17 out of 32 are movable). Image (e) is the method with 15 point definitions described in this paper. In each definition, these lines are standard bone pairs.
[0042] Output and motion chain space
[0043] The network output is defined as a combination of the root position of the body and the relative pose. The root position is usually defined at the key point of the pelvis. The relative pose is defined as the 3D position of other bones relative to the root position. Figure 4 Image (e) describes 15 key point definitions, where 0 is the pelvis to be used as the root position, and the others are to be estimated as relative positions relative to the root. Figure 5 Describes the definition of the 3D space as described in this paper. Figure 5 Represents the 3D camera coordinates of the design described in this paper according to some embodiments. Figure 3 Represents the 2D projection of the pose under different FOVs. The relative pose and the root position are estimated in the camera coordinates. In addition, KCS is used to decompose the relative pose into bone vectors and their lengths. The i-th joint of the kinematic chain is represented by a vector defined by including the x, y, and z coordinates of the joint position. By connecting the j joint vectors, a matrix representing the relative pose P r of:
[0044] P r = (p1, p2,..., p j ) (1)
[0045] And, the whole body pose P is expressed as:
[0046] P = (p0, p0,..., p r ) (2)
[0047] where p0 is the root position, and the relative pose is derived by subtracting the root pose. The k-th bone b k is defined as the vector between the r-th joint and the t-th joint,
[0048] b k = pr - p t = P rj d k , (3)
[0049] wherein,
[0050] d = (0, ..., 0, 1, 0, ..., 0, 0, -1, ..., 0) T , (4)
[0051] D=(d1,d2,...,d j ),
[0052] is 1 at position r and -1 at position t. d is the mapping vector of the r-th joint and the t-th joint, and through the connection of the entire joints, the entire mapping matrix D is expressed as Similar to Equation 1, the matrix can be defined as the matrix containing all b bones:
[0053] B = (b1, b2, ..., b b ); (5)
[0054] wherein, the matrix B is calculated from P r by the following formula:
[0055] B = P r D. (6)
[0056] Similar to D, the matrix r that maps b back to P
[0057] P r = BE. (7)
[0058] Then, the network can learn the mapping function:
[0059]
[0060] wherein, the input is the 2D pose u and the camera parameter c, and the output is the root position to be estimated bone length and its unit vector θ includes the network parameters. The reason for not directly estimating the bone vector b is to keep the output in the normalized space. In some embodiments, it is assumed that each bone length follows ||b k ||∈{0,1}, which never exceeds 1m. Any symbol with a hat is a prediction, and the symbol without a hat is the ground truth (e.g., label) to define how much the prediction loses compared to the truth.
[0061] root position Encoding and decoding are performed using the tanh form in the normalized space. Then it is decoded into the actual value. The encoding formula is:
[0062]
[0063]
[0064] And the decoding return will be formed as:
[0065]
[0066] where β and ε are constant values. Use β = 0.1e and ε = 1e -8 . Figure 6 Shows how to encode values in the normalized space. Figure 6 Represents encoding and decoding the root position. It provides greater granularity at distances near the camera and saturates at 20m. The z-axis value will be non-negative.
[0067] Since many pose regression models do not well consider how the output space and parameter space should be modeled, this normalization is very important. VP3D proposes to simultaneously estimate the root position and relative pose using two discrete networks with weighted losses on the root position, where the loss for the distant root position has a smaller weight. The method described in this paper includes forming granularity in the encoding space and including the parameter space. This is very important for propagating gradients and updating parameters, not only for the root position but also for the bone vectors in the end-to-end training manner.
[0068] Dataset and augmentation
[0069] The model described in this paper is purely trained from motion capture data, rather than using 2D pose annotations in a semi-supervised manner (many methods use this way to generalize well to outdoor videos and images). Human3.6M is used for the initial experiments for pure academic purposes, and for commercial purposes, motion capture data provided by Sony Interactive Entertainment (SIE) is used. However, the motion capture data may be too small to cover real-world scenarios. To solve this problem, several augmentations and perturbations are automatically adopted in the training data.
[0070] Algorithm 1: Pose Augmentation
[0071] Input:
[0072] FOV ← Random FOV set ∈ {°40; °100}
[0073] L ← Position range (limit)
[0074] x ∈ {-10, 10}, y ∈ {10, 10}, z ∈ {0, 10}
[0075] S ← Camera image size range
[0076] τ ← Variance threshold for rotational motion
[0077] Output:
[0078] Enhanced pose in camera coordinates: p'
[0079] 2D projected pose for model input: u'
[0080] Camera parameters: c = (f x ; f y )
[0081] Data: Pose sequence data P
[0082]
[0083] Set random FOV, diagonal focal length:
[0084] v ← Randomly select from FOV
[0085] fdiag ← 0.5 / tan(v * 0.5)
[0086] Set vertical and horizontal focal lengths according to random aspect ratio: f x , f y F f (f diag , U(0.5, 2.0), S)
[0087] Globally rotate the pose randomly along the Y-axis:
[0088]
[0089] Get the maximum and minimum positions:
[0090] p' min , p' max ← min(p'), max(p')
[0091] Determine the random camera position q on the Z-axis within the frustum using v: q z ← F z (p' maxx , p' minx ; L z , v)
[0092] Determine the random camera position q on the X-axis: q x ← F x (p' maxx , p' minx,q z )
[0093] Use v to determine a random camera position q on the Y-axis within the frustum: q y ←F y (q z ,L z ,v)
[0094] Offset the position p’ with q: p’ ← p’ + q
[0095] Calculate the trajectory variance: σ ← p’
[0096] If σ z < τ,
[0097] Then for each p’ t ∈ p’, linearly rotate along the Z-axis: p’ ← RotateZ T (p’)
[0098] Otherwise, if xσ x < τ
[0099] Then for each p’ t ∈ p’, linearly rotate along the X-axis:
[0100] p’ ← RotateX T (p’)
[0101] Project onto 2D: u’ ← ProjectLinear(p’, c)
[0102] Algorithm 1 is the simplified pseudocode for augmentation. Given the entire dataset P, each batch of samples contains time frames of length T, i.e., p t ∈ p, t = (0, 1, …, T). The FOV is randomly selected and the pose trajectory is fitted within the viewport such that there are no poses outside the line of sight from the camera view. Additionally, by analyzing the trajectory variations, flipping motions are randomly performed on the sequence p to simulate motions of the backflip or side flip types. Figure 7 Represents the root position distribution of the original Human3.6M and the distribution after data augmentation as described herein, where the embodiments described herein have a wider position distribution, thus making the dataset more suitable for real-world scenarios.
[0103] Figure 7 Represents the distribution of the root position in the X-Z and Z-Y planes of the camera coordinates according to some embodiments. Image (a) is the original Human3.6M and image (b) is the augmentation described herein.
[0104] Additionally, during the training phase, 2D keypoint dropout and perturbation are applied to the input. During data sampling, 3D poses are projected to 2D using perspective projection. However, due to occlusion, 2D pose detectors tend to have noisy and missed detections. Methods such as VP3D and others use 2D detector results as noisy 2D inputs to train the model to be noise-resistant. Instead, as described in this paper, the 2D projected keypoints are perturbed by using Gaussian noise and randomly dropping keypoints to simulate occlusion scenarios. The Gaussian radius is adaptive and depends on the size of the body in the UV space. All keypoints labeled "dropout" are set to zero. Figure 8 Denotes the perturbation of the input and keypoint dropout. Diagram (a) is the original clean 2D pose, while diagrams (b)-(d) are noisy 2D poses with random dropout and perturbation applied.
[0105] Network details
[0106] As described in Equation 8, the goal is to learn a mapping function for a given input 2d pose u and camera parameters c to output the root position Bone lengths and their unit vectors To this end, 1D convolution and LSTM are combined to achieve stable prediction of the sequence. Figure 9 Denotes a simplified block diagram of the network as described in this paper according to some embodiments. The reason for using LSTM for the root position is that, similar to the relative pose estimation in the KCS space, 1D convolution with a kernel size of 3 was experimented with. However, its stability is affected, especially on the z-axis, which is a common problem in monocular 3D pose estimation. It is hypothesized that this is because 1D convolution cannot guarantee the estimation of the output at time t conditioned on the previous time t-1, regardless of whether a temporal loss function is applied. However, LSTM is able to pass the previous time features into the current features, which stabilizes the overall root position estimation.
[0107] On the input u with 512 and 1024 feature maps, there are two feature extraction blocks with four stacked residual connections, and these residual connections use 1D convolutions with a kernel size of 1. Here, the 1D convolution with a kernel size of 1 includes all time frames processed in a discrete manner, thereby mapping the feature space of each time frame. Then, the outputs of each block are concatenated with 1D convolutions with a kernel size of 3, such that padding is applied to the convolutions at the edges. The 1D convolution with a kernel size of 3 is used to aggregate adjacent frames. Different from VP3D which only outputs one frame out of N frames (the output of 243 frames to 1 frame is the best model of VP3D), padding is applied to all convolutions with a kernel size of 3, making the number of output frames equal to the number of input frames. The concatenation order is designed to first predict the bone length, then predict the bone unit vector based on this, and finally predict the root position. Each output is mapped back to the feature space using a convolution with a kernel size of 1, and then concatenated with the features extracted in the early stage to estimate the next prediction. This stems from how humans intuitively estimate the distance of an object, which is by first estimating the overall size of the object and its surrounding environment. It is found that separating the first feature extraction block for bone length and bone unit vector from the root position achieves better accuracy. The LSTM block has two recurrent layers, with 128 hidden units and is unidirectional. In some embodiments, all activation functions use the parametric ReLU.
[0108] Loss formula
[0109] The loss formula is described herein. First, the L2 loss is mainly applied to each output in the following form:
[0110]
[0111] where B is the combination of the bone length ||B|| and its unit vector In addition, the relative pose P r derived from Equation 7 is added, which involves adding more weights to the bone length and vector. The p0 term is used to encode the root position in both the encoding space and the decoding space, such that a smooth L1 loss with an amplitude of x2 on the z-axis. The reason for applying smooth L1 to the root position is that the loss in the decoding space will be large and may affect other loss ranges due to large errors. Applying the loss only in the encoding space does not perform as well as applying the loss in both the encoding space and the decoding space. In addition, a time term is added to the bone B and the root position p0:
[0112]
[0113] Among them, since the bone length does not change over time, the first term Δ||B|| above is zero. This enforces the bone length to be consistent over the time frames. For the root position, not only the increments of adjacent frames are adopted, but also the increments up to the third - order adjacent order and up to the second - order time derivative are used. Since the time difference is used to adjust the relative motion between frames, even if there may still be offset errors in the root position, it can converge to a smaller loss. However, this is important for trajectory tracking, especially for motion capture scenarios. Applying the 2D reprojection error,
[0114]
[0115] Note that this u is not the 2D pose input after the above perturbation, but the clean 2D projection of the ground - truth 3D pose. From the predicted 3D pose Derive the prediction Finally, the total loss is shown as follows:
[0116] L = L 3D + L 3DT + L 2D (14)
[0117] where each loss is added in the same way.
[0118] Experimental evaluation
[0119] Dataset and evaluation
[0120] Human3.6M contains 3.6 million video frames for 11 subjects, of which 7 are annotated with 3D poses. Following the same rules as other methods, these rules are divided into 5 subjects (S1, S5, S6, S7, S8) for training and 2 subjects (S9 and S11) for evaluation. Each subject performs 15 actions, which are recorded at a frequency of 50 Hz using four synchronized cameras. The mean per - joint position error (MPJPE) in millimeters is used, that is, the average Euclidean distance between the predicted joint positions and the ground - truth joint positions. Nevertheless, a slight change is made to how to aggregate the MPJPE when not all actions are averaged, so as to process all actions at once. For the root position, the mean position error (MPE) is evaluated, which is also the average Euclidean distance of the entire evaluation data. Human3.6M is evaluated by using the enhanced 15 - key - point and 17 - key - point definitions described in this paper. Perturbations are only applied to adding noise to the training data and key - point dropout, and are used together with camera and position augmentation for the evaluation set. The differences in key points are as Figure 4 shown.
[0121] Figure 9A simplified block diagram of the model described herein according to some embodiments is shown. As described herein, the ablation model variant without KC combines the blocks of xB1 and xB2 into one block and directly estimates the relative pose in Euclidean space, and the model without LSTM replaces xP from LSTM with 1D convolution.
[0122] In step 900, camera parameters (e.g., focal lengths with x and y in 2D space) are fed into the network so that the network can output the state of the camera. The network also receives 2D poses. The 2D poses can come from any image or video.
[0123] In step 902, feature extraction as described herein is applied frame by frame. The feature extraction can be implemented in any way. The feature extraction includes the determination of residuals using 1D convolution. Additionally, in some embodiments, after concatenation, a padded 1D convolution is implemented. In step 904, the bone length is estimated as described herein. In step 906, the bone unit vector can be estimated based on the feature extraction conditioned on the bone length. In step 908, the relative pose is estimated from the bone length and the bone unit vector, and the root position is derived based on the feature extraction conditioned on the bone length and the bone unit vector. In some embodiments, the camera parameters are used for the estimation of the root pose. LSTM can be used to help estimate the root position to stabilize the root position.
[0124] In some embodiments, fewer or additional steps can be implemented. In some embodiments, the order of the steps is modified.
[0125] Network variant
[0126] Experiments for the root pose for ablation studies have been performed on models with and without KCS and models with and without LSTM. The model without KCS directly regresses the relative pose in Euclidean space by using a 1D convolutional block with a kernel size of 3 followed by a kernel size of 1, such that the output dimension is several key points x3. Similarly, the model without LSTM regresses the root pose by using 1D convolution. All models are trained under the same training process. To compare with other methods, the method described herein is compared with the current state-of-the-art method VP3D.
[0127] Training
[0128] For the optimizer, Adam with weight decay set to zero is used and it is trained for 100 generations. From 1e -3Starting from the first generation, the learning rate is exponentially decayed every 10 generations with a factor of 0.5, where the first generation is for learning rate warm-up. Samples with a batch size of 192 and 121-frame input are used. During batch sampling, the frames randomly jump from 1 (non-skipped) to 5 among the 50Hz sampled frames of Human3.6M. This is to make the model robust to frame rate variations in outdoor videos. VP3D has been retrained using the same strategy as the model described in this paper, except that VP3D only accepts an input of 243 frames, so an input of 243 frames is used for VP3D instead of 121 frames. Compared with the above training process, neither batch normalization decay nor using Amsgrad with a decay of 0.95 proposed in VP3D shows poor performance on all models.
[0129] Figure 10 Represents the visualization of the root position prediction of the target according to some embodiments. The Z-axis has a larger error compared to the other axes, and there is a larger error for people at a far distance.
[0130] Evaluation and ablation study
[0131] Figure 11 A table showing the results of Human3.6M described in this paper plus the data augmentation scheme, where the camera FOV varies and has a much wider distribution in the root position. Since there is no alternative method to provide root position estimation, there is an exact comparison of the relative pose MPJPE. In addition, VP3D uses 243 frames to estimate 1 frame, while the model described in this paper is trained using 121 frames. Although the model described in this paper can adopt any frame size, for comparison under the same conditions, the evaluation is performed under the condition of 243-frame input, and the middle frame (the 121st frame) is evaluated. There are two variants, one applying KCS and the other adopting direct relative pose estimation. The model with KCS described in this paper has better MPJPE performance with much fewer parameters. This indicates that it may not be possible to implicitly infer camera parameter differences without queues. In addition, by studying the variants, the KCS method shows a significant advantage for directly estimating relative poses. It is also worth noting that even the root localization block is equivalent in these two methods. The MPE performance shows differences. By studying the training curve and validation error, the current hypothesis is that there are still fluctuations in the root localization performance.
[0132] For MPE, the root position error still seems to have a large error of about 20cm. This indicates that there are still difficulties in solving uncertain depths from monocular, especially only from 2D pose input. Figure 10Represents the overall projection error of a 15-keypoint pose model. X and Y show a reasonably good fit to the target, but Z shows errors as the target moves away and also some large errors at close range. The large errors at close range mainly come from the whole body being invisible due to the subject being too close to the camera (e.g., parts of the body are visible), but these situations occur in real-world scenarios. Although the experiments show that there is a lot of room to improve the MPE, the overall trajectory tracking will be observed, which is very important for the motion capture scenario.
[0133] The model without LSTM shows MPE comparable to or better than the LSTM model. To compare motion tracking, another evaluation is performed, in which all output frames are used instead of taking an intermediate frame that matches the input of VP3D. Thus, as Figure 12 shown, by observing the average trajectory error defined as the second term in Equation 12, the LSTM version shows better trajectory performance. The difference may become more significant when trying to reduce the model parameters. Figure 13 Represents a backflip sequence applied to a reduced version of the model described in this paper. The reduced version model uses LSTM and 1D convolution for root pose estimation. The 1D convolution estimates a large drift, especially on the Z-axis, which is very important for motion recovery. Figure 13 Represents the visualization of the Z-axis root position tracking of a sample sequence to compare models using LSTM and 1D convolution. The 1D convolution tends to have large tracking errors, especially on dynamic motions.
[0134] Figure 14A ~B represents a backflip video from YouTube with AlphaPose applied as a 2D pose detector, which is then executed using the method described in this paper. As shown by the X-Z plane reprojection, although the motion itself is very dynamic and there are many occlusions and errors in the 2D pose detector, the overall root position on the Z-axis is very stable. Figure 14A ~B represents the visualization of the output of the model described in this paper on an outdoor video, which is placed in two grouped columns, with each group showing 4 frames. Starting from the left, the video frames are for 2D pose estimation, 3D pose in the X-Y plane, and 3D pose in the X-Z plane. The red line on the 3D graph indicates the global trajectory. The model described in this paper is able to output a trajectory with a stable z position for dynamic motions. The 6th frame above has large errors in the 2D pose detection results.
[0135] Conclusion
[0136] The method described herein enables the recovery of full-skeleton 3D pose from a monocular camera, where the full-skeleton includes both the root position and relative pose in 3D. The model has significant advantages compared to the current state-of-the-art in the academic community to cover various FOVs and dynamic motions, such as backflips trained using only motion capture data. By making the model based on human perception rather than brute-force modeling of a large network and regressing values, the use of KCS and forming the model in the normalized space described herein yields better performance.
[0137] The method described herein takes as input only the 2D pose input normalized in the UV space and the basic camera parameters. Bone length estimation is trained with a very small distribution, and it is difficult to estimate the true bone length without the support of other queues (such as RGB images (e.g., appearance features)). It is assumed that the bone length can be derived from the ratio of the 2D bone lengths, where children tend to have a longer torso than arm bones. The height of a person can be roughly estimated based on the surrounding environment. Game engines (such as Unreal Engine) can be used to render images with relevant 3D geometries and perform end-to-end estimation of human 3D pose from the images. An original adversarial module has been constructed, which enables semi-supervised training using 2D annotations.
[0138] Figure 15A block diagram showing an exemplary computing device configured to implement a full-skeleton 3D pose recovery method according to some embodiments. The computing device 1500 can be used to acquire, store, calculate, process, transmit, and / or display information such as images and videos. The computing device 1500 can implement any aspect of full-skeleton 3D pose recovery. Generally, the hardware structure suitable for implementing the computing device 1500 includes a network interface 1502, a memory 1504, a processor 1506, one or more I / O devices 1508, a bus 1510, and a storage device 1512. The choice of the processor is not important as long as a suitable processor with sufficient speed is selected. The memory 1504 can be any conventional computer memory known in the art. The storage device 1512 can include a hard disk drive, a CDROM, a CDRW, a DVD, a DVDRW, a high-definition optical disc / drive, an ultra-high-definition drive, a flash card, or any other storage device. The computing device 1500 can include one or more network interfaces 1502. Examples of network interfaces include network cards connected to Ethernet or other types of LANs. One or more of the I / O devices 1508 can include one or more of the following: a keyboard, a mouse, a monitor, a screen, a printer, a modem, a touch screen, a button interface, and other devices. The one or more full-skeleton 3D pose recovery applications 1530 for implementing the full-skeleton 3D pose recovery method may be stored in the storage device 1512 and the memory 1504 and are generally processed as applications. More or fewer components as shown in Figure 15 may be included in the computing device 1500. In some embodiments, a full-skeleton 3D pose recovery hardware 1520 is included. Although Figure 15 the computing device 1500 in includes the application 1530 and the hardware 1520 for the full-skeleton 3D pose recovery method, the full-skeleton 3D pose recovery method can be implemented on the computing device in hardware, firmware, software, or any combination thereof. For example, in some embodiments, the full-skeleton 3D pose recovery application 1530 is programmed in the memory and executed by using the processor. In another example, in some embodiments, the full-skeleton 3D pose recovery hardware 1520 is a programmed hardware logic including gates specifically designed to implement the full-skeleton 3D pose recovery method.
[0139] In some embodiments, the one or more full-skeleton 3D pose recovery applications 1530 include several applications and / or modules. In some embodiments, a module further includes one or more sub-modules. In some embodiments, fewer or more modules may be included.
[0140] Examples of suitable computing devices include personal computers, laptop computers, computer workstations, servers, host computers, handheld computers, personal digital assistants, cellular / mobile phones, smart devices, gaming consoles, digital cameras, digital video cameras, camera phones, smartphones, portable music players, tablet computers, mobile devices, video players, video disc writers / players (e.g., DVD writers / players, high-definition disc writers / players, ultra-high-definition disc writers / players), televisions, home entertainment systems, augmented reality devices, virtual reality devices, smart jewelry (e.g., smartwatches), vehicles (e.g., autonomous vehicles), or any other suitable computing device.
[0141] To utilize the full-skeleton 3D pose recovery method described herein, a device such as a digital camera / video recorder is used to obtain content. The full-skeleton 3D pose recovery method can be implemented with user assistance or automatically without user participation to perform pose estimation.
[0142] In operation, the full-skeleton 3D pose recovery method provides a more accurate and efficient post-estimation implementation. Results show that better pose estimation can be achieved compared to previous implementations.
[0143] Some embodiments of full-skeleton 3D pose recovery from a monocular camera
[0144] 1. A method, comprising:
[0145] Receiving camera information, wherein the camera information includes two-dimensional poses and camera parameters including focal lengths;
[0146] Applying feature extraction to the camera information, including determining using one-dimensional convolutional residuals;
[0147] Estimating bone lengths based on the feature extraction;
[0148] Estimating bone unit vectors based on the feature extraction conditioned on the bone lengths; and
[0149] Estimating relative poses from the bone lengths and the bone unit vectors, and deriving root positions based on the feature extraction conditioned on the bone lengths and the bone unit vectors.
[0150] 2. The method according to clause 1, further comprising receiving one or more frames as input.
[0151] 3. The method according to clause 1, wherein it is assumed that each bone length does not exceed a length of 1 meter.
[0152] 4. The method according to clause 1, wherein long short-term memory is used to estimate the root position to stabilize the root position.
[0153] 5. The method according to clause 1 further includes applying auto-augmentation to the global position and rotation to simulate dynamic motion.
[0154] 6. The method according to clause 1 further includes randomly changing the field of view of the camera for each batch of samples to estimate any video using different camera parameters.
[0155] 7. The method according to clause 1 further includes performing perturbation of the two-dimensional pose input with Gaussian noise and random key point dropout to simulate the noise and occlusion conditions of two-dimensional pose prediction.
[0156] 8. An apparatus, comprising:
[0157] A non-transitory memory for storing an application that is configured to:
[0158] Receive camera information, where the camera information includes two-dimensional pose and camera parameters including focal length;
[0159] Apply feature extraction to the camera information, including determination using one-dimensional convolutional residuals;
[0160] Estimate bone length based on the feature extraction;
[0161] Estimate bone unit vectors based on the feature extraction conditional on the bone length; and
[0162] Estimate the relative pose from the bone length and the bone unit vectors, and derive the root position based on the feature extraction conditional on the bone length and the bone unit vectors; and
[0163] A processor coupled to the memory, the processor being configured to process the application.
[0164] 9. The apparatus according to clause 8 further includes receiving one or more frames as input.
[0165] 10. The apparatus according to clause 8, wherein it is assumed that the length of each bone does not exceed 1 meter.
[0166] 11. The apparatus according to clause 8, wherein long short-term memory is used to estimate the root position to stabilize the root position.
[0167] 12. The apparatus according to clause 8 further includes applying auto-augmentation to the global position and rotation to simulate dynamic motion.
[0168] 13. The apparatus according to clause 8 further includes randomly changing the field of view of the camera for each batch of samples to estimate any video using different camera parameters.
[0169] 14. The apparatus according to clause 8 further includes perturbing the two-dimensional pose input to perform a two-dimensional pose with Gaussian noise and random keypoint dropout to simulate the noise and occlusion conditions of two-dimensional pose prediction.
[0170] 15. A system includes:
[0171] a camera configured to acquire content; and
[0172] a computing device configured to:
[0173] receive camera information from the camera, where the camera information includes a two-dimensional pose and camera parameters including a focal length;
[0174] apply feature extraction to the camera information, including determining using one-dimensional convolutional residuals;
[0175] estimate bone length based on the feature extraction;
[0176] estimate a bone unit vector based on the feature extraction conditional on the bone length; and
[0177] estimate a relative pose from the bone length and the bone unit vector, and derive a root position based on the feature extraction conditional on the bone length and the bone unit vector.
[0178] 16. The system according to clause 15 further includes receiving one or more frames as input.
[0179] 17. The system according to clause 15, wherein it is assumed that each bone length does not exceed a length of 1 meter.
[0180] 18. The system according to clause 15, wherein long short-term memory is used to estimate the root position to stabilize the root position.
[0181] 19. The system according to clause 15 further includes applying auto-augmentation to the global position and rotation to simulate dynamic motion.
[0182] 20. The system according to clause 15 further includes randomly changing the field of view of the camera for each batch of samples to estimate any video using different camera parameters.
[0183] 21. The system according to clause 15 further includes perturbing the two-dimensional pose input to perform a two-dimensional pose with Gaussian noise and random keypoint dropout to simulate the noise and occlusion conditions of two-dimensional pose prediction.
[0184] The present invention has been described in terms of specific embodiments incorporating details, so as to facilitate understanding of the construction and operating principles of the present invention. Such reference to specific embodiments and their details herein is not intended to limit the scope of the claims appended hereto. It will be apparent to those skilled in the art that various other modifications can be made in the embodiments selected for illustration without departing from the spirit and scope of the present invention as defined by the claims.
Claims
1. A method for three-dimensional pose estimation, comprising: Receiving camera information, wherein the camera information includes a two-dimensional pose and camera parameters including a focal length; Applying feature extraction to the camera information, including determination using a one-dimensional convolutional residual; Estimating bone lengths based on the feature extraction; Estimating bone unit vectors based on the feature extraction conditional on the bone lengths; and Estimating a relative pose from the bone lengths and the bone unit vectors, and deriving a root position based on the feature extraction conditional on the bone lengths and the bone unit vectors.
2. The method according to claim 1, further comprising receiving one or more frames as an input.
3. The method according to claim 1, wherein Assuming that each bone length does not exceed a length of 1 meter.
4. The method according to claim 1, wherein Using a long short-term memory for estimating the root position to stabilize the root position.
5. The method according to claim 1, further comprising applying auto-augmentation to global position and rotation to simulate dynamic motion.
6. The method according to claim 1, further comprising randomly changing the camera field of view of each batch of samples to estimate any video using different camera parameters.
7. The method according to claim 1, further comprising performing perturbation of the two-dimensional pose with Gaussian noise and random keypoint dropout on the two-dimensional pose input to simulate noise and occlusion conditions in two-dimensional pose prediction.
8. An apparatus for three-dimensional pose estimation, comprising: A non-transitory memory for storing an application that is configured to: Receive camera information, wherein the camera information includes a two-dimensional pose and camera parameters including a focal length; Apply feature extraction to the camera information, including determination using a one-dimensional convolutional residual; Estimate bone lengths based on the feature extraction; Estimate bone unit vectors based on the feature extraction conditional on the bone lengths; and Estimate a relative pose from the bone lengths and the bone unit vectors, and derive a root position based on the feature extraction conditional on the bone lengths and the bone unit vectors; and A processor coupled to the memory, the processor being configured to process the application.
9. The apparatus according to claim 8, further comprising receiving one or more frames as an input.
10. The device according to claim 8, wherein, Assuming that each bone length does not exceed a length of 1 meter.
11. The device according to claim 8, wherein, Using a long short-term memory for estimating the root position to stabilize the root position.
12. The apparatus according to claim 8, further comprising applying auto-augmentation to global position and rotation to simulate dynamic motion.
13. The apparatus according to claim 8, further comprising randomly changing the camera field of view of each batch of samples to estimate any video using different camera parameters.
14. The apparatus according to claim 8, further comprising performing perturbation of the two-dimensional pose with Gaussian noise and random keypoint dropout on the two-dimensional pose input to simulate noise and occlusion conditions in two-dimensional pose prediction.
15. A system for three-dimensional pose estimation, comprising: A camera configured to acquire content; And A computing device configured to: Receive camera information from the camera, wherein the camera information includes a two-dimensional pose and camera parameters including a focal length; Apply feature extraction to the camera information, including determination using a one-dimensional convolutional residual; Estimate bone lengths based on the feature extraction; Estimate a bone unit vector based on feature extraction conditional on the bone length; and Estimate a relative pose from the bone length and the bone unit vector, and derive a root position based on feature extraction conditional on the bone length and the bone unit vector.
16. The system according to claim 15, further comprising receiving one or more frames as input.
17. The system according to claim 15, wherein, Assume that each bone length does not exceed a length of 1 meter.
18. The system according to claim 15, wherein Use a long short-term memory for estimating the root position to stabilize the root position.
19. The system according to claim 15, further comprising applying auto-augmentation to global position and rotation to simulate dynamic motion.
20. The system according to claim 15, further comprising randomly changing the camera field of view of each batch of samples to estimate any video using different camera parameters.
21. The system according to claim 15, further comprising performing perturbation of a two-dimensional pose with Gaussian noise and random keypoint dropout on a two-dimensional pose input to simulate noise and occlusion conditions of two-dimensional pose prediction.