Monocular 3D Human Posture Estimation From Smartphone Video
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for 3-D posture estimation from 2-D images are inaccurate, require depth sensors or multiple cameras, and fail to leverage temporal information effectively, leading to high resource consumption and poor temporal posture transitions.
Innovation Solution
A neural network model using an autoencoder and a second encoder, trained with monocular video from a single low-end device, employs residual network blocks and 1-D convolution to estimate 3-D posture in real-time with high accuracy, eliminating the need for depth sensors and leveraging temporal information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional CV-based 3-D posture monitoring techniques are used, then depth information can be obtained, but device complexity increases due to requirement of depth sensor along with RGB camera
Solution Approach 1:
The patent replaces the mechanical depth sensing system (depth sensor + RGB camera) with a computational approach using a single RGB camera. The neural network model processes 2D images and temporal sequences to infer 3D posture, substituting physical depth measurement hardware with algorithmic reconstruction from image data.
Solution Approach 2:
The patent creates a virtual 3D representation of the human body by copying spatial relationships from 2D image sequences. The neural network learns to map 2D pixel coordinates to 3D joint positions by analyzing temporal patterns and geometric constraints, effectively copying 3D structure from 2D projections without depth sensors.
2Adaptability or versatility
If parametric forms such as SMPL models are used, then entire 3-D shape including postures can be recovered, but memory and processing power requirements increase due to larger parameter size
Solution Approach 1:
The patent extracts only the essential posture information from the complex SMPL parametric model. Instead of recovering entire 3D shapes with numerous parameters, the neural network focuses on estimating key joint positions and skeletal structure, taking out only the necessary spatial transformation parameters while discarding redundant shape parameters.
Solution Approach 2:
The patent segments the complex posture estimation problem into independent joint position estimation tasks. The neural network processes each joint's 3D coordinates separately based on its 2D image position and temporal patterns, dividing the overall 3D reconstruction into manageable spatial segments rather than using a monolithic parametric model.
3Loss of time
If techniques estimating 3-D posture from static images are used, then computation time is reduced, but temporal information is not exploited leading to poor smooth transition of postures
Solution Approach 1:
The patent performs preliminary temporal pattern learning during the training phase by analyzing sequences of 3D postures from video data. The neural network pre-learns the temporal relationships and smooth transition patterns between consecutive postures, storing this information in the model's weight parameters for efficient inference.
Solution Approach 2:
The patent maintains continuous temporal context by processing video frames as sequences rather than independently. The neural network uses recurrence or attention mechanisms to maintain continuous representation of body posture over time, ensuring smooth transitions by leveraging temporal continuity in the input video data.
4Reliability
If techniques exploiting temporal information are used, then posture transition smoothness improves, but time requirements increase due to two state computations
Solution Approach 1:
The patent merges spatial and temporal processing into a unified neural network architecture that simultaneously handles both dimensions. Instead of separate stages for spatial posture estimation and temporal smoothing, the model integrates both operations, processing spatial and temporal relationships in parallel within the same computational graph.
Solution Approach 2:
The patent transforms the temporal processing problem into a spatial transformation problem by learning 3D pose embeddings that inherently capture temporal information. The neural network maps 2D image sequences to 3D posture representations in a latent space where temporal relationships are encoded as geometric transformations, converting temporal computation into more efficient spatial operations.
Data Source
AI summary
The disclosure relates generally to methods and systems for real time video driven human 3-dimensional (3-D) posture estimation during physical activities. Conventional techniques do not exploit temporal information, they do not give smooth transition of postures over time. Furthermore, the techniques that exploit the temporal information suffer from higher time requirements due to two state computations. The present disclosure solves the technical problems in the art with the methods and systems for real time video driven human 3-D posture estimation during physical activities. The present invention discloses a smart-phone camera based automatic posture monitoring system designed with an auto-encoder based architecture. The disclosed auto-encoder based cross-modal method uses monocular video (2-D image sequences) from a single low-end mobile device (for example, smart-phone camera) for estimating human 3-D posture in real time (˜5 fps) with high accuracy (less than 1 cm error per joint location).


