Monocular 3D Human Posture Estimation From Smartphone Video

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for 3-D posture estimation from 2-D images are inaccurate, require depth sensors or multiple cameras, and fail to leverage temporal information effectively, leading to high resource consumption and poor temporal posture transitions.

Innovation Solution

A neural network model using an autoencoder and a second encoder, trained with monocular video from a single low-end device, employs residual network blocks and 1-D convolution to estimate 3-D posture in real-time with high accuracy, eliminating the need for depth sensors and leveraging temporal information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional CV-based 3-D posture monitoring techniques are used, then depth information can be obtained, but device complexity increases due to requirement of depth sensor along with RGB camera

Engineering Contradiction:
Improvedepth estimation accuracyVSAvoidcamera system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces the mechanical depth sensing system (depth sensor + RGB camera) with a computational approach using a single RGB camera. The neural network model processes 2D images and temporal sequences to infer 3D posture, substituting physical depth measurement hardware with algorithmic reconstruction from image data.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent creates a virtual 3D representation of the human body by copying spatial relationships from 2D image sequences. The neural network learns to map 2D pixel coordinates to 3D joint positions by analyzing temporal patterns and geometric constraints, effectively copying 3D structure from 2D projections without depth sensors.

Inventive Principle:
Principle #26Copying

2Adaptability or versatility

If parametric forms such as SMPL models are used, then entire 3-D shape including postures can be recovered, but memory and processing power requirements increase due to larger parameter size

Engineering Contradiction:
Improve3-D shape recovery capabilityVSAvoidparameter size
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent extracts only the essential posture information from the complex SMPL parametric model. Instead of recovering entire 3D shapes with numerous parameters, the neural network focuses on estimating key joint positions and skeletal structure, taking out only the necessary spatial transformation parameters while discarding redundant shape parameters.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent segments the complex posture estimation problem into independent joint position estimation tasks. The neural network processes each joint's 3D coordinates separately based on its 2D image position and temporal patterns, dividing the overall 3D reconstruction into manageable spatial segments rather than using a monolithic parametric model.

Inventive Principle:
Principle #1Segmentation

3Loss of time

If techniques estimating 3-D posture from static images are used, then computation time is reduced, but temporal information is not exploited leading to poor smooth transition of postures

Engineering Contradiction:
Improvecomputation timeVSAvoidtemporal posture transition smoothness
Core Design Contradiction:
Loss of timeVSReliability

Solution Approach 1:

The patent performs preliminary temporal pattern learning during the training phase by analyzing sequences of 3D postures from video data. The neural network pre-learns the temporal relationships and smooth transition patterns between consecutive postures, storing this information in the model's weight parameters for efficient inference.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent maintains continuous temporal context by processing video frames as sequences rather than independently. The neural network uses recurrence or attention mechanisms to maintain continuous representation of body posture over time, ensuring smooth transitions by leveraging temporal continuity in the input video data.

Inventive Principle:
Principle #20Continuity of useful action

4Reliability

If techniques exploiting temporal information are used, then posture transition smoothness improves, but time requirements increase due to two state computations

Engineering Contradiction:
Improvetemporal posture transition smoothnessVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent merges spatial and temporal processing into a unified neural network architecture that simultaneously handles both dimensions. Instead of separate stages for spatial posture estimation and temporal smoothing, the model integrates both operations, processing spatial and temporal relationships in parallel within the same computational graph.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent transforms the temporal processing problem into a spatial transformation problem by learning 3D pose embeddings that inherently capture temporal information. The neural network maps 2D image sequences to 3D posture representations in a latent space where temporal relationships are encoded as geometric transformations, converting temporal computation into more efficient spatial operations.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS20250336236A1Methods and systems for real time video driven human 3-d posture estimation
Publication Date: 2025.10.30 TATA CONSULTANCY SERVICES LTD
  • US20250336236A1 patent drawing
  • US20250336236A1 patent drawing
  • US20250336236A1 patent drawing

AI summary

The disclosure relates generally to methods and systems for real time video driven human 3-dimensional (3-D) posture estimation during physical activities. Conventional techniques do not exploit temporal information, they do not give smooth transition of postures over time. Furthermore, the techniques that exploit the temporal information suffer from higher time requirements due to two state computations. The present disclosure solves the technical problems in the art with the methods and systems for real time video driven human 3-D posture estimation during physical activities. The present invention discloses a smart-phone camera based automatic posture monitoring system designed with an auto-encoder based architecture. The disclosed auto-encoder based cross-modal method uses monocular video (2-D image sequences) from a single low-end mobile device (for example, smart-phone camera) for estimating human 3-D posture in real time (˜5 fps) with high accuracy (less than 1 cm error per joint location).