Markerless 3D Human Pose Capture for Real-Time Mobile Avatars
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current human pose estimation methods require expensive industrial-grade equipment and physical markers, which are inconvenient and limit outdoor use, and existing markerless algorithms are unsuitable for real-time processing on handheld devices.
Innovation Solution
A method for human pose estimation on mobile devices that identifies 2D and 3D joint positions from monocular RGB images, using a differentiable spatial-to-numerical transform layer and a 3D human pose estimation network to convert 2D poses into 3D, enabling real-time avatar rendering without markers.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If industrial grade imaging equipment with physical markers is used, then measurement precision is improved, but device complexity and ease of operation deteriorate
Solution Approach 1:
The patent extracts and eliminates the physical markers from the system, achieving markerless pose estimation using only standard RGB cameras. This removes the complexity of marker attachment and tracking while maintaining pose estimation accuracy through advanced algorithms that process natural human body appearance features.
Solution Approach 2:
The patent creates a virtual copy of the physical marker system by using image processing and deep learning models to infer pose information from 2D camera images. The system synthesizes 3D pose data from 2D projections, effectively replacing physical markers with computational models that achieve similar measurement precision without the physical infrastructure.
2Measurement precision
If multiple optical or depth cameras with multiple viewing angles are used, then measurement precision is improved, but device complexity and ease of operation worsen
Solution Approach 1:
The patent extracts the pose estimation function from complex multi-camera systems and implements it using a single standard RGB camera. By removing the need for multiple viewing angles and depth cameras, the system maintains accuracy through sophisticated 2D-to-3D pose conversion algorithms while dramatically reducing system complexity.
Solution Approach 2:
The patent replaces the mechanical multi-camera system with a computational approach using a single camera. Instead of using multiple physical sensors to capture pose information, the system uses deep learning models and spatial transform algorithms to derive 3D pose from 2D images, substituting mechanical complexity with computational intelligence.
3Measurement precision
If markerless algorithms are executed offline on personal computers, then measurement precision is improved, but productivity and ease of operation worsen
Solution Approach 1:
The patent transforms the pose estimation system from static offline processing to dynamic real-time processing. By optimizing the computational algorithms and leveraging mobile device hardware capabilities, the system achieves real-time pose estimation on handheld devices, enabling dynamic capture and processing without requiring powerful desktop computers.
Solution Approach 2:
The patent changes the computational parameters and optimization levels to enable real-time execution on mobile devices. By adjusting model complexity, using quantized precision, and optimizing for mobile hardware architectures, the system maintains high measurement precision while achieving real-time processing speeds suitable for handheld devices.
4Ease of operation
If optical cameras are used for markerless pose estimation, then ease of operation is improved, but measurement precision deteriorates in outdoor environments
Solution Approach 1:
The patent changes the operational parameters of the pose estimation system to work effectively in outdoor lighting conditions. By adjusting the deep learning models to be robust against varying illumination, using techniques like image normalization and adaptive feature extraction, the system maintains high measurement precision while preserving the ease of markerless operation in diverse environments.
Data Source
AI summary
A computer system obtains an image of a scene captured by a camera and identifies a two-dimensional (2D) pose of the person in the image. The 2D pose includes a plurality of 2D joint positions in the image. The 2D pose is converted to a three-dimensional (3D) pose of the person including a plurality of 3D joint positions. The computer system determines a rotation angle of each joint relative to a T-pose of the person based on the plurality of 3D joint positions. The rotation angle of each joint is applied to a skeleton template of the avatar. The computer system renders the skeleton template of the avatar having the rotation angle for each joint.


