Markerless Hand Motion Capture via Multi-Engine Pose Estimation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Markerless motion capture systems struggle to accurately track the fine movements of hands, resulting in lower fidelity and fewer tracked joints compared to marker-based systems, which can lead to unnatural movement representations.
Innovation Solution
A system utilizing multiple computer vision-based pose estimation engines processes multiple views to capture hand motion without markers, generating a pose for the entire subject and performing additional pose estimation on extracted hand regions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If markerless motion capture is used, then hardware usage is reduced and cost is lowered, but measurement precision and tracking fidelity of hand movements deteriorate
Solution Approach 1:
The system segments the motion capture task into two parts: a first pose estimation engine processes the entire image to provide overall pose information, while a second pose estimation engine specifically processes extracted hand regions to provide detailed hand pose information. This segmentation allows the system to maintain high measurement precision for hand movements while avoiding the need for complex marker-based hardware.
Solution Approach 2:
The system extracts the hand region of interest from the full-body pose estimation results and subjects it to additional pose estimation processing. By taking out the hand region for specialized analysis, the system achieves detailed hand movement tracking without requiring markers on the subject's hands, thus resolving the contradiction between hardware reduction and measurement precision.
2Measurement precision
If marker-based motion capture is used, then measurement precision and tracking fidelity are improved, but device complexity and hardware requirements increase
Solution Approach 1:
The patent replaces the mechanical marker-based system with a computational vision-based system. Instead of using physical markers attached to the subject's hands, the system uses computer vision algorithms (pose estimation engines) to detect and track hand movements through image processing, thereby reducing hardware complexity while maintaining or improving measurement precision.
Solution Approach 2:
The system creates a digital copy of the hand region from the full-body pose estimation and applies additional pose estimation to this copied region. This copying approach allows detailed hand analysis without requiring physical markers on the actual subject, reducing hardware requirements while maintaining measurement precision.
3Device complexity
If single pose estimation engine is used for entire subject, then device complexity is reduced, but measurement precision for fine movements like hands deteriorates
Solution Approach 1:
The system segments the pose estimation process into two specialized engines: one for overall body pose and another specifically for hand pose estimation. This segmentation enables the hand-specific engine to focus computational resources on detailed hand joint tracking, improving measurement precision without requiring an overly complex monolithic system.
Solution Approach 2:
The patent applies local quality by using a second pose estimation engine specifically for the hand region, which receives enhanced detail from the first engine's full-body analysis. This localized specialized processing ensures high measurement precision for hand movements while keeping the overall system architecture manageable through modular design.
Data Source
AI summary
An example of an apparatus for markerless motion capture is provided. The apparatus includes cameras to capture images of a subject from different perspectives. In addition, the apparatus includes a pose estimation engines to receive the images. Each pose estimation engine is to generate a coarse skeletons of the received image and is to identify a region of the image based on the coarse skeleton. Furthermore, the apparatus includes pose estimation engines to receive the regions of interest previously identified. Each of these pose estimation engines is to generate a fine skeleton of the region of interest. In addition, the apparatus includes attachment engines to generate a whole skeletons. Each whole skeleton is to include a fine skeleton attached to a coarse skeleton. The apparatus further includes an aggregator to receive the whole skeletons. The aggregator is to generate a three-dimensional skeleton from the whole skeletons.


