Head-Mounted Camera and IMU Pose Estimation in Unrestricted Environments
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for tracking human poses are limited to defined spaces and require high computational effort, especially when using third-person cameras or wearable sensors, and fail to accurately estimate skeletal data in unrestricted environments.
Innovation Solution
A method using a head-mounted camera and body sensors, such as IMUs, to create depth maps and skeletal data, aligning them through kinematic chains and coordinate frames to estimate poses, allowing tracking in unknown environments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a third-person camera is used to track human movements, then skeletal data can be determined, but tracking is limited to defined spaces and suffers from occlusion problems
Solution Approach 1:
The patent inverts the traditional third-person camera approach by using a first-person head-mounted camera. This inversion allows the system to work in unrestricted environments including outdoor and unknown spaces, eliminating occlusion issues while maintaining skeletal data accuracy through sensor fusion with IMUs and depth maps.
Solution Approach 2:
The patent introduces depth maps as an intermediary element to bridge the first-person camera view and skeletal pose estimation. The depth map provides three-dimensional spatial information that, when combined with IMU sensor data, enables accurate pose estimation in environments where traditional camera-based methods fail.
2Adaptability or versatility
If a wide-angle camera attached to the person's chest is used to track body parts, then motion tracking is enabled in wider spaces, but high computational effort is required to estimate skeletal data with appropriate accuracy
Solution Approach 1:
The patent segments the pose estimation problem into multiple components: depth map generation from first-person camera, skeletal pose estimation from IMU sensors, and fusion of both data sources. This segmentation allows each component to be optimized independently, reducing overall computational effort compared to using only wide-angle camera-based methods.
Solution Approach 2:
The patent replaces the purely optical wide-angle camera system with a hybrid system that incorporates inertial sensors (IMUs). This substitution leverages mechanical sensing to capture motion data, reducing the computational burden on the camera system while maintaining accuracy in wide spatial coverage.
3Adaptability or versatility
If only parts of the target's body are tracked using wearable sensors, then tracking is enabled in unrestricted environments, but high computational effort is required to estimate complete skeletal data
Solution Approach 1:
The patent makes the first-person camera system multi-functional by using it both for navigation/scene understanding (through depth maps) and for pose estimation (when combined with IMU data). This universality allows the system to achieve complete skeletal tracking with reduced computational complexity compared to dedicated sensor-only approaches.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
Enables accurate estimation of human poses and interactions with surrounding objects in unrestricted environments, improving working quality and safety by providing real-time guidance and warnings.
Implementation Method 1
camera data from the camera attached to a head of the person is received and a depth map is created based on the received camera data
Implementation Method 2
one or more signals from the one or more sensors attached to a body of the person are received based on which skeletal data is created. The one or more sensors may preferably be inertial measurement units (IMU)
Data Source
Figure 1
Figure 2
Figure 3
AI summary
The present subject matter particularly relates to a method, a computer program product and an apparatus 10 for estimating one or more poses of a person and a system including the apparatus 10. The method comprising the steps of receiving camera data from a camera 100 attached to a head of the person, creating a depth map based on the received camera data to generate three-dimensional camera data 1100, receiving one or more signals from the one or more sensors 200 attached to a body of the person, creating skeletal data 2100 based on the one or more received signals, determining the orientation of the camera 100 attached to the head of the person using the created skeletal data 2100, aligning the three-dimensional camera data 1100 and the skeletal data 2100 based on the determined orientation of the camera 100 and a position of the camera 100, and estimating a pose of the person based on the aligned three-dimensional camera data 1100 and the skeletal data 100.