XR Camera Selection Using Predicted Skeletons for Gesture Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing XR devices with multiple image sensors face resource-intensive gesture recognition processing due to occlusion issues, necessitating a more efficient method to select the optimal camera for high accuracy without excessive computational load.
Innovation Solution
The XR device selects a camera for gesture recognition based on predicted skeletons using skeleton-based rules or neural-network predictions, optimizing the choice to minimize occlusion and reduce resource usage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If all cameras are processed simultaneously for gesture recognition, then measurement precision is improved, but use of energy increases
Solution Approach 1:
The patent segments the gesture recognition process into two distinct phases: a coarse filtering phase that processes all cameras to identify candidate cameras, and a fine recognition phase that processes only the selected candidate camera. This segmentation allows the system to maintain high recognition accuracy while significantly reducing computational resource consumption by avoiding full processing of all cameras.
Solution Approach 2:
The patent performs preliminary action by generating predicted skeletons for all cameras and using these predictions to pre-select candidate cameras before the actual gesture recognition process. This preliminary selection based on skeleton predictions filters out cameras with poor viewing angles or occlusions, ensuring that only promising candidates undergo resource-intensive classification processing.
2Reliability
If all cameras are processed simultaneously for gesture recognition, then reliability is improved, but productivity deteriorates
Solution Approach 1:
The patent divides the processing workflow into segmentation stages: initial skeleton prediction for all cameras, candidate selection based on prediction quality, and final gesture classification only for selected candidates. This segmented approach maintains reliability by ensuring comprehensive initial assessment while improving productivity through focused processing of only the most promising camera feeds.
Solution Approach 2:
The system performs preliminary skeleton prediction and camera selection before committing to full gesture recognition processing. This preliminary action identifies and filters out cameras with poor viewing conditions, ensuring that subsequent processing is performed only on reliable candidates, thus maintaining high recognition reliability while significantly improving processing efficiency.
3Device complexity
If camera selection is performed without optimization, then device complexity is reduced, but measurement precision deteriorates
Solution Approach 1:
The patent implements a relatively simple preliminary action of generating predicted skeletons for each camera and selecting candidates based on prediction quality metrics. This approach maintains low device complexity by using straightforward skeleton-based filtering rather than complex optimization algorithms, while still achieving high measurement precision by ensuring that selected cameras have good viewing angles and minimal occlusions.
Data Source
AI summary
A head-worn device system includes multiple image sensors (e.g., cameras), one or more display devices and one or more processors. The system also includes a memory storing instructions that, when executed by the one or more processors, configure the system to obtain a first image captured by a first image sensor of the device; generate, based on obtaining the first image, a respective predicted skeleton corresponding to a respective view from each of the first image sensor and the one or more second image sensors; select, based on generating the respective predicted skeletons, an image sensor from among the first image sensor and the one or more second image sensors, the selected image sensor being used for classifying the first image; and determine, based on the selected image sensor, a classification for the first image.


