XR Input Sign Tracking via Handheld Device and Hand Fusion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing extended reality (XR) interaction methods, such as using virtual buttons in head-mounted devices, are not straightforward for users and can be exhausting.
Innovation Solution
A system that tracks an input sign using a combination of image capture, bounding box detection, and data fusion of hand and handheld device positions, allowing for intuitive interaction with XR environments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If virtual buttons in head-mounted devices are used for interaction, then the user can operate the device, but the interaction method becomes complex and exhausting for the user
Solution Approach 1:
The patent replaces the mechanical interaction system (physical virtual buttons in head-mounted devices) with an optical recognition system that detects hand gestures and handheld device positions through image capture and bounding box analysis. This substitution eliminates the need for users to physically interact with complex virtual interfaces, thereby improving ease of operation while reducing interaction complexity.
Solution Approach 2:
The patent uses image capture to create a visual copy of the user's hand and handheld device, then processes this copy through bounding box detection and data fusion to recognize input signs. This copying approach allows the system to interpret natural hand movements without requiring users to learn complex virtual button operations, resolving the contradiction between ease of operation and interaction complexity.
2Ease of operation
If traditional virtual object interaction is used, then the device can be operated, but the method is not straightforward and may exhaust the user
Solution Approach 1:
The patent enables the system to automatically detect and interpret user input signs through image processing and data fusion of hand and handheld device positions. The system serves itself by autonomously recognizing gestures and converting them into commands without requiring users to navigate complex virtual interfaces, thereby making interaction more straightforward and reducing the time and effort users must invest.
Solution Approach 2:
The patent introduces an intermediary processing layer that fuses data from hand bounding boxes and handheld device bounding boxes to recognize input signs. This intermediary data fusion mechanism translates natural hand movements into meaningful commands, making the interaction process more straightforward and efficient compared to direct virtual button manipulation.
3Ease of operation
If hand gesture recognition is implemented, then interaction becomes more intuitive, but tracking accuracy may be insufficient without additional data fusion
Solution Approach 1:
The patent merges multiple data sources by performing data fusion between hand bounding box information and handheld device bounding box information. This combination of multiple tracking sources enhances measurement precision and tracking accuracy while maintaining the intuitiveness of hand gesture-based interaction, thereby resolving the contradiction between ease of operation and measurement precision.
Data Source
AI summary
A system and a method of tracking an input sign for an extended reality are provided, wherein the method including: obtaining an image; detecting for a handheld device and a hand in the image; in response to a first bounding box of the hand and a second bounding box of the handheld device being detected, detecting at least one joint of the hand from the image; performing a data fusion of the first bounding box and the second bounding box according to the at least one joint to obtain the input sign; and outputting a command corresponding to the input sign via the output device.


