Virtual Hand Interaction via Segmented Back-of-Hand Projection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing virtual reality and mixed reality display systems face challenges in providing immersive, low-latency, and natural touch-like interactions due to high latency, uncomfortable user positioning, and counterintuitive hand representation, which affects the realism and intuitiveness of hand interactions with virtual objects.
Innovation Solution
The system captures the back of the user's hand using RGB and depth cameras, generates segmented images, and projects them into the virtual environment, maintaining a consistent virtual orientation and size, allowing for realistic and intuitive hand interactions with minimal latency by avoiding computationally intensive processes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If precise hand tracking is used to enable direct physical-virtual interactions, then interaction accuracy is improved, but latency increases significantly
Solution Approach 1:
The patent extracts only the essential information needed for interaction (hand position and orientation) rather than performing complete precise hand tracking. This selective extraction reduces computational load and latency while maintaining sufficient interaction accuracy, directly resolving the contradiction between measurement precision and time loss.
2Ease of operation
If the user raises their hand toward the capture device for silhouette capture, then hand interaction is enabled, but user comfort deteriorates and view obstruction increases
Solution Approach 1:
Instead of capturing the inner-palm view by requiring the user to raise their hand toward the device, the system inverts the approach by capturing the back-of-hand view with a downward-facing camera. This allows the user to keep their hand in a natural, comfortable position below the display while still enabling hand interaction, eliminating view obstruction and discomfort.
3Loss of information
If inner-palm silhouette is captured and presented, then hand representation is achieved, but interaction intuitiveness deteriorates
Solution Approach 1:
The system creates a visual copy of the user's hand (the back-of-hand image) and presents it in the virtual environment. This copy maintains the natural appearance and orientation expectations of the user's hand, making the interaction intuitive. The copied hand image is displayed at the appropriate virtual position corresponding to the actual hand position, preserving interaction intuitiveness while enabling hand representation.
4Measurement precision
If hand moves closer to sensor for better capture, then capture quality improves, but silhouette realism deteriorates due to counterintuitive size change
Solution Approach 1:
The system resolves the size contradiction by changing the dimensional reference frame. Instead of allowing silhouette size to change with distance from the sensor (2D projection effect), the system maintains the silhouette size consistent with the user's perspective (as if viewing their own hand). This creates a perceptually accurate representation where the hand silhouette size remains constant regardless of distance from the sensor, matching real-world visual expectations.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Embodiments that relate to providing a low-latency interaction in a virtual environment are provided. In one embodiment an initial image of a hand (74) and initial depth information representing (76) an initial actual position are received (350). The initial image is projected into the virtual environment (38) to an initial virtual position (78). A segmented version (80) of the initial image (74) is provided for display (312) in the virtual environment (308) at the initial virtual position (78). A subsequent image (82) of the hand and depth information (84) representing a subsequent actual position (370) are received. The subsequent image (82) is projected into the virtual environment (308) to a subsequent virtual position (86). A segmented version (92) of the subsequent image (82) is provided for display at the subsequent virtual position (86). A collision is detected between a three-dimensional representation of the hand and a virtual or physical object.