Virtual Scene Input Recognition via Binocular Hand Tracking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing virtual reality and augmented reality systems require users to interact with real-world controllers or special sensor devices for input operations, reducing immersion and realism.
Innovation Solution
An input recognition method using a binocular camera to identify hand key points, calculate fingertip coordinates through a binocular positioning algorithm, and compare them with virtual input interfaces to determine user input operations without additional hardware.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a controller or special sensor device is used to determine user interaction with the virtual interface, then the input recognition accuracy is improved, but the hardware cost increases and user immersion deteriorates
Solution Approach 1:
The patent uses a binocular camera to capture images of the user's hand and creates a virtual model of the hand with key points corresponding to finger joints and tips. This virtual copy allows the system to track hand movements and determine input operations without requiring physical controllers or special sensor devices, thereby reducing hardware costs while maintaining input recognition accuracy.
Solution Approach 2:
The patent replaces the mechanical system of physical controllers and special sensors with an optical system using a binocular camera. By capturing images and processing them through algorithms that identify hand key points and calculate three-dimensional coordinates, the system substitutes mechanical input devices with optical detection, reducing hardware complexity while preserving input recognition capabilities.
2Measurement precision
If a controller or special sensor device is used to determine user interaction with the virtual interface, then the input recognition accuracy is improved, but the user immersion deteriorates
Solution Approach 1:
The system creates a virtual representation of the user's hand with key points that correspond to actual anatomical landmarks. This virtual copy enables direct tracking of natural hand movements within the virtual scene, allowing users to interact with virtual interfaces using their own hands rather than external controllers, thereby maintaining input accuracy while significantly improving immersion.
Solution Approach 2:
The user's own hand serves as the input device. By tracking the user's natural hand movements and finger positions through the binocular camera system, the hand itself becomes the sensor, eliminating the need for separate control devices and enhancing the sense of presence and immersion in the virtual environment.
3Ease of operation
If hand key points are identified from binocular images and fingertip coordinates are calculated, then user immersion is improved, but the algorithm complexity increases
Solution Approach 1:
The patent segments the hand into multiple key points, including finger joint points and fingertip points. By dividing the hand tracking task into identification of discrete key points followed by calculation of three-dimensional coordinates for each point, the algorithm processes manageable segments rather than attempting to track the entire hand as a single complex object, thereby reducing computational complexity while maintaining immersion.
Solution Approach 2:
The patent introduces hand key points as intermediary elements between the raw binocular images and the final three-dimensional fingertip coordinates. These key points serve as intermediate representations that simplify the coordinate calculation process, acting as mediators that bridge image processing and spatial localization without requiring complex direct computation.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
Enhances user immersion and realism by allowing input operations directly in the virtual scene without the need for external devices, reducing hardware costs and improving interaction efficiency.
Implementation Method 1
a binocular image obtained by taking a hand by a binocular camera
Implementation Method 2
calculating a fingertip coordinate by using a binocular positioning algorithm, based on a position of the hand key point in the binocular image
Data Source
AI summary
Provided in the embodiments of the present disclosure are an input recognition method in a virtual scene, a device and a storage medium. On the basis of recognized position of hand key point, fingertip coordinate can be calculated by using a binocular positioning algorithm, the fingertip coordinate is compared with at least one virtual input interface in the virtual scene, and if fingertip position and a target virtual input interface in the at least one virtual input interface satisfy a set position rule, it is determined that a user executes an input operation by means of the target virtual input interface. In this way, a fingertip position of a user can be calculated, and the user does not need to interact with a controller or a special sensor device in the real world, further enhancing a sense of immersion and a sense of reality of the virtual scene.


