Head-Mounted Display Sound Identification via Motion and Voice Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing head-mounted display devices face challenges in accurately identifying voice intonation, responding to user intentions, and integrating virtual images with real-world scenes, leading to reduced convenience and increased size and cost.
Innovation Solution
A transmission type head-mounted display device equipped with a sound acquiring unit, sound identifying unit, image display unit, image storing unit, and function executing unit, which combines execution function images with specific sound images to improve sound identification accuracy and user convenience, allowing for the execution of functions via short sounds and motion detection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If voice intonation is used to identify user intent, then the system can respond to user commands, but the intonation identification accuracy deteriorates
Solution Approach 1:
The patent segments the sound identification task by combining multiple sound sources (microphone for voice, acceleration sensor for motion) to identify user intent. Instead of relying solely on voice intonation, the system divides the identification into multiple channels: voice commands and physical motions, thereby improving overall identification accuracy while maintaining ease of operation
Solution Approach 2:
The patent introduces motion detection as an intermediary element between voice command and system response. The acceleration sensor detects physical motions (such as shaking the device) that serve as alternative or complementary commands, mediating the interaction between user intent and system execution, thereby reducing reliance on difficult-to-identify voice intonation
2Adaptability or versatility
If multiple functions are assigned to different sound types, then the system can execute diverse functions, but the number of sound images increases
Solution Approach 1:
The patent makes the motion detection capability universal by using the same acceleration sensor for multiple functions. The same sensor that detects device shaking can also detect head movements, hand gestures, or other physical actions, allowing a single hardware component to serve multiple command functions, thereby maintaining function diversity without proportionally increasing the number of sound images
Solution Approach 2:
The patent changes the parameter space for command input by introducing motion parameters (acceleration, direction, duration) alongside or instead of sound parameters. This parameter change allows the system to execute diverse functions through physical motions, reducing the burden on the sound image library while maintaining adaptability
3Loss of information
If character images and outside scenes are displayed separately, then the display unit can show virtual information, but the user convenience deteriorates
Solution Approach 1:
The patent merges the display of virtual information with the outside scene by overlaying or integrating character images with the real-world view. Instead of displaying character images on a separate screen, the system combines them in a unified display field, allowing users to perceive both virtual information and real environment simultaneously, thereby improving convenience while preserving information
Data Source
AI summary
A transmission type head-mounted display device includes a sound acquiring unit configured to acquire sound on the outside, a sound identifying unit configured to identify specific sound in the acquired sound, an image display unit capable of displaying an image and capable of transmitting an outside scene, an image storing unit configured to store an execution function image representing a function executable by the head-mounted display device and a specific sound image associated with the specific sound, a display-image setting unit configured to cause the image display unit to display a combined image obtained by combining the execution function image and the specific sound image, and a function executing unit configured to execute a function corresponding to the execution function image combined with the specific sound image associated with the acquired specific sound.


