Voice Control Dictionary Selection via User Gaze Direction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current voice-controlled device systems face challenges in accurately identifying the intended device for operation, especially when multiple devices are involved, leading to user inconvenience and reduced voice recognition precision due to the need for explicit device identification and the complexity of managing large dictionaries.
Innovation Solution
A control method that utilizes line-of-sight information to select the appropriate dictionary for voice command processing, allowing users to control devices without explicitly mentioning them, by detecting the user's gaze direction and switching between dictionaries accordingly.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a single large dictionary is used to cover multiple devices, then the system can handle various device operations, but the voice recognition precision decreases and the dictionary management becomes complex
Solution Approach 1:
The patent divides a single large dictionary into multiple device-specific dictionaries, each dedicated to a particular device type. This segmentation allows the system to select the appropriate dictionary based on the user's line-of-sight direction, thereby improving voice recognition precision while maintaining the ability to handle various device operations through multiple specialized dictionaries.
2Reliability
If device identification is made explicit through user input, then the intended device can be accurately identified, but the ease of operation decreases due to the additional user burden
Solution Approach 1:
The patent introduces line-of-sight direction detection as an intermediary mechanism between the user and the device identification process. The camera detects the user's gaze direction, and the controller uses this information to automatically select the appropriate device dictionary, thereby accurately identifying the intended device without requiring explicit user input and maintaining ease of operation.
3Measurement precision
If multiple dictionaries are prepared for different devices, then the voice recognition precision can be improved, but the device complexity increases due to dictionary management
Solution Approach 1:
The patent implements a dynamic dictionary selection mechanism where the controller automatically switches between different device dictionaries based on the real-time line-of-sight direction detected by the camera. This dynamic approach allows the system to use multiple specialized dictionaries for improved voice recognition precision while avoiding the complexity of manual dictionary management, as the selection process is automated based on user gaze direction.
4Reliability
If the user must explicitly mention the device name, then the control accuracy is improved, but the productivity decreases due to the longer speech required
Solution Approach 1:
The patent performs preliminary action by detecting the user's line-of-sight direction before processing the voice command. The camera captures the gaze direction, and the controller pre-selects the appropriate device dictionary based on this information. This preliminary action allows the system to accurately identify the intended device without requiring the user to explicitly mention the device name, thereby maintaining control accuracy while improving operation speed and productivity.
Data Source
AI summary
A first device is installed at a first location in a first space visible to a user. A second device is installed at a second location in a second space not visible to the user. A control method acquires line-of-sight information indicating a line-of-sight direction of the user from a camera. The line-of-sight direction of the user is determined based on the line-of-sight information. In a case where the line-of-sight direction indicates a third location other than the first location, a second dictionary corresponding to the second device is acquired from a plurality of dictionaries including a first dictionary corresponding to the first device and the second dictionary. Sound data indicating speech of the user is acquired from a microphone, a control command corresponding to the sound data is generated using the second dictionary, and the control command is transmitted to the second device.


