Display Device Voice Recognition Context Awareness
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition services in display devices do not consider the usage environment of the device, leading to limitations in accurately interpreting user voice commands.
Innovation Solution
A display device equipped with a network interface and control unit that receives voice commands, acquires usage information, and transmits this information to a server to determine the user's intention, allowing for dynamic operation based on the voice command and usage context.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If the speech recognition service depends only on the user's utterance, then the system complexity is reduced, but the accuracy of intention recognition deteriorates
Solution Approach 1:
The system performs preliminary actions by acquiring usage information (channel information, volume information, input device information) before processing the voice command. This preliminary data collection enables the server to accurately interpret user intention by combining it with the current usage context, resolving the contradiction between system simplicity and recognition accuracy.
Solution Approach 2:
The patent introduces usage information as an intermediary element that mediates between the user's voice command and the system's intention recognition. This intermediary provides contextual background that helps the server accurately understand user intent without requiring the client device to be overly complex.
2Measurement precision
If usage information is transmitted to the server for each voice command, then the accuracy of speech recognition is improved, but the network communication load increases
Solution Approach 1:
The patent extracts only the necessary usage information (channel, volume, input device) that is relevant for speech recognition accuracy, rather than transmitting all possible device data. This selective extraction reduces network communication load while maintaining recognition accuracy.
3Measurement precision
If a separate dictionary is built in the server for all real-time information, then the completeness of recognition is improved, but the server load increases
Solution Approach 1:
The patent extracts and transmits only the specific usage information needed for context-aware recognition (channel, volume, input device) rather than maintaining a comprehensive dictionary of all possible real-time information in the server. This approach achieves recognition completeness for relevant contexts while minimizing server load.
Solution Approach 2:
The client device self-services by collecting and transmitting its own usage information to the server, eliminating the need for the server to maintain and update a separate dictionary for all real-time device states. This shifts the information gathering responsibility to the client, reducing server burden.
Data Source
AI summary
A display device for providing a speech recognition service according to an embodiment of the present disclosure can include a display unit, a network interface unit configured to communicate with a server, and a control unit configured to receive a voice command uttered by a user, acquire usage information of the display device, transmit the voice command and the usage information of the display device to the server through the network interface unit, receive, from the server, an utterance intention based on the voice command and the usage information of the display device, and perform an operation corresponding to the received utterance intention.


