Context-Aware Voice Command Recognition Without a Wake Word
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing electronic devices with voice recognition functions require a call word for operation, leading to malfunctions when recognizing user utterances unrelated to device functions without a call word.
Innovation Solution
An electronic apparatus and method that identify a dedicated language model based on displayed content, recognize user utterances, and evaluate candidate texts for suitability using a geometric mean of entropy and threshold values to determine relevant operations without a call word.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If the electronic apparatus operates by recognizing user utterance without a call word, then the ease of operation is improved, but malfunctions occur due to performing operations for utterances unrelated to device functions
Solution Approach 1:
The patent introduces a call word as an intermediary element between the user and the voice recognition system. The call word serves as a mediator that activates the voice recognition mode, allowing the system to distinguish between casual speech and intended commands. This resolves the contradiction by maintaining ease of operation through voice recognition while preventing malfunctions through the call word activation mechanism.
Solution Approach 2:
The system performs preliminary action by requiring the user to utter a call word before the voice recognition function is activated. This preliminary step prepares the system for accurate operation by establishing a clear boundary between activation and execution phases, thereby preventing misrecognition of unrelated utterances while maintaining user-friendly voice-based operation.
2Reliability
If the electronic apparatus requires a call word for voice recognition, then the reliability is improved by preventing malfunctions, but the ease of operation deteriorates as users must utter the call word every time
Solution Approach 1:
The patent extracts the call word requirement from the continuous operation phase and applies it only to the activation phase. By separating the activation mechanism (requiring call word) from the execution phase (direct voice recognition), the system maintains reliability through call word verification while improving ease of operation during actual command execution, as users only need to utter the call word once per session rather than with every command.
3Ease of operation
If the electronic apparatus recognizes all user utterances without filtering, then the ease of operation is improved, but harmful factors increase due to unintended operations
Solution Approach 1:
The patent applies preliminary anti-action by using the call word as a preventive measure before voice recognition is activated. The call word serves as an anti-action that counteracts the potential harmful effect of recognizing unrelated utterances. By requiring this preliminary verification step, the system allows easy voice-based operation while preventing unintended operations that would constitute harmful factors.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Provided are an electronic apparatus and method for performing an operation based on recognizing a user's command utterance without a call word. The method includes identifying a dedicated language model related to a displayed content; receiving an utterance of a user; recognizing the received utterance and identifying candidate texts of the recognized utterance; identifying a similarity between the recognized utterance and the identified candidate texts; identifying, based on the identified dedicated language model and a predetermined threshold value, a suitability of a predetermined number of candidate texts with a high identified similarity, among the candidate texts; based on the identified suitability being outside a predetermined suitability range, ignoring the recognized utterance; and based on the identified suitability being in the predetermined suitability range, identifying a candidate text having a highest suitability, among the candidate texts, as the recognized utterance, and performing a corresponding operation.