Voice Command Generation via Visual Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Voice control of electronic devices often fails due to ambiguous or unrecognized voice commands, particularly when keywords are from different languages, not in the vocabulary, or pronounced unclearly, leading to unintended outcomes.
Innovation Solution
A method that combines voice input with visual selection from a screen to generate a command, using speech recognition for voice input and text recognition for selected content, allowing for more accurate command creation and avoiding improper recognition.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If speech recognition techniques are used for voice control, then the device can be controlled verbally, but recognition accuracy deteriorates when keywords are from different languages, not in vocabulary, or pronounced unclearly
Solution Approach 1:
The patent combines voice input with visual selection from a displayed list of recognized keywords. The system displays multiple possible interpretations of the voice command and allows the user to visually select the correct one, merging acoustic recognition with visual confirmation to improve overall accuracy.
Solution Approach 2:
The system provides feedback by displaying the recognized keywords and their possible interpretations on the screen. This feedback loop allows users to review and correct recognition errors by selecting from the displayed options, improving the reliability of command recognition.
2Adaptability or versatility
If the vocabulary for speech recognition is limited to the default language, then the recognition engine remains simple, but it cannot recognize terms from different languages
Solution Approach 1:
The system achieves multi-language capability without requiring separate speech recognition engines for each language. By displaying recognized keywords and allowing visual selection, the system universally handles different languages, pronunciations, and vocabulary limitations through a single interface mechanism.
3Productivity
If ambiguous terms are pronounced unclearly, then the command can be spoken faster, but the transcription becomes incorrect
Solution Approach 1:
The system performs preliminary recognition and displays multiple possible interpretations before final command execution. This allows users to correct ambiguous transcriptions by selecting from the displayed options, ensuring accuracy before the command is carried out.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A technique for generating a command to be processed by a voice-controlled electronic device is disclosed. A method implementation of the technique comprises receiving (S202) a voice input representative of a first portion of a command to be processed by the electronic device, receiving (S204) a selection of content displayed on a screen of the electronic device, the selected content being representative of a second portion of the command to be processed by the electronic device, and generating (S206) the command based on a combination of the voice input and the selected content.