Voice Command Generation via Visual Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Voice control of electronic devices often fails due to ambiguous or unrecognized voice commands, particularly when keywords are from different languages, not in the vocabulary, or pronounced unclearly, leading to unintended outcomes.

Innovation Solution

A method that combines voice input with visual selection from a screen to generate a command, using speech recognition for voice input and text recognition for selected content, allowing for more accurate command creation and avoiding improper recognition.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If speech recognition techniques are used for voice control, then the device can be controlled verbally, but recognition accuracy deteriorates when keywords are from different languages, not in vocabulary, or pronounced unclearly

Engineering Contradiction:
Improvevoice control capabilityVSAvoidrecognition accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent combines voice input with visual selection from a displayed list of recognized keywords. The system displays multiple possible interpretations of the voice command and allows the user to visually select the correct one, merging acoustic recognition with visual confirmation to improve overall accuracy.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system provides feedback by displaying the recognized keywords and their possible interpretations on the screen. This feedback loop allows users to review and correct recognition errors by selecting from the displayed options, improving the reliability of command recognition.

Inventive Principle:
Principle #23Feedback

2Adaptability or versatility

If the vocabulary for speech recognition is limited to the default language, then the recognition engine remains simple, but it cannot recognize terms from different languages

Engineering Contradiction:
Improvemulti-language recognitionVSAvoidvocabulary complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system achieves multi-language capability without requiring separate speech recognition engines for each language. By displaying recognized keywords and allowing visual selection, the system universally handles different languages, pronunciations, and vocabulary limitations through a single interface mechanism.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Productivity

If ambiguous terms are pronounced unclearly, then the command can be spoken faster, but the transcription becomes incorrect

Engineering Contradiction:
Improvecommand input speedVSAvoidtranscription accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system performs preliminary recognition and displays multiple possible interpretations before final command execution. This allows users to correct ambiguous transcriptions by selecting from the displayed options, ensuring accuracy before the command is carried out.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentEP3891730B1Technique for generating a command for a voice-controlled electronic device
Publication Date: 2023.07.05 VESTEL ELEKTRONIK SANAYI & TICARET ANONIM SIRKETI
  • EP3891730B1 patent drawingFigure 1
  • EP3891730B1 patent drawingFigure 2
  • EP3891730B1 patent drawingFigure 3

AI summary

A technique for generating a command to be processed by a voice-controlled electronic device is disclosed. A method implementation of the technique comprises receiving (S202) a voice input representative of a first portion of a command to be processed by the electronic device, receiving (S204) a selection of content displayed on a screen of the electronic device, the selected content being representative of a second portion of the command to be processed by the electronic device, and generating (S206) the command based on a combination of the voice input and the selected content.