Visual Content Modification for Speech Recognition Ambiguity
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional Automatic Speech Recognition (ASR) systems face difficulties in recognizing spoken utterances with large vocabularies, unfamiliar accents or dialects, and in noisy environments, such as crowded airports or moving vehicles, due to ambiguity and lack of suitable training data.
Innovation Solution
The system modifies visual content on a display by adjusting the distance between potentially confusing visual elements based on their acoustic or topical similarity, using visual attention tracking to disambiguate user intent and customize the ASR system, thereby improving recognition accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Area of stationary object
If visual elements are placed close together to maximize display space utilization, then display efficiency is improved, but speech recognition accuracy deteriorates due to increased ambiguity between similarly pronounced words
Solution Approach 1:
The patent applies local quality by dynamically adjusting the spatial distance between visual elements based on their acoustic similarity. Elements with similar pronunciations (high ambiguity) are placed farther apart, while elements with distinct pronunciations can be closer together. This localized adjustment optimizes both display space utilization and speech recognition accuracy by adapting the layout to the specific acoustic properties of the content being displayed.
2Adaptability or versatility
If the ASR system uses a large vocabulary to handle diverse speech inputs, then system versatility is improved, but recognition accuracy deteriorates due to increased ambiguity and lack of suitable training data
Solution Approach 1:
The patent introduces visual elements as an intermediary between the user's speech and the ASR system. By displaying visual representations of possible speech inputs and adjusting their spatial arrangement based on acoustic similarity, the system provides contextual disambiguation that helps the ASR system accurately recognize speech even when using a large vocabulary. This visual intermediary resolves ambiguity without requiring the system to reduce its vocabulary size.
Data Source
Figure 1
Figure 2~3
Figure 4
AI summary
Technologies described herein relate to modifying visual content for presentment on a display to facilitate improving performance of an automatic speech recognition (ASR) system. The visual content is modified to move elements further away from one another, wherein the moved elements give rise to ambiguity from the perspective of the ASR system. The visual content is modified to take into consideration accuracy of gaze tracking. When a user views an element in the modified visual content, the ASR system is customized as a function of the element being viewed by the user.