Visual Content Modification for Speech Recognition Ambiguity

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional Automatic Speech Recognition (ASR) systems face difficulties in recognizing spoken utterances with large vocabularies, unfamiliar accents or dialects, and in noisy environments, such as crowded airports or moving vehicles, due to ambiguity and lack of suitable training data.

Innovation Solution

The system modifies visual content on a display by adjusting the distance between potentially confusing visual elements based on their acoustic or topical similarity, using visual attention tracking to disambiguate user intent and customize the ASR system, thereby improving recognition accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Area of stationary object

If visual elements are placed close together to maximize display space utilization, then display efficiency is improved, but speech recognition accuracy deteriorates due to increased ambiguity between similarly pronounced words

Engineering Contradiction:
Improvedisplay space utilizationVSAvoidspeech recognition accuracy
Core Design Contradiction:
Area of stationary objectVSMeasurement precision

Solution Approach 1:

The patent applies local quality by dynamically adjusting the spatial distance between visual elements based on their acoustic similarity. Elements with similar pronunciations (high ambiguity) are placed farther apart, while elements with distinct pronunciations can be closer together. This localized adjustment optimizes both display space utilization and speech recognition accuracy by adapting the layout to the specific acoustic properties of the content being displayed.

Inventive Principle:
Principle #3Local quality

2Adaptability or versatility

If the ASR system uses a large vocabulary to handle diverse speech inputs, then system versatility is improved, but recognition accuracy deteriorates due to increased ambiguity and lack of suitable training data

Engineering Contradiction:
Improvevocabulary coverageVSAvoidrecognition accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent introduces visual elements as an intermediary between the user's speech and the ASR system. By displaying visual representations of possible speech inputs and adjusting their spatial arrangement based on acoustic similarity, the system provides contextual disambiguation that helps the ASR system accurately recognize speech even when using a large vocabulary. This visual intermediary resolves ambiguity without requiring the system to reduce its vocabulary size.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentEP3152754B1Modification of visual content to facilitate improved speech recognition
Publication Date: 2018.01.10 MICROSOFT TECHNOLOGY LICENSING LLC
  • EP3152754B1 patent drawingFigure 1
  • EP3152754B1 patent drawingFigure 2~3
  • EP3152754B1 patent drawingFigure 4

AI summary

Technologies described herein relate to modifying visual content for presentment on a display to facilitate improving performance of an automatic speech recognition (ASR) system. The visual content is modified to move elements further away from one another, wherein the moved elements give rise to ambiguity from the perspective of the ASR system. The visual content is modified to take into consideration accuracy of gaze tracking. When a user views an element in the modified visual content, the ASR system is customized as a function of the element being viewed by the user.