Visual Cue Processing for Voice Recognition Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Voice recognition programs and auto-complete applications on mobile devices often misinterpret user inputs due to inaccuracies in processing audible cues and incomplete text entries, despite attempts to improve accuracy using history and location data.

Innovation Solution

Processing visual cues, including images and video, to extract identification data and update a probable words dictionary, which enhances the understanding of user inputs by adding relevant words with priority values based on the context of the visual cues.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If voice recognition programs and auto-complete applications use history and location data to improve accuracy, then the understanding of user input improves, but the accuracy still deteriorates due to insufficient contextual information

Engineering Contradiction:
Improveaccuracy of user input understandingVSAvoidcontextual information
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent transitions from using only audio and text data to incorporating visual data as an additional dimension. The camera captures images of the target object, and image recognition extracts visual features that are integrated with voice and text inputs. This multi-dimensional approach provides richer contextual information that resolves ambiguities in user input that history and location data alone cannot address.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If visual cues are integrated into the probable words dictionary, then the accuracy of voice recognition and auto-complete improves, but the device complexity increases

Engineering Contradiction:
Improveaccuracy of voice recognitionVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system is divided into independent modular components: a camera module for capturing visual cues, an image recognition module for extracting visual features, a probable words dictionary module for storing and updating word lists, and a voice recognition module for processing audio inputs. Each module operates independently but contributes to the overall system function, making the complex system manageable and maintainable.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The probable words dictionary serves as an intermediary that integrates information from multiple sources (visual cues from image recognition, audio from voice recognition, and text from keyboard input). It acts as a central repository that combines data from different modules and provides unified output, simplifying the integration process and reducing direct complexity between components.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Loss of information

If multiple data sources (visual, audio, text) are processed simultaneously, then the contextual understanding improves, but the processing time increases

Engineering Contradiction:
Improvecontextual understandingVSAvoidprocessing time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The system performs preliminary processing of visual cues by capturing images and extracting visual features in advance, before voice recognition or text input occurs. The probable words dictionary is pre-populated with words associated with recognized objects in the visual field. This preliminary action reduces the processing burden during the actual input phase, as the system already has visual context ready to combine with audio and text inputs.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10146979B2Processing visual cues to improve device understanding of user input
Publication Date: 2018.12.04 LENOVO GLOBAL TECHNOLOGIES SWITZERLAND INTERNATIONAL GMBH
  • US10146979B2 patent drawing
  • US10146979B2 patent drawing
  • US10146979B2 patent drawing

AI summary

Processing visual cues to improve understanding of an input is described herein, including receiving a visual cue, the visual cue including visual media of a target; storing a list of words representing the target; and updating a probable words dictionary to include the list of words.