Visual Cue Processing for Voice Recognition Accuracy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Voice recognition programs and auto-complete applications on mobile devices often misinterpret user inputs due to inaccuracies in processing audible cues and incomplete text entries, despite attempts to improve accuracy using history and location data.
Innovation Solution
Processing visual cues, including images and video, to extract identification data and update a probable words dictionary, which enhances the understanding of user inputs by adding relevant words with priority values based on the context of the visual cues.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If voice recognition programs and auto-complete applications use history and location data to improve accuracy, then the understanding of user input improves, but the accuracy still deteriorates due to insufficient contextual information
Solution Approach 1:
The patent transitions from using only audio and text data to incorporating visual data as an additional dimension. The camera captures images of the target object, and image recognition extracts visual features that are integrated with voice and text inputs. This multi-dimensional approach provides richer contextual information that resolves ambiguities in user input that history and location data alone cannot address.
2Measurement precision
If visual cues are integrated into the probable words dictionary, then the accuracy of voice recognition and auto-complete improves, but the device complexity increases
Solution Approach 1:
The system is divided into independent modular components: a camera module for capturing visual cues, an image recognition module for extracting visual features, a probable words dictionary module for storing and updating word lists, and a voice recognition module for processing audio inputs. Each module operates independently but contributes to the overall system function, making the complex system manageable and maintainable.
Solution Approach 2:
The probable words dictionary serves as an intermediary that integrates information from multiple sources (visual cues from image recognition, audio from voice recognition, and text from keyboard input). It acts as a central repository that combines data from different modules and provides unified output, simplifying the integration process and reducing direct complexity between components.
3Loss of information
If multiple data sources (visual, audio, text) are processed simultaneously, then the contextual understanding improves, but the processing time increases
Solution Approach 1:
The system performs preliminary processing of visual cues by capturing images and extracting visual features in advance, before voice recognition or text input occurs. The probable words dictionary is pre-populated with words associated with recognized objects in the visual field. This preliminary action reduces the processing burden during the actual input phase, as the system already has visual context ready to combine with audio and text inputs.
Data Source
AI summary
Processing visual cues to improve understanding of an input is described herein, including receiving a visual cue, the visual cue including visual media of a target; storing a list of words representing the target; and updating a probable words dictionary to include the list of words.


