Voice Input Correction Using Visual Lip Movement Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Voice input interfaces often misinterpret voice inputs due to similarities in consonant sounds, leading to erroneous speech detection, and existing correction mechanisms require manual user intervention or rely on contextual data that may not always be available or accurate.
Innovation Solution
The system accepts voice input, identifies ambiguities through a processor, and uses non-audible inputs from sensors like cameras to capture lip movements and other visual cues to adjust the interpretation, providing a more intuitive and accurate correction mechanism.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If voice input interface is used for text input, then input speed and convenience are improved, but misinterpretation accuracy worsens due to similar consonant sounds
Solution Approach 1:
The patent combines audio input from a microphone with visual input from a camera to create a multi-modal input system. The audio receiver captures voice input while the camera captures lip movements, and both inputs are processed together to determine the user's intended speech, thereby maintaining high input speed while improving recognition accuracy.
Solution Approach 2:
The system introduces an intermediary visual channel (camera capturing lip movements) that mediates between the ambiguous audio input and the final text output. This intermediary provides additional information about the user's speech that helps disambiguate similar-sounding consonants without requiring manual correction.
2Measurement precision
If manual correction is implemented for misinterpreted text, then recognition accuracy is improved, but user operation complexity and time consumption worsen
Solution Approach 1:
The system performs self-correction by automatically using visual input from the camera to correct misinterpretations of audio input. The correction process happens automatically without requiring the user to manually select or type corrections, making the system self-sufficient in resolving recognition errors.
Solution Approach 2:
The system implements a feedback loop where the camera continuously captures lip movement data that is compared with the audio recognition results. When a mismatch is detected, the system uses the visual feedback to automatically adjust and correct the text input, providing real-time correction without user intervention.
3Measurement precision
If contextual data is used to correct voice input, then recognition accuracy is improved, but system complexity and data processing requirements worsen
Solution Approach 1:
The patent replaces complex contextual analysis and data processing mechanisms with a more direct sensory substitution approach. Instead of analyzing contextual data from multiple sources, the system substitutes the ambiguous audio signal with visual lip movement data, which provides direct information about the user's intended speech without requiring complex processing.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach allows for real-time, user-friendly correction of misinterpreted voice inputs by leveraging non-audible data to improve speech recognition accuracy, especially in cases where contextual data is limited or unreliable.
Implementation Method 1
accepting, at an audio receiver of an information handling device, voice input of a user
Implementation Method 2
a sensor that captures input... uses non-audible inputs from sensors like cameras to capture lip movements and other visual cues
Data Source
AI summary
An embodiment provides a method, including: accepting, at an audio receiver of an information handling device, voice input of a user; interpreting, using a processor, the voice input; identifying, using a processor, at least one ambiguity in interpreting the voice input; thereafter accessing stored non-audible input associated in time with the at least one ambiguity; and adjusting an interpretation of the voice input using non-audible input. Other aspects are described and claimed.


