Lip-Reading Visual Command Confirmation for Noisy Voice Interfaces
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Voice-controlled systems in automotive contexts can be unreliable in noisy environments, leading to incorrect execution of user commands due to background noise or music, and lack of visual confirmation of intended actions.
Innovation Solution
An interactive system that generates visual images in response to spoken commands using lip detection and a text-to-image model, allowing users to confirm intended actions, particularly in noisy environments, by analyzing lip movements and generating informative images to assist the user.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If voice-controlled systems are used in noisy environments, then user interaction convenience is improved, but system reliability deteriorates due to background noise causing incorrect command detection
Solution Approach 1:
The patent introduces an intermediary verification mechanism between voice input and action execution. A visual confirmation interface is displayed to the user, showing the interpreted command and intended action. The system waits for explicit user confirmation (e.g., through follow-up voice command or touch input) before executing the action. This intermediary step filters out false positives caused by background noise while maintaining the convenience of voice-controlled interaction.
2Reliability
If visual confirmation is added to voice-controlled systems, then system reliability is improved by reducing incorrect detections, but device complexity increases
Solution Approach 1:
The patent leverages the existing display interface of the vehicle's infotainment system to provide visual confirmation, rather than adding a separate dedicated visual device. The same touchscreen display used for navigation and media control is repurposed to show command interpretation and seek confirmation. This multi-functional use of existing hardware adds reliability without significantly increasing device complexity.
3Reliability
If lip detection is used instead of audio processing, then reliability in noisy environments is improved, but loss of audio information occurs
Solution Approach 1:
The patent combines multiple detection modalities rather than choosing one over the other. Both audio processing and lip detection are performed simultaneously, and their results are fused to determine the user's intended command. The audio channel captures the full speech signal while the lip detection channel provides visual verification, especially useful in noisy environments. This merged approach maintains audio information while improving reliability through cross-validation.
Data Source
AI summary
A system for instructing execution of a function in response to a spoken input includes one or more cameras, a lip detection module, a text-to-image generation module, and an output module. The cameras are configured to capture images of a user. The lip detection module is configured to process the captured images to determine one or more words corresponding to lip movements of the user's spoken input. The text-to-image generation module is configured to generate one or more images representing a function responsive to the spoken input, based on the words determined by the lip detection module. The output module is configured to output the generated images for display. The cameras are arranged to capture the user's response to the displayed images. The system is arranged to instruct execution of the function in response to the captured response.


