Lip-Reading Visual Command Confirmation for Noisy Voice Interfaces

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Voice-controlled systems in automotive contexts can be unreliable in noisy environments, leading to incorrect execution of user commands due to background noise or music, and lack of visual confirmation of intended actions.

Innovation Solution

An interactive system that generates visual images in response to spoken commands using lip detection and a text-to-image model, allowing users to confirm intended actions, particularly in noisy environments, by analyzing lip movements and generating informative images to assist the user.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If voice-controlled systems are used in noisy environments, then user interaction convenience is improved, but system reliability deteriorates due to background noise causing incorrect command detection

Engineering Contradiction:
Improveuser interaction convenienceVSAvoidsystem reliability
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent introduces an intermediary verification mechanism between voice input and action execution. A visual confirmation interface is displayed to the user, showing the interpreted command and intended action. The system waits for explicit user confirmation (e.g., through follow-up voice command or touch input) before executing the action. This intermediary step filters out false positives caused by background noise while maintaining the convenience of voice-controlled interaction.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If visual confirmation is added to voice-controlled systems, then system reliability is improved by reducing incorrect detections, but device complexity increases

Engineering Contradiction:
Improvedetection accuracyVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent leverages the existing display interface of the vehicle's infotainment system to provide visual confirmation, rather than adding a separate dedicated visual device. The same touchscreen display used for navigation and media control is repurposed to show command interpretation and seek confirmation. This multi-functional use of existing hardware adds reliability without significantly increasing device complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Reliability

If lip detection is used instead of audio processing, then reliability in noisy environments is improved, but loss of audio information occurs

Engineering Contradiction:
Improvecommand detection reliabilityVSAvoidaudio information loss
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent combines multiple detection modalities rather than choosing one over the other. Both audio processing and lip detection are performed simultaneously, and their results are fused to determine the user's intended command. The audio channel captures the full speech signal while the lip detection channel provides visual verification, especially useful in noisy environments. This merged approach maintains audio information while improving reliability through cross-validation.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS20250322829A1System and Method for Generating Visual Response to Spoken Input
Publication Date: 2025.10.16 APTIV TECHNOLOGIES AG
  • US20250322829A1 patent drawing
  • US20250322829A1 patent drawing
  • US20250322829A1 patent drawing

AI summary

A system for instructing execution of a function in response to a spoken input includes one or more cameras, a lip detection module, a text-to-image generation module, and an output module. The cameras are configured to capture images of a user. The lip detection module is configured to process the captured images to determine one or more words corresponding to lip movements of the user's spoken input. The text-to-image generation module is configured to generate one or more images representing a function responsive to the spoken input, based on the words determined by the lip detection module. The output module is configured to output the generated images for display. The cameras are arranged to capture the user's response to the displayed images. The system is arranged to instruct execution of the function in response to the captured response.