Lip-Reading Visual Confirmation for In-Car Voice Commands

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Voice-controlled systems in automotive contexts can be unreliable in noisy environments, leading to incorrect execution of user commands due to background noise or music, and lack of intuitive confirmation mechanisms.

Innovation Solution

An interactive system that generates visual images based on lip detection and text-to-image processing, allowing users to confirm and verify the intended action through image output, reducing reliance on audio input and enhancing user understanding.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If voice-controlled systems are used in noisy environments, then hands-free operation is enabled, but reliability of command recognition deteriorates due to background noise

Engineering Contradiction:
Improvehands-free operationVSAvoidcommand recognition accuracy
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent introduces an intermediary visual confirmation mechanism between voice input and system execution. The system displays visual representations of detected commands and offers confirmation options to the user, acting as a mediator that verifies the accuracy of voice recognition before executing actions, thereby resolving the reliability issue in noisy environments while maintaining hands-free operation

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system implements feedback by visually displaying the interpreted command to the user and seeking confirmation. This feedback loop allows the user to verify that the system correctly understood the voice input despite background noise, and to correct misinterpretations before execution, thus improving reliability without sacrificing hands-free convenience

Inventive Principle:
Principle #23Feedback

2Reliability

If visual confirmation is added to voice-controlled systems, then command accuracy is improved, but system complexity increases

Engineering Contradiction:
Improvecommand execution accuracyVSAvoidsystem architecture complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system employs a multi-functional display device that serves both as a visual confirmation interface and as a confirmation input device. The same display shows the detected command and provides interactive confirmation options, eliminating the need for separate confirmation hardware and reducing overall system complexity while improving command execution accuracy

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Reliability

If lip detection is used instead of audio processing, then noise resistance is improved, but information loss occurs due to inability to capture audio semantics

Engineering Contradiction:
Improvenoise resistanceVSAvoidaudio semantic information
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent merges lip detection technology with audio processing in a hybrid system. The lip detection provides noise-resistant visual information about mouth movements, while the audio processing captures semantic content. By combining both inputs, the system achieves both noise resistance and accurate semantic understanding, preventing information loss while maintaining reliability in noisy environments

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentEP4632699A1System and method for generating visual response to spoken input
Publication Date: 2025.10.15 APTIV TECHNOLOGIES AG
  • EP4632699A1 patent drawingFigure 1
  • EP4632699A1 patent drawingFigure 2
  • EP4632699A1 patent drawingFigure 3

AI summary

A text-to-image generative model is used to render an informative image to a user according to a spoken input interpreted using lip reading analysis. In automotive contexts, functional information can be delivered to a driver visually, responsive to perceived semantic context of a spoken input. The driver can instruct or verify the information from the generated images that the system perceives and shows in order to execute a particular function.