Depth-Camera Speech Recognition for Noisy In-Vehicle Voice Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In-vehicle speech recognition systems face challenges due to environmental noise changes, affecting accuracy and potentially causing navigation errors and increased driving risk.

Innovation Solution

A method that combines audio and video features by using a depth camera to capture facial depth images, extracting mouth shape features, and fusing them with voice features to improve recognition accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If voice interaction is used in in-vehicle environment, then convenience and safety are improved, but speech recognition accuracy deteriorates due to environmental noise

Engineering Contradiction:
ImproveconvenienceVSAvoidspeech recognition accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent combines audio features extracted from voice signals with visual features extracted from facial depth images to form fused features for speech recognition. This multi-modal fusion approach leverages both auditory and visual information to overcome the limitations of single-modality recognition in noisy in-vehicle environments, thereby maintaining high recognition accuracy while preserving the convenience of voice interaction.

Inventive Principle:
Principle #5Merging (Combining)

2Reliability

If voice interaction is used in in-vehicle environment, then safety is improved by freeing hands, but speech recognition accuracy deteriorates due to environmental noise

Engineering Contradiction:
ImprovesafetyVSAvoidspeech recognition accuracy
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The system fuses audio and visual modalities to enhance speech recognition reliability. By combining voice signals with facial depth image features, the system creates a more robust recognition mechanism that is less susceptible to environmental noise interference, thereby maintaining high safety standards while enabling hands-free operation.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces facial depth images as an intermediary visual modality to supplement voice recognition. The depth information from facial images serves as an additional channel that helps disambiguate voice commands in noisy environments, acting as a mediator that enhances the reliability of speech recognition without requiring physical contact or visual line-of-sight to the driver.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Device complexity

If audio features alone are used for speech recognition, then system complexity is reduced, but recognition accuracy deteriorates in noisy environments

Engineering Contradiction:
Improvesystem complexityVSAvoidrecognition accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent integrates audio feature extraction with visual feature extraction from facial depth images. The system processes both modalities simultaneously and fuses their features to achieve superior recognition accuracy in noisy in-vehicle environments, demonstrating that the increased complexity is justified by the significant improvement in recognition reliability.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentEP4191579B1Electronic device and speech recognition method therefor, and medium
Publication Date: 2026.01.28 HUAWEI TECH CO LTD
  • EP4191579B1 patent drawingFigure 1
  • EP4191579B1 patent drawingFigure 2
  • EP4191579B1 patent drawingFigure 3

AI summary

Embodiments of this application provide an electronic device, a speech recognition method therefor, and a medium, and relate to a speech recognition technology in the field of artificial intelligence (Artificial Intelligence, AI). The speech recognition method in this application includes: obtaining a facial depth image and a to-be-recognized voice of a user, where the facial depth image is an image collected by using a depth camera; recognizing a mouth shape feature from the facial depth image, and recognizing a voice feature from a to-be-recognized audio; and fusing the voice feature and the mouth shape feature into an audio-video feature, and recognizing, based on the audio-video feature, a voice uttered by the user. According to the method, because the mouth shape feature extracted from the facial depth image is not affected by light of an environment, the mouth shape feature can more accurately reflect a mouth shape change obtained when the user utters the voice. The mouth shape feature extracted from the facial depth image and the voice feature are fused, so that speech recognition accuracy can be improved.