Camera-Assisted Audio Processing for Far-Field Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing far-field sound pickup technologies in intelligent devices suffer from noise interference, reverberation, and echo, leading to reduced speech recognition efficiency, especially in large open rooms with high acoustic reflection and diverse noise sources.

Innovation Solution

A signal processing method using a microphone array and a camera to determine the target sound source direction, combined with a voice quality enhancement model that integrates audio and video information to enhance speech recognition by suppressing noise and echo, and restoring clear audio signals based on lip shape semantics.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If beamforming technology and echo cancellation algorithm are used to suppress noise and echo, then audio signal clarity is improved, but speech recognition rate is greatly reduced due to severe reverberation and environmental noise

Engineering Contradiction:
Improveaudio signal clarityVSAvoidspeech recognition rate
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent introduces video information as a new dimension to complement audio processing. By capturing lip motion from video frames and aligning it with audio signals, the system creates a multi-modal input that provides additional spatial and temporal cues for speech recognition, overcoming the limitations of audio-only processing in reverberant environments

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent merges audio signals with video-based lip motion information into a unified processing framework. The lip video frames are extracted, aligned with audio signals in time, and fed together into the speech recognition model, combining complementary information from both modalities to improve recognition accuracy despite noise and reverberation

Inventive Principle:
Principle #5Merging (Combining)

2Speed

If microphone array is used for far-field sound pickup, then speech can be captured from distance, but definition of sound is greatly reduced due to noise and reverberation

Engineering Contradiction:
Improvefar-field sound pickup capabilityVSAvoidsound definition
Core Design Contradiction:
SpeedVSMeasurement precision

Solution Approach 1:

The patent adds the video dimension to complement far-field audio capture. By capturing the speaker's lip motions visually and synchronizing them with the audio signal, the system recovers speech information that would otherwise be lost in reverberation and noise, maintaining sound definition despite far-field pickup challenges

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent uses lip motion video as an intermediary to bridge the gap between far-field audio capture and clear speech recognition. The visual lip information serves as a mediator that provides clean speech cues independent of the degraded audio signal, enabling accurate recognition even when audio definition is reduced

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If voice quality enhancement model with lip shape semantics is used, then speech recognition accuracy is improved, but system complexity increases due to integration of audio and video processing

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the processing pipeline into distinct modules: video frame extraction, lip region detection, lip video frame generation, audio-video alignment, and unified speech recognition. This segmentation allows each component to be optimized independently while maintaining overall system efficiency and managing complexity through modular design

Inventive Principle:
Principle #1Segmentation

Data Source

PatentEP4207186B1Signal processing method and electronic device
Publication Date: 2025.07.30 HUAWEI TECH CO LTD
  • EP4207186B1 patent drawingFigure 1
  • EP4207186B1 patent drawingFigure 2~3
  • EP4207186B1 patent drawingFigure 4

AI summary

A signal processing method and an electronic device (10) are disclosed. A first video (S220) is obtained by using a camera (140). A target sound source direction in which a target user performing speech interaction with the electronic device (10) is located is determined based on a first audio signal (S210) obtained by using a microphone array (130), so that estimation accuracy of the target sound source direction can be greatly improved. In addition, by using a preset voice quality enhancement model and a user lip video (S250) obtained in the target sound source direction by using the camera (140), voice quality enhancement is performed on a second audio signal (S260) obtained by using the microphone array (130). Because the voice quality enhancement model integrates a correspondence between a semantic meaning and a lip shape, a clean third audio signal can be restored based on the user lip video and the voice quality enhancement model, and finally, speech recognition efficiency can be effectively improved.