Camera-Assisted Audio Processing for Far-Field Speech Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing far-field sound pickup technologies in intelligent devices suffer from noise interference, reverberation, and echo, leading to reduced speech recognition efficiency, especially in large open rooms with high acoustic reflection and diverse noise sources.
Innovation Solution
A signal processing method using a microphone array and a camera to determine the target sound source direction, combined with a voice quality enhancement model that integrates audio and video information to enhance speech recognition by suppressing noise and echo, and restoring clear audio signals based on lip shape semantics.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If beamforming technology and echo cancellation algorithm are used to suppress noise and echo, then audio signal clarity is improved, but speech recognition rate is greatly reduced due to severe reverberation and environmental noise
Solution Approach 1:
The patent introduces video information as a new dimension to complement audio processing. By capturing lip motion from video frames and aligning it with audio signals, the system creates a multi-modal input that provides additional spatial and temporal cues for speech recognition, overcoming the limitations of audio-only processing in reverberant environments
Solution Approach 2:
The patent merges audio signals with video-based lip motion information into a unified processing framework. The lip video frames are extracted, aligned with audio signals in time, and fed together into the speech recognition model, combining complementary information from both modalities to improve recognition accuracy despite noise and reverberation
2Speed
If microphone array is used for far-field sound pickup, then speech can be captured from distance, but definition of sound is greatly reduced due to noise and reverberation
Solution Approach 1:
The patent adds the video dimension to complement far-field audio capture. By capturing the speaker's lip motions visually and synchronizing them with the audio signal, the system recovers speech information that would otherwise be lost in reverberation and noise, maintaining sound definition despite far-field pickup challenges
Solution Approach 2:
The patent uses lip motion video as an intermediary to bridge the gap between far-field audio capture and clear speech recognition. The visual lip information serves as a mediator that provides clean speech cues independent of the degraded audio signal, enabling accurate recognition even when audio definition is reduced
3Reliability
If voice quality enhancement model with lip shape semantics is used, then speech recognition accuracy is improved, but system complexity increases due to integration of audio and video processing
Solution Approach 1:
The patent segments the processing pipeline into distinct modules: video frame extraction, lip region detection, lip video frame generation, audio-video alignment, and unified speech recognition. This segmentation allows each component to be optimized independently while maintaining overall system efficiency and managing complexity through modular design
Data Source
Figure 1
Figure 2~3
Figure 4
AI summary
A signal processing method and an electronic device (10) are disclosed. A first video (S220) is obtained by using a camera (140). A target sound source direction in which a target user performing speech interaction with the electronic device (10) is located is determined based on a first audio signal (S210) obtained by using a microphone array (130), so that estimation accuracy of the target sound source direction can be greatly improved. In addition, by using a preset voice quality enhancement model and a user lip video (S250) obtained in the target sound source direction by using the camera (140), voice quality enhancement is performed on a second audio signal (S260) obtained by using the microphone array (130). Because the voice quality enhancement model integrates a correspondence between a semantic meaning and a lip shape, a clean third audio signal can be restored based on the user lip video and the voice quality enhancement model, and finally, speech recognition efficiency can be effectively improved.