Audio-Visual Source Tracking for Mobile Voice Capture
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional audio-only source tracking systems face challenges in noisy environments, particularly in mobile communication devices, where device movements and changing sound source positions lead to processing delays and poor signal-to-noise ratios, making it difficult to separate the user's voice from background noise, especially during silent periods and in environments with competing talkers.
Innovation Solution
An audio-visual source tracking system that uses a combination of video cameras and microphone arrays to detect and track the user's face, adjusting the microphone array's sensitivity and beam direction in real-time to enhance audio capture quality by determining the mouth reference point and orienting the microphone array towards the user's mouth, thereby reducing background noise and improving voice and video call quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If audio-only source tracking techniques are used to improve signal-to-noise ratio, then unwanted signals are eliminated, but processing delays occur and tracking fails during silent periods
Solution Approach 1:
The patent combines audio signal processing with visual information from cameras to create an audio-visual source tracking system. The visual data provides continuous tracking information during silent periods when audio-only methods fail, while audio data refines the source location when available. This merging of sensory modalities resolves the contradiction by maintaining reliable tracking without processing delays.
Solution Approach 2:
The system uses visual information as an intermediary to bridge the gaps in audio-based tracking. During silent periods or when audio processing delays occur, the visual tracking data serves as a mediator to maintain continuous source location information, ensuring neither reliability nor timing is compromised.
2Ease of operation
If device movements are accommodated to maintain user contact, then communication continuity is preserved, but source tracking becomes challenging due to fast DOA changes
Solution Approach 1:
The system dynamically adapts to device movements by continuously updating the relationship between camera and microphone array positions. The audio-visual fusion algorithm dynamically adjusts to fast changes in direction of arrival caused by device movement, maintaining accurate source tracking despite the complexity introduced by continuous motion compensation.
Solution Approach 2:
The integrated audio-visual system serves multiple functions simultaneously: it tracks the user for communication continuity, compensates for device movements, and maintains accurate source localization. This multi-functionality resolves the contradiction by handling device mobility complexities through a unified system that performs both communication and tracking tasks.
3Reliability
If microphone array directivity is adjusted to attenuate background noise, then signal-to-noise ratio improves, but voice capture quality may be compromised during silent periods
Solution Approach 1:
The system performs preliminary tracking using visual information from cameras to establish the user's position and orientation before audio processing begins. This preliminary visual data allows the microphone array to be pre-oriented toward the user, ensuring that when voice signals are present, the system is already optimized for capture while maintaining noise attenuation. This prevents voice information loss during the transition into speech.
Data Source
Figure 1~2
Figure 3
Figure 4
AI summary
Disclosed herein is an apparatus. The apparatus includes a housing, electronic circuitry, and an audio-visual source tracking system. The electronic circuitry is in the housing. The audio-visual source tracking system includes a first video camera and an array of microphones. The first video camera and the array of microphones are attached to the housing. The audio- visual source tracking system is configured to receive video information from the first video camera. The audio-visual source tracking system is configured to capture audio information from the array of microphones at least partially in response to the video information. The audio-visual source tracking system might include a second video camera that is attached to the housing, wherein the first and second video cameras together estimate the beam orientation of the array of microphones.