TOF and HOA Array Fusion for Active Speaker Localization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio and video sensor platforms struggle with isolating and identifying sound sources in dynamic environments due to limitations in sound source discrimination and localization, particularly in the presence of external noise and occlusions, leading to inefficiencies in real-time operation and high processing complexity.
Innovation Solution
A system comprising a camera array and a microphone array, integrated with a computing device, uses DOA estimation, sound source identification, and sensor fusion to accurately locate and isolate sound sources by converting sound vectors and source positions into a common data representation, employing beamforming to record audio from active sources while discarding noise.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If multiple sensors tailored to different domains are used to analyze sound sources, then measurement precision and reliability are improved, but device complexity increases
Solution Approach 1:
The patent combines camera arrays and microphone arrays into an integrated multi-sensor system. The camera array captures visual data while the microphone array captures audio data, and both are processed together through sensor fusion to identify sound sources with high precision. This merging of different sensor types resolves the contradiction by achieving superior measurement precision through complementary data while managing complexity through unified system architecture.
Solution Approach 2:
The system employs sensors that serve multiple functions: the camera array not only captures visual information but also helps localize sound sources through visual tracking, while the microphone array provides both audio capture and spatial positioning data. This multi-functionality improves sound source identification accuracy without proportionally increasing device complexity, as each sensor contributes to multiple aspects of the analysis.
2Productivity
If real-time sound source localization is implemented using audio data only, then processing speed is improved, but measurement precision deteriorates due to inability to distinguish sources in dynamic environments
Solution Approach 1:
The system merges camera and microphone arrays to process both visual and audio data simultaneously in real-time. The camera array provides visual confirmation of sound sources while the microphone array captures audio signals, and their combined processing through sensor fusion achieves both real-time performance and high discrimination accuracy, resolving the contradiction between processing speed and measurement precision.
Solution Approach 2:
The sensor fusion module acts as an intermediary that integrates and correlates data from both camera and microphone arrays. It combines the temporal precision of audio processing with the spatial localization capabilities of visual processing, enabling real-time sound source identification with high precision by mediating between the two data streams and synthesizing their complementary strengths.
3Measurement precision
If sensor fusion techniques are used to correct occlusion issues, then measurement precision is improved, but processing time increases
Solution Approach 1:
The system performs preliminary tracking of sound sources using the camera array to establish visual positions before audio processing occurs. This preliminary visual localization allows the audio processing to focus on confirming and refining the position rather than searching from scratch, reducing processing time while maintaining high localization accuracy through the preliminary action of visual tracking.
Solution Approach 2:
The sensor fusion module implements feedback mechanisms where the camera's visual tracking data continuously updates and refines the sound source position estimates from the microphone array. This feedback loop allows the system to maintain high localization accuracy by constantly correcting positional estimates with visual information while processing efficiently, as the feedback guides rather than resets the processing.
Data Source
AI summary
A time-of-flight (TOF) array including a plurality of cameras is positioned with a coplanar Higher Order Ambisonics (HOA) array including a plurality of microphones above a target area such as a room or space for analyzing speaker activity in that area. The HOA microphone array iteratively samples sound energy level frames captured by the microphones to identify a sound vector associated with the global maximum energy level in the sound energy level frame. This sound vector is fused with the data from the TOF camera array that identified the positions of sound sources in the target area to associate produced sound corresponding with the sound vector with a physical active sound source in the target area. A beamformer can then be used to save audio corresponding to the sound vector and discard audio not associated with the sound vector to produce a sound recording associated with the active sound source.


