Virtual Microphone Audio Synthesis via Visual Object Tracking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional audio tracking techniques are inaccurate due to echoes, noise, and directional microphone characteristics, making it difficult to track audio sources, especially when they are silent or move outside the microphone's sensitive region, leading to compromised location determination.
Innovation Solution
A system and method that processes video and audio signals to identify and match visual objects with audio sources, using geometric modeling and computer vision to track audio sources indirectly by tracking visual objects, and synthesizes audio signals for a virtual microphone position, simulating the sound recorded by an actual microphone at that position.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional audio tracking techniques are used, then the system can track audio sources, but the accuracy is compromised due to echoes, noise, and directional microphone characteristics
Solution Approach 1:
The patent introduces visual objects as an intermediary to track audio sources. Instead of directly tracking audio sources using microphones, the system tracks visual objects and uses geometric modeling to infer audio source positions. This intermediary approach eliminates the harmful effects of echoes, noise, and directional microphone characteristics on tracking accuracy.
Solution Approach 2:
The patent replaces the acoustic field-based tracking mechanism with a visual field-based tracking mechanism. By substituting acoustic signals with visual signals for tracking purposes, the system avoids the limitations of acoustic field measurements including echoes and noise interference.
2Adaptability or versatility
If directional microphones are used, then the system can capture audio directionally, but audio sources moving outside the sensitive region or silent sources cannot be tracked accurately
Solution Approach 1:
The patent makes the tracking system universal by using visual objects that can detect all audio sources regardless of their position, direction, or whether they are currently emitting sound. The visual tracking system does not have directional sensitivity limitations like acoustic systems, enabling consistent tracking across all scenarios.
Solution Approach 2:
The system performs preliminary tracking of visual objects to establish their positions before audio tracking is needed. This preliminary visual tracking ensures that even when audio sources are silent or move outside the microphone's sensitive region, the system already has the spatial information needed for accurate location determination.
3Ease of manufacture
If directional microphone characteristics are used, then the system can provide directional audio capture, but coloration and delays make it difficult to precisely determine audio source location
Solution Approach 1:
The patent extracts the tracking function from the acoustic field and places it in the visual field. By separating tracking from acoustic measurement, the system eliminates the coloration and delays introduced by directional microphone characteristics while preserving the ability to provide directional audio capture through the visual tracking and geometric modeling approach.
Data Source
AI summary
Disclosed is a system and method for generating a model of the geometric relationships between various audio sources recorded by a multi-camera system. The spatial audio scene module associates source signals, extracted from recorded audio, of audio sources to visual objects identified in videos recorded by one or more cameras. This association may be based on estimated positions of the audio sources based on relative signal gains and delays of the source signal received at each microphone. The estimated positions of audio sources are tracked indirectly by tracking the associated visual objects with computer vision. A virtual microphone module may receive a position for a virtual microphone and synthesize a signal corresponding to the virtual microphone position based on the estimated positions of the audio sources.


