Mobile Audio Beamforming with Inertial-Camera Sensor Fusion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Adaptive beamforming techniques for mobile microphone arrays face challenges in accurately targeting acoustic sources due to relative movement between the array and the object, especially when the source is intermittent or has low magnitude, leading to computational inefficiencies and lag.
Innovation Solution
Incorporating sensor fusion using inertial sensors and cameras to detect and measure relative positioning between the microphone array and the targeted object, enabling real-time adjustment of beamforming parameters.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If adaptive beamforming techniques are used to detect signal of interest in audio signals, then audio signal processing capability is improved, but computational complexity increases and causes processing lag
Solution Approach 1:
The system segments the signal processing task by separating adaptive beamforming (performed on audio signals) from movement detection (performed on sensor data). The sensor fusion module independently processes inertial sensor and camera data to detect relative movement, then feeds this information to the beamforming module. This segmentation allows each module to operate with optimized computational complexity while maintaining overall detection accuracy.
Solution Approach 2:
The patent introduces sensor fusion data (from inertial sensors and cameras) as an intermediary that mediates between the physical movement of the microphone array and the beamforming processing. This intermediary provides movement information without requiring the beamforming algorithm to directly analyze audio signals for movement detection, thereby reducing computational complexity while preserving detection precision.
2Measurement precision
If adaptive beamforming relies on detecting signal of interest in audio signals, then beamforming accuracy is improved, but processing speed decreases due to computational lag
Solution Approach 1:
The system performs preliminary action by detecting relative movement between the microphone array and targeted objects using inertial sensors and cameras before the beamforming process begins. This pre-detection of movement allows the beamforming module to receive already-processed movement information, eliminating the need to perform computationally intensive movement analysis during the time-critical beamforming stage, thus improving processing speed while maintaining accuracy.
3Adaptability or versatility
If the microphone array moves relative to the targeted object, then mobility is improved, but beamforming accuracy deteriorates due to relative movement
Solution Approach 1:
The system implements feedback by continuously monitoring relative movement between the microphone array and targeted objects using sensor fusion (inertial sensors and cameras). This movement information is fed back to the beamforming module, which then adjusts its parameters in real-time to compensate for the relative movement. This closed-loop feedback mechanism enables the system to maintain high beamforming accuracy even during mobile operation.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Audio receive beamforming is performed by a computing system. A set of audio signals are obtained via a microphone array (230, 310) and a set of inertial signals are obtained via a set of inertial sensors (232, 312) of a mobile device (210). A location of a targeted object (220) to beamform is identified within a camera feed captured via a set of one or more cameras (234, 314) imagining an environment of the mobile device. A parameter of a beamforming function is determined (324) that defines a beamforming region containing the targeted object (220) based on the set of inertial signals and the location of the targeted object. The beamforming function is applied (330) to the set of audio signals using the parameter to obtain a set of processed audio signals that increases a signal-to-noise ratio of an audio source within the beamforming region relative to the set of audio signals.