Acoustic Beamforming for Future Visual Frames in Autonomous Vehicles
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing object detection and tracking technologies in autonomous vehicles are limited by adverse weather conditions and low reflectance, as electromagnetic radiation-based sensors struggle in fog and low-light scenarios, leading to reduced performance and inability to detect objects outside their direct line-of-sight.
Innovation Solution
Employing a microphone array for long-range acoustic beamforming and synthetic aperture expansion to generate spatial beamforming maps, which are combined with visual signals for enhanced object detection and future visual frame prediction, leveraging passive sound from traffic participants.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If electromagnetic radiation-based sensors (camera, radar, LiDAR) are used for object detection, then detection capability in normal conditions is achieved, but performance deteriorates in adverse weather conditions such as fog and low-light scenarios
Solution Approach 1:
The patent combines electromagnetic radiation-based sensors (camera, radar, LiDAR) with acoustic sensors (microphone array) to create a multi-modal sensing system. This merging allows the system to leverage both visual data and acoustic data, maintaining reliable object detection in adverse weather conditions where either modality alone would fail.
Solution Approach 2:
The patent creates a composite sensing system that integrates multiple sensing modalities (electromagnetic and acoustic) into a unified detection framework. This composite approach enables the system to overcome the limitations of individual sensor types by combining their complementary strengths.
2Measurement precision
If acoustic beamforming is applied to improve sound source localization, then spatial resolution is enhanced, but computational complexity and processing requirements increase
Solution Approach 1:
The patent segments the acoustic signal processing into distinct stages: beamforming to generate spatial maps, synthetic aperture expansion to enhance resolution, and temporal information extraction for prediction. This segmentation allows each processing stage to be optimized independently, managing computational complexity while maintaining precision.
Solution Approach 2:
The patent applies synthetic aperture expansion to transition from the physical aperture dimension to an expanded synthetic aperture dimension, effectively increasing spatial resolution without proportionally increasing the physical microphone array size or computational burden.
3Measurement precision
If future visual frame prediction is generated using only visual signals, then processing speed is maintained, but accuracy deteriorates in non-line-of-sight and occluded scenarios
Solution Approach 1:
The patent performs preliminary beamforming to generate spatial acoustic maps and extract temporal information in advance. This preliminary processing of acoustic data, combined with visual signals, enables more accurate future frame prediction without significantly increasing real-time processing requirements, as the acoustic temporal patterns are pre-analyzed.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
Improves object detection and tracking in challenging conditions by providing high-resolution spatial information and future frame prediction, complementing electromagnetic radiation-based sensors with acoustic data, especially in non-line-of-sight and partially occluded scenarios.
Implementation Method 1
acoustic signals received at the plurality of microphones of the microphone array
Implementation Method 2
generate spatial beamforming maps locating a sound source using a beamforming model corresponding to acoustic signals
Implementation Method 3
apply a synthetic aperture expansion to the acoustic signals to increase resolution of the spatial beamforming maps
Data Source
AI summary
An autonomous vehicle including a microphone array of a plurality of microphones, a visual sensor network configured to receive visual signals, at least one processor, and at least one memory storing instructions is disclosed. The instructions, when executed by the at least one processor, cause the at least one processor to: (i) generate spatial beamforming maps locating a sound source using a beamforming model corresponding to acoustic signals received at the plurality of microphones of the microphone array; (ii) apply a synthetic aperture expansion to the acoustic signals to increase resolution of the spatial beamforming maps; and (iii) generate a future visual frame based at least partially upon temporal information extracted from the spatial beamforming maps and visualization maps generated based on the visual signals received by the visual sensor network.


