Acoustic Depth Map Using Echo Spectrograms
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing depth perception methods using acoustic signals face challenges in generating accurate depth maps, particularly in environments with limited visibility and directional representation, as they rely on narrow field of view data and struggle with omnidirectional acoustic information.
Innovation Solution
A depth sensing apparatus and method that employs an audio output device to emit omnidirectional audio signals, captures echo signals with multiple audio sensors, generates spectrograms, and applies them to a computational model trained with reference echo signals and omnidirectional reference depth images to produce accurate depth maps, utilizing techniques like pre-processing, augmentation, and recurrent modules for improved accuracy and robustness.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional vision and lidar based sensing systems are used, then high fidelity information is provided for navigation, but they become unreliable in low light, smoke, dust or fog conditions
Solution Approach 1:
The patent replaces optical-based sensing systems (vision and lidar) with an acoustic-based echolocation system. The audio output device emits acoustic signals that propagate through the environment and reflect off objects, with the audio sensors capturing these echoes. This acoustic substitution enables reliable operation in adverse conditions (smoke, dust, fog, low light) where optical systems fail, while maintaining depth perception capability through computational processing of acoustic echo data
2Ease of manufacture
If narrow field of view ground truth data is used for training, then the computational model can be trained with available data, but the depth estimation is limited to the narrow field of view and creates a mismatch with omnidirectional acoustic information
Solution Approach 1:
The patent transforms the training data representation from narrow field of view 2D images to omnidirectional 360-degree depth maps. The audio output device emits omnidirectional acoustic signals and the system captures echoes from all directions around the device. The computational model is trained to process this omnidirectional acoustic data and generate complete 360-degree depth maps, enabling full environmental awareness rather than limited field of view coverage
Solution Approach 2:
The patent creates a universal training framework where the computational model learns to process omnidirectional acoustic information from multiple audio sensors and generate comprehensive depth maps of the entire environment. The system uses multiple audio sensors positioned to capture echoes from all directions, and the trained model can universally handle omnidirectional acoustic inputs to produce complete spatial awareness, making the system adaptable to any directional scenario
3Measurement precision
If multiple audio sensors are used to capture omnidirectional echo signals, then complete environmental coverage is achieved, but the device complexity increases
Solution Approach 1:
The patent merges the functionality of multiple audio sensors into a unified omnidirectional acoustic processing system. The multiple audio sensors are positioned around the audio output device to capture echoes from all directions simultaneously. The computational model integrates the signals from all sensors and processes them together to generate a single comprehensive omnidirectional depth map, effectively combining multiple sensing functions into one unified output that provides complete environmental awareness
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
The solution enables the generation of accurate omnidirectional depth maps, improving accuracy and robustness in various environments, even in conditions where traditional sensing modalities like lidar or vision fail, and allows for multi-modal sensing with enhanced performance and reduced computational requirements.
Implementation Method 1
acquire echo signals indicative of reflected audio signals captured by the at least one audio sensors in response to reflection of the emitted audio signal from the environment surrounding the depth sensing apparatus
Data Source
AI summary
A depth sensing apparatus configured to generate a depth map of an environment, the apparatus including an audio output device, at least one audio sensor and one or more processing devices configured to cause the audio output device to emit an omnidirectional emitted audio signal, acquire echo signals indicative of reflected audio signals captured by the at least one audio sensors in response to reflection of the emitted audio signal from the environment surrounding the depth sensing apparatus, generate spectrograms using the echo signals and apply the spectrograms to a computational model to generate a depth map, the computational model being trained using reference echo signals and omnidirectional reference depth images.


