Spherical Microphone Array GPU Beamforming
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies face challenges in achieving frame-rate performance for sound image creation and beamforming with spherical microphone arrays, and in automatically identifying sound sources in conjunction with video, requiring advanced calibration and processing methods to align audio and video data effectively.
Innovation Solution
A spherical-camera array system is calibrated to perform vision-guided beamforming, using graphics processors to process audio and video data in real-time, allowing for the creation of sound images and the transfer of audio intensity information onto video images, enabling precise source localization and noise identification.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If beamforming is performed using weighted sum of Fourier coefficients of all microphone signals, then sound image quality is improved, but processing speed deteriorates and frame-rate performance cannot be achieved
Solution Approach 1:
The patent divides the complex beamforming computation into multiple processing stages: FFT computation for each microphone channel, spectral weighting application, and spatial summation. This segmentation allows parallel processing of individual microphone signals before combining them, significantly improving processing speed while maintaining sound image quality.
Solution Approach 2:
The patent applies spectral weighting selectively to enhance specific frequency ranges that are most relevant for sound source identification. By focusing computational resources on partial frequency bands rather than processing all frequencies equally, the system achieves frame-rate performance while maintaining adequate sound image quality for practical applications.
2Adaptability or versatility
If spherical microphone arrays are used to capture sound field representation, then sound source identification capability is improved, but system complexity increases
Solution Approach 1:
The patent employs spherical microphone arrays that serve multiple functions: capturing the complete sound field representation for 3D sound source localization, providing directional information through beamforming, and enabling both qualitative identification and quantitative analysis. This multi-functionality justifies the increased device complexity by delivering comprehensive acoustic analysis capabilities.
Solution Approach 2:
The patent introduces computational processing intermediaries including FFT algorithms, spectral weighting functions, and beamforming calculations that transform the raw multi-channel microphone data into meaningful sound images. These computational intermediaries bridge the gap between the complex physical array and the simplified visual output, making the system manageable despite its physical complexity.
3Measurement precision
If sound images are captured in conjunction with video for automatic analysis, then source localization precision is improved, but processing complexity increases
Solution Approach 1:
The patent merges audio and video processing pipelines by synchronizing frame rates and integrating the sound image overlay with video frames. This combination allows automatic source localization by correlating acoustic data with visual information, improving precision while managing processing complexity through unified system architecture.
Solution Approach 2:
The patent creates a simplified 2D representation of the 3D sound field by projecting acoustic intensity data onto the video image plane. This copying approach allows automatic analysis by translating complex spatial acoustic information into a format that can be directly overlaid and analyzed with video data, reducing processing complexity while maintaining localization precision.
Data Source
AI summary
Spherical microphone arrays provide an ability to compute the acoustical intensity corresponding to different spatial directions in a given frame of audio data. These intensities may be exhibited as an image and these images are generated at a high frame rate to achieve a video image if the data capture and intensity computations can be performed sufficiently quickly, thereby creating a frame-rate audio camera. A description is provided herein regarding how such a camera is built and the processing done sufficiently quickly using graphics processors. The joint processing of and captured frame-rate audio and video images enables applications such as visual identification of noise sources, beamforming and noise-suppression in video conferencing and others, by accounting for the spatial differences in the location of the audio and the video cameras. Based on the recognition that the spherical array can be viewed as a central projection camera, such joint analysis can be performed.


