Spatial Audio Processing Using Frequency-Domain Whitening
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing microphone array systems struggle to effectively separate target sound sources from multiple sound sources in complex acoustic environments, particularly when sound sources are in the same direction relative to the array, leading to limitations in sound acquisition and source localization.
Innovation Solution
A spatial audio processing system that uses an adaptive whitening filter and an environmental and physical model of the acoustic space as a waveguide to process audio inputs. This system estimates model parameters for the acoustic space, allowing for the detection and location of new sound sources and providing significant separation of target sources even with fewer microphones than noise sources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional beamforming techniques are used to separate target sound sources from multiple sound sources, then sound acquisition and source localization can be achieved, but the signal-to-noise ratio is limited and the system requires many microphones to handle complex acoustic environments
Solution Approach 1:
The patent transforms the audio signal from time domain to frequency domain using Fourier transform, and applies frequency-dependent filtering parameters to separate target sounds from noise. This parameter transformation enables effective sound separation with fewer microphones by exploiting spectral characteristics rather than relying solely on spatial arrangement.
Solution Approach 2:
The patent replaces the mechanical/spatial approach of conventional beamforming (which relies on microphone array geometry and time-delay processing) with a spectral processing approach using Fourier transform and frequency-domain filtering. This substitution allows achieving better sound separation performance with fewer microphones by operating in the frequency domain where noise and target signals have distinct spectral signatures.
2Reliability
If more microphones are used to improve sound source separation in complex acoustic environments, then the signal-to-noise ratio improves, but the device complexity and cost increase
Solution Approach 1:
The patent replaces complex mechanical microphone array configurations with a simplified spectral processing system. By using Fourier transform to convert time-domain signals to frequency-domain representations, the system can apply frequency-selective filtering to enhance target sounds and suppress noise without requiring complex spatial arrangements of many microphones.
Solution Approach 2:
The patent changes the processing domain from time domain to frequency domain, enabling the use of frequency-dependent parameters for noise reduction. This parameter transformation allows the system to achieve reliable sound acquisition quality by exploiting spectral differences between target signals and noise, rather than relying on increased microphone quantity or complex array geometry.
3Object-affected harmful factors
If conventional beamforming is used to reduce reverberation and improve sound quality, then directional audio pickup is achieved, but the system cannot effectively separate sources in the same direction and requires knowledge of array configuration
Solution Approach 1:
The patent replaces conventional beamforming's spatial filtering approach with frequency-domain spectral processing. By applying Fourier transform and frequency-selective filtering, the system can separate target sounds from reverberation and noise based on their spectral characteristics rather than spatial direction. This enables effective source separation even when sources are in the same direction, without requiring knowledge of microphone array configuration or orientation.
Solution Approach 2:
The patent transforms the audio signal to the frequency domain and applies frequency-dependent filtering parameters to differentiate target sounds from reverberation. This parameter change from spatial-domain to frequency-domain processing enables the system to achieve adaptability in separating co-directional sources by exploiting spectral differences, without being constrained by array geometry or requiring directional information.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
The system achieves a 15 dB or more improvement in signal-to-noise ratio (SNR) compared to prior art beamforming techniques, while using fewer transducers and without requiring knowledge of the array configuration or orientation. It effectively reduces non-desired sounds and enhances target sounds in complex acoustic environments.
Implementation Method 1
converting the audio input from a time domain to a frequency domain according to at least one transform function, the at least one transform function being selected from the group consisting of: a Fourier transform, a Fast Fourier transform, a Short Time Fourier transform and a modulated complex lapped transform
Implementation Method 2
determining at least one acoustic propagation model for at least one source location within the acoustic environment according to a normalized cross power spectral density calculation, the at least one acoustic propagation model comprising at least one Green's Function estimation
Implementation Method 3
applying a whitening filter to a spatially filtered target audio signal to derive at least one separated audio output signal
Data Source
AI summary
A spatial audio processing system operable to enable audio signals to be spatially extracted from, or transmitted to, discrete locations within an acoustic space. Embodiments of the present disclosure enable an array of transducers being installed in an acoustic space to combine their signals via inverting physical and environmental models that are measured, learned, tracked, calculated, or estimated. The models may be combined with a whitening filter to establish a cooperative or non-cooperative information-bearing channel between the array and one or more discrete, targeted physical locations in the acoustic space by applying the inverted models with whitening filter to the received or transmitted acoustical signals. The spatial audio processing system may utilize a model of the combination of direct and indirect reflections in the acoustic space to receive or transmit acoustic information, regardless of ambient noise levels, reverberation, and positioning of physical interferers.


