Binaural Cue Preservation via Interaural Coherence Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Multichannel audio signal processing methods that aim to enhance target sources while suppressing interfering sources often collapse all input channels into a mono channel, resulting in distortion or loss of spatial information for both target and interfering sources.
Innovation Solution
The method employs interaural coherence to control the tradeoff between signal-to-noise ratio (SNR) and binaural quality by generating sound filters that prioritize either SNR or binaural information based on the interaural coherence, allowing for the preservation of binaural information of interfering sources without requiring high computational complexity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If all input channels are collapsed into a mono channel to suppress interfering sources, then signal-to-noise ratio is improved, but spatial information of target and interfering sources is distorted or completely removed
Solution Approach 1:
The patent segments the audio signal processing into multiple independent channels (left and right) rather than collapsing to mono. Each channel is processed separately with its own sound filters, allowing spatial information to be preserved while still enabling interference suppression through channel-specific filtering operations.
Solution Approach 2:
The patent transitions from a single-channel (mono) approach to a multi-channel (stereo/binaural) approach, adding the spatial dimension back into the signal processing. This dimensional change allows the system to suppress interfering sources while preserving spatial cues by operating in the additional channel dimension rather than reducing to a single dimension.
2Loss of information
If sound filters are designed to preserve binaural information of interfering sources, then spatial cues are maintained, but signal-to-noise ratio decreases
Solution Approach 1:
The patent applies local quality by making the sound filters channel-specific rather than uniform across all channels. Each channel (left and right) has its own optimized filters that are tailored to preserve local binaural characteristics while suppressing local interfering sources, allowing different filtering strategies for different spatial locations.
Solution Approach 2:
The patent implements dynamic control of the tradeoff between SNR and binaural quality through the interaural coherence parameter. The system can adaptively adjust the filtering strength and binaural preservation level based on the calculated interaural coherence, making the filter behavior dynamic rather than static to optimize performance under varying acoustic conditions.
3Loss of information
If complex processing methods are used to preserve binaural information, then spatial quality is improved, but computational complexity increases
Solution Approach 1:
The patent extracts the interaural coherence as a separate control parameter that governs the tradeoff between SNR and binaural quality. By taking out this specific metric and using it to control the filtering process, the system avoids the need for more complex joint optimization methods, simplifying the computational requirements while still achieving adaptive binaural preservation.
Data Source
AI summary
An audio signal is enhanced using interaural coherence to control noise reduction and binaural cue preservation. Sounds from a local area are detected via an acoustic array. An interaural coherence is determined using the detected sounds. Sound filters for an audio signal are generated based on the interaural coherence. The sound filters implement a tradeoff between increasing signal-to-noise ratio (SNR) between a target source and an interfering source and preserving of binaural information of the interfering source. The tradeoff is controlled based on the interaural coherence. The sound filters are applied to the audio signal to generate audio content. The audio content is presented via a speaker array.


