Audio Signal Mixing with Sub-Band Separation for Speech Intelligibility
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio processing technologies fail to effectively improve speech intelligibility for listeners using mono-channel devices and struggle to distinguish multiple audio signals, especially when signals originate from similar positions or lack spatial cues.
Innovation Solution
The proposed audio processing method involves spectral, spatial, and temporal separation techniques, including sub-band suppression, spatial auditory property assignment, and time scaling, to enhance intelligibility by minimizing energetic and informational masking effects, and improving the perception of audio signals' spatial origins.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple audio signals are mixed and transmitted through mono-channel devices, then communication coverage is expanded, but speech intelligibility deteriorates due to masking effects
Solution Approach 1:
The audio spectrum is segmented into multiple sub-bands, and different audio signals are assigned to different sub-bands. This frequency domain segmentation allows multiple signals to coexist in a mono-channel without mutual masking, as each signal occupies a distinct frequency region that the human ear can distinguish.
Solution Approach 2:
The patent transforms the mixing problem from the time domain to the frequency domain by applying sub-band coding. This dimensional transformation allows multiple signals that would be masking each other in time to be separated and transmitted simultaneously in frequency, maintaining intelligibility while using a single channel.
2Loss of information
If spectral filtering is applied to separate audio signals, then speech intelligibility is improved, but audio quality and frequency information are lost
Solution Approach 1:
Different sub-bands are allocated to different audio signals based on their characteristics and requirements. Each signal receives optimized frequency allocation tailored to its needs, allowing high-quality transmission of speech in critical frequency regions while maintaining overall audio fidelity through selective filtering rather than aggressive compression.
3Loss of information
If spatial cues are added to distinguish multiple speakers, then speech separation is improved for mono-channel listeners, but device complexity increases
Solution Approach 1:
The patent uses spatial audio processing as an intermediary layer that translates multi-speaker scenes into separable spectral patterns. By introducing spatial cues that manifest as distinct frequency characteristics, the system enables mono-channel devices to separate speakers without requiring complex hardware modifications, as the separation is achieved through signal processing rather than physical device complexity.
Data Source
AI summary
An audio processing method and apparatus are described. In one embodiment, at least one first sub-band of a first audio signal is suppressed to obtain a reduced first audio signal with reserved sub-bands; suppressing at least one second sub-band of the at least one second audio signal to obtain at least one reduced second audio signal with reserved sub-bands; and mixing the reduced first audio signal and at least one reduced second audio signal. Alternatively, a first spatial auditory property is assigned to a first audio signal so that the first audio signal may be perceived as originating from a first position. Alternatively, rhythmic similarity between at least two audio signals is detected, and time scaling is applied to an audio signal in response to relatively high rhythmic similarity between the audio signal and the other audio signal(s); and then at least two audio signals are mixed.


