Audio Signal Mixing with Sub-Band Separation for Speech Intelligibility

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio processing technologies fail to effectively improve speech intelligibility for listeners using mono-channel devices and struggle to distinguish multiple audio signals, especially when signals originate from similar positions or lack spatial cues.

Innovation Solution

The proposed audio processing method involves spectral, spatial, and temporal separation techniques, including sub-band suppression, spatial auditory property assignment, and time scaling, to enhance intelligibility by minimizing energetic and informational masking effects, and improving the perception of audio signals' spatial origins.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If multiple audio signals are mixed and transmitted through mono-channel devices, then communication coverage is expanded, but speech intelligibility deteriorates due to masking effects

Engineering Contradiction:
Improvedevice compatibilityVSAvoidspeech intelligibility
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The audio spectrum is segmented into multiple sub-bands, and different audio signals are assigned to different sub-bands. This frequency domain segmentation allows multiple signals to coexist in a mono-channel without mutual masking, as each signal occupies a distinct frequency region that the human ear can distinguish.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transforms the mixing problem from the time domain to the frequency domain by applying sub-band coding. This dimensional transformation allows multiple signals that would be masking each other in time to be separated and transmitted simultaneously in frequency, maintaining intelligibility while using a single channel.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Loss of information

If spectral filtering is applied to separate audio signals, then speech intelligibility is improved, but audio quality and frequency information are lost

Engineering Contradiction:
Improvespeech intelligibilityVSAvoidaudio quality
Core Design Contradiction:
Loss of informationVSManufacturing precision

Solution Approach 1:

Different sub-bands are allocated to different audio signals based on their characteristics and requirements. Each signal receives optimized frequency allocation tailored to its needs, allowing high-quality transmission of speech in critical frequency regions while maintaining overall audio fidelity through selective filtering rather than aggressive compression.

Inventive Principle:
Principle #3Local quality

3Loss of information

If spatial cues are added to distinguish multiple speakers, then speech separation is improved for mono-channel listeners, but device complexity increases

Engineering Contradiction:
Improvespeaker distinguishabilityVSAvoidprocessing complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent uses spatial audio processing as an intermediary layer that translates multi-speaker scenes into separable spectral patterns. By introducing spatial cues that manifest as distinct frequency characteristics, the system enables mono-channel devices to separate speakers without requiring complex hardware modifications, as the separation is achieved through signal processing rather than physical device complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS9602943B2Audio processing method and audio processing apparatus
Publication Date: 2017.03.21 DOLBY LABORATORIES LICENSING CORP
  • US9602943B2 patent drawing
  • US9602943B2 patent drawing
  • US9602943B2 patent drawing

AI summary

An audio processing method and apparatus are described. In one embodiment, at least one first sub-band of a first audio signal is suppressed to obtain a reduced first audio signal with reserved sub-bands; suppressing at least one second sub-band of the at least one second audio signal to obtain at least one reduced second audio signal with reserved sub-bands; and mixing the reduced first audio signal and at least one reduced second audio signal. Alternatively, a first spatial auditory property is assigned to a first audio signal so that the first audio signal may be perceived as originating from a first position. Alternatively, rhythmic similarity between at least two audio signals is detected, and time scaling is applied to an audio signal in response to relatively high rhythmic similarity between the audio signal and the other audio signal(s); and then at least two audio signals are mixed.