Sub-band Mixing for Consistent Timbre in Multi-Mic Arrays

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio mixing technologies using multiple microphones face challenges with noticeable fluctuations in timbre and sound level due to gating, and increased noise and reverberation when multiple talkers are active simultaneously.

Innovation Solution

The implementation of subband mixing techniques that transform microphone signals into frequency audio data, apply weights based on smoothed spectral powers and noise floors, and integrate signals across multiple microphones to enhance direct speech and suppress noise and reverberations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If gating is used to mute microphones when active talkers are not detected, then microphone signal processing is simplified, but fluctuations in perceived timbre and sound level become noticeable

Engineering Contradiction:
Improvemicrophone signal processing complexityVSAvoidperceived timbre consistency
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent divides the audio frequency spectrum into multiple subbands and processes each subband separately. This segmentation allows the system to apply different gain values to different frequency regions, smoothing transitions when microphones are switched and maintaining consistent perceived timbre while still enabling selective microphone muting based on talker detection.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different processing characteristics to different subbands rather than uniformly across the entire frequency spectrum. By calculating subband-specific gain values based on local energy distribution and talker presence, the system maintains optimal signal quality in each frequency region while reducing audible artifacts during microphone transitions.

Inventive Principle:
Principle #3Local quality

2Object-affected harmful factors

If multiple microphones are processed by gating, then noise and reverberation can be reduced, but when multiple talkers are active simultaneously, noise and reverberation levels in the audio mix increase

Engineering Contradiction:
Improvenoise and reverberation levelsVSAvoidhandling of multiple simultaneous talkers
Core Design Contradiction:
Object-affected harmful factorsVSAdaptability or versatility

Solution Approach 1:

The patent segments the audio spectrum into multiple subbands and independently processes each subband. This allows the system to identify and suppress noise and reverberation in each frequency region separately, while still preserving multiple simultaneous talkers by maintaining their respective subband energy distributions. The segmented approach prevents the accumulation of noise and reverberation that occurs with traditional full-band gating methods.

Inventive Principle:
Principle #1Segmentation

3Device complexity

If traditional audio mixing is used with multiple microphones, then implementation is simpler, but speech intelligibility and direct sound enhancement are reduced

Engineering Contradiction:
Improveaudio mixing implementationVSAvoidspeech intelligibility
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent implements subband processing that divides the audio spectrum into multiple frequency regions, allowing selective enhancement of direct speech components in each subband. This segmentation enables the system to amplify speech-related frequency components while suppressing noise and reverberation, significantly improving speech intelligibility compared to traditional full-band mixing approaches.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10623854B2Sub-band mixing of multiple microphones
Publication Date: 2020.04.14 DOLBY LABORATORIES LICENSING CORP
  • US10623854B2 patent drawing
  • US10623854B2 patent drawing
  • US10623854B2 patent drawing

AI summary

Input audio data portions of a common time window index value generated by multiple microphones at a location are received. Subband portions are generated from the input audio data portions. Peak powers, noise floors, etc., are individually determined for the subband portions. Weights for the subband portions are computed based on the peak powers, the noise floors, etc., for the subband portions. An integrated audio data portion of the common time window index is generated based on the subband portions and the weight values for the subband portions. An integrated signal may be generated based at least in part on the integrated audio data portion.