Frequency-Selective Microphone Combining for Smooth Speaker Switching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In multi-microphone systems, especially in noisy environments like cars, existing technologies face challenges in combining microphone signals effectively, leading to noticeable noise and speech artifacts during speaker changes due to varying noise levels and characteristics among microphones.

Innovation Solution

The implementation of frequency-selective signal combining using noise reduction and automatic gain control processes, where microphone signals are transformed into the frequency subband domain, and frequency-based channel selection is performed using speaker activity detection information to choose the channel with the best signal-to-noise ratio, reducing processing resources and complexity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If hard switching between active speakers is done, then processing resources are reduced, but noise jumps and switching artifacts become noticeable

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidsignal quality
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent applies frequency-selective signal combination where different frequency subbands are processed differently. For each frequency subband, the system selects the microphone signal with the best signal-to-noise ratio, rather than applying a uniform switching approach across all frequencies. This local optimization in the frequency domain resolves the contradiction by maintaining high signal quality in each frequency band while keeping processing efficient through selective combination rather than full multi-channel processing.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent segments the audio signal into multiple frequency subbands using a filter bank. This segmentation allows independent processing of each frequency band, enabling the system to select the best microphone for each subband based on noise characteristics. By dividing the broadband signal into narrowband components, the system achieves smooth transitions and avoids switching artifacts while maintaining processing efficiency.

Inventive Principle:
Principle #1Segmentation

2Stability of the object's composition

If soft mixing functions include higher noise level, then speech continuity is maintained, but the resulting noise level increases

Engineering Contradiction:
Improvespeech continuityVSAvoidnoise level
Core Design Contradiction:
Stability of the object's compositionVSObject-affected harmful factors

Solution Approach 1:

The patent processes each frequency subband independently and selects the microphone signal with the best signal-to-noise ratio for each subband. This frequency-selective approach allows the system to maintain speech continuity through smooth transitions while avoiding the inclusion of high-noise frequency components in the final output. The local optimization in each frequency band ensures that only clean speech components are combined, resolving the contradiction between continuity and noise level.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent converts the harmful effect of varying noise levels across microphones into a beneficial selection criterion. By analyzing the noise characteristics of each microphone in each frequency subband, the system identifies and selects the microphones with the lowest noise levels. This transforms the problem of noise variation into an opportunity for optimal signal selection, maintaining continuity while minimizing overall noise.

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

3Reliability

If one noise reduction process and one automatic gain control process per input channel are used, then signal quality is maintained, but processing resources and system complexity increase significantly

Engineering Contradiction:
Improvesignal qualityVSAvoidprocessing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent merges the noise reduction and automatic gain control processes into a single frequency-selective signal combination step. Instead of applying separate noise reduction and AGC processes to each microphone channel independently, the system combines the microphone signals in the frequency domain by selecting the best signal for each frequency subband. This unified approach maintains signal quality through optimal selection while dramatically reducing processing complexity by eliminating redundant per-channel processing.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent segments the processing into frequency subbands and performs selection independently in each band. This segmentation allows the system to apply a simple selection rule (choose the microphone with best SNR) in each frequency band rather than complex per-channel processing. The segmented approach maintains signal quality by preserving frequency-specific characteristics while reducing overall processing complexity through parallel independent decisions.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10536773B2Methods and apparatus for selective microphone signal combining
Publication Date: 2020.01.14 CERENCE OPERATING CO
  • US10536773B2 patent drawing
  • US10536773B2 patent drawing
  • US10536773B2 patent drawing

AI summary

Methods and apparatus for frequency selective signal mixing for speech enhancement. In one embodiment frequency-based channel selection is performed for signal magnitude, signal energy, and noise estimate using speaker activity detection information, signal-to-noise ratio, and/or signal level, Frequency-based channel selection is performed for a dynamic spectral floor to adjust the noise estimate using speaker dominance information. Noise reduction is performed on the signal for the selected channel.