Frequency-Selective Microphone Combining for Smooth Speaker Switching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In multi-microphone systems, especially in noisy environments like cars, existing technologies face challenges in combining microphone signals effectively, leading to noticeable noise and speech artifacts during speaker changes due to varying noise levels and characteristics among microphones.
Innovation Solution
The implementation of frequency-selective signal combining using noise reduction and automatic gain control processes, where microphone signals are transformed into the frequency subband domain, and frequency-based channel selection is performed using speaker activity detection information to choose the channel with the best signal-to-noise ratio, reducing processing resources and complexity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If hard switching between active speakers is done, then processing resources are reduced, but noise jumps and switching artifacts become noticeable
Solution Approach 1:
The patent applies frequency-selective signal combination where different frequency subbands are processed differently. For each frequency subband, the system selects the microphone signal with the best signal-to-noise ratio, rather than applying a uniform switching approach across all frequencies. This local optimization in the frequency domain resolves the contradiction by maintaining high signal quality in each frequency band while keeping processing efficient through selective combination rather than full multi-channel processing.
Solution Approach 2:
The patent segments the audio signal into multiple frequency subbands using a filter bank. This segmentation allows independent processing of each frequency band, enabling the system to select the best microphone for each subband based on noise characteristics. By dividing the broadband signal into narrowband components, the system achieves smooth transitions and avoids switching artifacts while maintaining processing efficiency.
2Stability of the object's composition
If soft mixing functions include higher noise level, then speech continuity is maintained, but the resulting noise level increases
Solution Approach 1:
The patent processes each frequency subband independently and selects the microphone signal with the best signal-to-noise ratio for each subband. This frequency-selective approach allows the system to maintain speech continuity through smooth transitions while avoiding the inclusion of high-noise frequency components in the final output. The local optimization in each frequency band ensures that only clean speech components are combined, resolving the contradiction between continuity and noise level.
Solution Approach 2:
The patent converts the harmful effect of varying noise levels across microphones into a beneficial selection criterion. By analyzing the noise characteristics of each microphone in each frequency subband, the system identifies and selects the microphones with the lowest noise levels. This transforms the problem of noise variation into an opportunity for optimal signal selection, maintaining continuity while minimizing overall noise.
3Reliability
If one noise reduction process and one automatic gain control process per input channel are used, then signal quality is maintained, but processing resources and system complexity increase significantly
Solution Approach 1:
The patent merges the noise reduction and automatic gain control processes into a single frequency-selective signal combination step. Instead of applying separate noise reduction and AGC processes to each microphone channel independently, the system combines the microphone signals in the frequency domain by selecting the best signal for each frequency subband. This unified approach maintains signal quality through optimal selection while dramatically reducing processing complexity by eliminating redundant per-channel processing.
Solution Approach 2:
The patent segments the processing into frequency subbands and performs selection independently in each band. This segmentation allows the system to apply a simple selection rule (choose the microphone with best SNR) in each frequency band rather than complex per-channel processing. The segmented approach maintains signal quality by preserving frequency-specific characteristics while reducing overall processing complexity through parallel independent decisions.
Data Source
AI summary
Methods and apparatus for frequency selective signal mixing for speech enhancement. In one embodiment frequency-based channel selection is performed for signal magnitude, signal energy, and noise estimate using speaker activity detection information, signal-to-noise ratio, and/or signal level, Frequency-based channel selection is performed for a dynamic spectral floor to adjust the noise estimate using speaker dominance information. Noise reduction is performed on the signal for the selected channel.


