Multi-Channel Audio Voice Enhancement via Gain Function Weighting

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for enhancing voice components in multi-channel audio signals, such as boosting the center channel or using dynamic range compression, suffer from low performance and fail to effectively isolate voice components from non-voice components, especially in complex audio environments.

Innovation Solution

A signal processing apparatus and method that filters multi-channel audio signals using a gain function determined from all channels, employing a Wiener filtering approach to weight and combine left, center, and right channel audio signals, while incorporating voice activity detection to enhance voice components and suppress non-speech signals.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If the center channel audio signal is boosted to enhance voice component, then voice enhancement is achieved, but the performance of voice enhancement is low and non-voice components are not effectively suppressed

Engineering Contradiction:
Improvevoice enhancement performanceVSAvoidnon-voice component interference
Core Design Contradiction:
Measurement precisionVSObject-generated harmful factors

Solution Approach 1:

The audio signal is segmented into center channel and side channel components. The center channel is extracted and processed separately to enhance voice, while side channels are attenuated to suppress non-voice components. This segmentation allows selective enhancement of voice without amplifying interfering non-voice content.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Different processing strategies are applied to different spatial locations in the audio field. The center channel receives gain boosting for voice enhancement, while side channels receive attenuation for non-voice suppression. This local differentiation optimizes voice enhancement performance while minimizing interference from non-voice components in specific spatial regions.

Inventive Principle:
Principle #3Local quality

2Measurement precision

If dynamic range compression is applied to attenuate loud non-voice components and boost soft voice components, then loudness level is improved, but the nature of multi-channel audio signal is not considered and voice enhancement performance is limited

Engineering Contradiction:
Improvevoice component detection accuracyVSAvoidsignal processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The audio signal is divided into center and side channel components, allowing separate processing of voice and non-voice content. This segmentation enables accurate voice component detection by focusing on the center channel where voice is predominantly located, while simplifying the overall processing complexity through channel-specific operations.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system dynamically adjusts processing based on voice activity detection. Voice activity is detected in the center channel, and processing parameters are dynamically modified accordingly - boosting center channel during voice activity while attenuating side channels. This dynamic adaptation improves voice detection accuracy without requiring complex static processing of all channels.

Inventive Principle:
Principle #15Dynamics

3Measurement precision

If voice activity detection is performed on all channels to account for voice variation over time, then voice enhancement accuracy is improved, but the computational complexity increases

Engineering Contradiction:
Improvevoice activity detection accuracyVSAvoidcomputational energy consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by stationary object

Solution Approach 1:

Voice activity detection is performed locally on the center channel where voice is predominantly located, rather than uniformly across all channels. This localized approach maintains high voice detection accuracy while reducing computational energy consumption by focusing processing resources on the most relevant channel for voice activity.

Inventive Principle:
Principle #3Local quality

Data Source

PatentEP3204945B1A signal processing apparatus for enhancing a voice component within a multi-channel audio signal
Publication Date: 2019.10.16 HUAWEI TECH CO LTD
  • EP3204945B1 patent drawingFigure 1
  • EP3204945B1 patent drawingFigure 2
  • EP3204945B1 patent drawingFigure 3

AI summary

The invention relates to a signal processing apparatus (100) for enhancing a voice component within a multi-channel audio signal, the multi-channel audio signal comprising a left channel audio signal (L), a center channel audio signal (C), and a right channel audio signal (R), the signal processing apparatus (100) comprising a filter (101) and a combiner (103); wherein the filter (101) is configured to determine a measure representing an overall magnitude of the multi-channel audio signal over frequency upon the basis of the left channel audio signal (L), the center channel audio signal (C), and the right channel audio signal (R), to obtain a gain function (G) based on a ratio between a measure of magnitude of the center channel audio signal (C) and the measure representing the overall magnitude of the multi- channel audio signal, and to weight the left channel audio signal (L) by the gain function (G) to obtain a weighted left channel audio signal (LE), to weight the center channel audio signal (C) by the gain function (G) to obtain a weighted center channel audio signal (CE), and to weight the right channel audio signal (R) by the gain function (G) to obtain a weighted right channel audio signal (RE); and wherein the combiner (103) is configured to combine the left channel audio signal (L) with the weighted left channel audio signal (LE) to obtain a combined left channel audio signal (LEV), to combine the center channel audio signal (C) with the weighted center channel audio signal (CE) to obtain a combined center channel audio signal (CEV), and to combine the right channel audio signal (R) with the weighted right channel audio signal (RE) to obtain a combined right channel audio signal (REV).