Multi-Channel Audio Voice Enhancement via Gain Function Weighting
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for enhancing voice components in multi-channel audio signals, such as boosting the center channel or using dynamic range compression, suffer from low performance and fail to effectively isolate voice components from non-voice components, especially in complex audio environments.
Innovation Solution
A signal processing apparatus and method that filters multi-channel audio signals using a gain function determined from all channels, employing a Wiener filtering approach to weight and combine left, center, and right channel audio signals, while incorporating voice activity detection to enhance voice components and suppress non-speech signals.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the center channel audio signal is boosted to enhance voice component, then voice enhancement is achieved, but the performance of voice enhancement is low and non-voice components are not effectively suppressed
Solution Approach 1:
The audio signal is segmented into center channel and side channel components. The center channel is extracted and processed separately to enhance voice, while side channels are attenuated to suppress non-voice components. This segmentation allows selective enhancement of voice without amplifying interfering non-voice content.
Solution Approach 2:
Different processing strategies are applied to different spatial locations in the audio field. The center channel receives gain boosting for voice enhancement, while side channels receive attenuation for non-voice suppression. This local differentiation optimizes voice enhancement performance while minimizing interference from non-voice components in specific spatial regions.
2Measurement precision
If dynamic range compression is applied to attenuate loud non-voice components and boost soft voice components, then loudness level is improved, but the nature of multi-channel audio signal is not considered and voice enhancement performance is limited
Solution Approach 1:
The audio signal is divided into center and side channel components, allowing separate processing of voice and non-voice content. This segmentation enables accurate voice component detection by focusing on the center channel where voice is predominantly located, while simplifying the overall processing complexity through channel-specific operations.
Solution Approach 2:
The system dynamically adjusts processing based on voice activity detection. Voice activity is detected in the center channel, and processing parameters are dynamically modified accordingly - boosting center channel during voice activity while attenuating side channels. This dynamic adaptation improves voice detection accuracy without requiring complex static processing of all channels.
3Measurement precision
If voice activity detection is performed on all channels to account for voice variation over time, then voice enhancement accuracy is improved, but the computational complexity increases
Solution Approach 1:
Voice activity detection is performed locally on the center channel where voice is predominantly located, rather than uniformly across all channels. This localized approach maintains high voice detection accuracy while reducing computational energy consumption by focusing processing resources on the most relevant channel for voice activity.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
The invention relates to a signal processing apparatus (100) for enhancing a voice component within a multi-channel audio signal, the multi-channel audio signal comprising a left channel audio signal (L), a center channel audio signal (C), and a right channel audio signal (R), the signal processing apparatus (100) comprising a filter (101) and a combiner (103); wherein the filter (101) is configured to determine a measure representing an overall magnitude of the multi-channel audio signal over frequency upon the basis of the left channel audio signal (L), the center channel audio signal (C), and the right channel audio signal (R), to obtain a gain function (G) based on a ratio between a measure of magnitude of the center channel audio signal (C) and the measure representing the overall magnitude of the multi- channel audio signal, and to weight the left channel audio signal (L) by the gain function (G) to obtain a weighted left channel audio signal (LE), to weight the center channel audio signal (C) by the gain function (G) to obtain a weighted center channel audio signal (CE), and to weight the right channel audio signal (R) by the gain function (G) to obtain a weighted right channel audio signal (RE); and wherein the combiner (103) is configured to combine the left channel audio signal (L) with the weighted left channel audio signal (LE) to obtain a combined left channel audio signal (LEV), to combine the center channel audio signal (C) with the weighted center channel audio signal (CE) to obtain a combined center channel audio signal (CEV), and to combine the right channel audio signal (R) with the weighted right channel audio signal (RE) to obtain a combined right channel audio signal (REV).