Audio Signal Processing Device for Clear Dialogue in Downmixed Multichannel Audio
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
During the downmixing of multichannel audio data, dialogue sounds often become unclear due to gain suppression, leading to reduced sound reproduction quality.
Innovation Solution
An audio signal processing device that selects dialogue sound channels and downmixes them separately, allowing for gain correction and targeted addition to specific channels, ensuring clear dialogue sound reproduction.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If gain suppression correction is applied to suppress clipping in downmixing, then signal addition stability is improved, but dialogue sound volume is reduced
Solution Approach 1:
The audio signal processing is segmented into two distinct paths: one for dialogue sound channels and another for other audio channels. Dialogue sound channels are processed separately with gain correction applied specifically to them, while other channels undergo standard downmixing. This segmentation allows the system to maintain dialogue sound volume through targeted gain adjustment without causing clipping in the overall mixed signal.
Solution Approach 2:
Different gain correction strategies are applied to different parts of the audio signal based on their characteristics. Dialogue sound channels receive specific gain correction to maintain their volume and clarity, while other channels are processed with standard downmixing gain control. This local quality approach ensures that each part of the audio signal is optimized according to its specific requirements.
2Adaptability or versatility
If downmixing is applied to convert multichannel signals to fewer channels, then reproduction compatibility is improved, but dialogue sound localization becomes unclear
Solution Approach 1:
Dialogue sound channels are extracted from the multichannel audio signal before downmixing. By separating dialogue channels and processing them independently with gain correction, the system preserves the spatial localization information of dialogue sounds while still achieving downmixing compatibility for other audio content. The extracted dialogue channels are then added back to the downmixed signal.
Solution Approach 2:
The audio signal is segmented into dialogue channels and other channels, with different processing applied to each segment. This segmentation allows the system to maintain localization precision for dialogue sounds by preserving their channel separation, while still achieving reproduction compatibility through downmixing of the overall signal.
3Adaptability or versatility
If dialogue sound channels are distributed to multiple channels during downmixing, then reproduction flexibility is improved, but dialogue sound clarity is reduced
Solution Approach 1:
Dialogue sound channels are extracted and processed separately from the downmixing process. This extraction prevents the distribution of dialogue signals across multiple channels that would occur in standard downmixing, thereby maintaining dialogue sound clarity. The extracted channels are then independently gain-corrected and added back to the downmixed signal.
Solution Approach 2:
Different processing quality is applied to different parts of the audio signal. Dialogue channels receive specialized processing that maintains their clarity and prevents distribution, while other channels undergo standard downmixing that provides reproduction flexibility. This local quality differentiation resolves the contradiction between flexibility and clarity.
Data Source
AI summary
The present technology relates to an audio signal processing device and method, an encoding device and method, and a program, which are capable of obtaining a higher quality sound. A selection unit selects, from supplied multichannel audio signals, audio signals of a channel of a dialogue sound and audio signals of a channel to be downmixed. A downmixing unit downmixes the audio signals of the channel to be downmixed. An addition unit adds the audio signals of the channel of a dialogue sound to audio signals of a predetermined channel among audio signals of one or more channels obtained in the downmixing. The present technology can be applied to a decoder.


