Audio Signal Processing Device for Clear Dialogue in Downmixed Multichannel Audio

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

During the downmixing of multichannel audio data, dialogue sounds often become unclear due to gain suppression, leading to reduced sound reproduction quality.

Innovation Solution

An audio signal processing device that selects dialogue sound channels and downmixes them separately, allowing for gain correction and targeted addition to specific channels, ensuring clear dialogue sound reproduction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If gain suppression correction is applied to suppress clipping in downmixing, then signal addition stability is improved, but dialogue sound volume is reduced

Engineering Contradiction:
Improvesignal addition stabilityVSAvoiddialogue sound volume
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The audio signal processing is segmented into two distinct paths: one for dialogue sound channels and another for other audio channels. Dialogue sound channels are processed separately with gain correction applied specifically to them, while other channels undergo standard downmixing. This segmentation allows the system to maintain dialogue sound volume through targeted gain adjustment without causing clipping in the overall mixed signal.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Different gain correction strategies are applied to different parts of the audio signal based on their characteristics. Dialogue sound channels receive specific gain correction to maintain their volume and clarity, while other channels are processed with standard downmixing gain control. This local quality approach ensures that each part of the audio signal is optimized according to its specific requirements.

Inventive Principle:
Principle #3Local quality

2Adaptability or versatility

If downmixing is applied to convert multichannel signals to fewer channels, then reproduction compatibility is improved, but dialogue sound localization becomes unclear

Engineering Contradiction:
Improvereproduction compatibilityVSAvoiddialogue sound localization
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

Dialogue sound channels are extracted from the multichannel audio signal before downmixing. By separating dialogue channels and processing them independently with gain correction, the system preserves the spatial localization information of dialogue sounds while still achieving downmixing compatibility for other audio content. The extracted dialogue channels are then added back to the downmixed signal.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The audio signal is segmented into dialogue channels and other channels, with different processing applied to each segment. This segmentation allows the system to maintain localization precision for dialogue sounds by preserving their channel separation, while still achieving reproduction compatibility through downmixing of the overall signal.

Inventive Principle:
Principle #1Segmentation

3Adaptability or versatility

If dialogue sound channels are distributed to multiple channels during downmixing, then reproduction flexibility is improved, but dialogue sound clarity is reduced

Engineering Contradiction:
Improvereproduction flexibilityVSAvoiddialogue sound clarity
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

Dialogue sound channels are extracted and processed separately from the downmixing process. This extraction prevents the distribution of dialogue signals across multiple channels that would occur in standard downmixing, thereby maintaining dialogue sound clarity. The extracted channels are then independently gain-corrected and added back to the downmixed signal.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

Different processing quality is applied to different parts of the audio signal. Dialogue channels receive specialized processing that maintains their clarity and prevents distribution, while other channels undergo standard downmixing that provides reproduction flexibility. This local quality differentiation resolves the contradiction between flexibility and clarity.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS10621994B2Audio signal processing device and method, encoding device and method, and program
Publication Date: 2020.04.14 SONY GROUP CORP
  • US10621994B2 patent drawing
  • US10621994B2 patent drawing
  • US10621994B2 patent drawing

AI summary

The present technology relates to an audio signal processing device and method, an encoding device and method, and a program, which are capable of obtaining a higher quality sound. A selection unit selects, from supplied multichannel audio signals, audio signals of a channel of a dialogue sound and audio signals of a channel to be downmixed. A downmixing unit downmixes the audio signals of the channel to be downmixed. An addition unit adds the audio signals of the channel of a dialogue sound to audio signals of a predetermined channel among audio signals of one or more channels obtained in the downmixing. The present technology can be applied to a decoder.