Dialog Enhancement in Low Complexity Audio Decoders
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Low complexity decoders in channel-based audio systems cannot directly apply existing dialog enhancement methods as they lack the full channel configuration necessary for dialog enhancement parameters, which are defined with respect to the full channel configuration.
Innovation Solution
A method that allows dialog enhancement in low complexity decoders by receiving downmix signals, dialog enhancement parameters defined for a subset of channels, and reconstruction parameters to upmix and enhance only the required channels, reducing computational complexity and improving audio quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If full channel configuration is decoded to enable dialog enhancement, then dialog enhancement quality is improved, but decoder complexity increases
Solution Approach 1:
The patent divides the channel configuration into two parts: downmixed channels that are decoded by the low complexity decoder, and virtual channels that are synthesized through parametric upmixing. This segmentation allows dialog enhancement to be applied selectively to specific channels without requiring full channel configuration decoding, thus maintaining low complexity while preserving dialog enhancement quality.
Solution Approach 2:
The patent uses parametric reconstruction parameters to transform downmixed channels into virtual channels that approximate the full channel configuration. By changing the representation from actual decoded channels to parametrically synthesized channels, the system enables dialog enhancement without the computational burden of full decoding.
2Device complexity
If low complexity decoder is used to reduce computational complexity, then decoder complexity is reduced, but dialog enhancement cannot be applied directly
Solution Approach 1:
The patent introduces parametric upmixing as an intermediary process between downmix decoding and dialog enhancement. This intermediary step synthesizes virtual channels that serve as a bridge, allowing dialog enhancement parameters (defined for full channel configuration) to be applied to downmixed content without requiring full channel decoding.
Solution Approach 2:
The patent performs parametric upmixing in advance to create virtual channels before applying dialog enhancement. This preliminary action prepares the channel structure in a form that is compatible with dialog enhancement parameters, enabling low complexity decoders to apply dialog enhancement without directly decoding the full channel configuration.
3Manufacturing precision
If full channel configuration is decoded, then dialog enhancement parameters can be applied, but processing time increases
Solution Approach 1:
The patent applies partial action by decoding only the downmixed channels rather than the full channel configuration, then using parametric upmixing to synthesize only the necessary virtual channels for dialog enhancement. This partial decoding approach reduces processing time while maintaining sufficient quality for dialog enhancement applications.
Data Source
AI summary
There is provided a method for enhancing dialog in a decoder of an audio system. The method comprises receiving a plurality of downmix signals being a downmix of a larger plurality of channels; receiving parameters for dialog enhancement being defined with respect to a subset of the plurality of channels that is downmixed into a subset of the plurality of downmix signals; upmixing the subset of downmix signals parametrically in order to reconstruct the subset of the plurality of channels with respect to which the parameters for dialog enhancement are defined; applying dialog enhancement to the subset of the plurality of channels with respect to which the parameters for dialog enhancement are defined using the parameters for dialog enhancement to provide at least one dialog enhanced signal; and subjecting the at least one dialog enhanced signal to mixing to provide dialog enhanced versions of the subset of downmix signals.


