Audio Decoder DRC Profile Selection for Loudness-Adjusted Playback
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio signal processing technologies face challenges in maintaining high-quality and intelligibility across a broad range of rendering devices with different capabilities and environments, as they often fail to adapt dynamic range appropriately, leading to distortion or inaudibility.
Innovation Solution
The method involves embedding multiple Dynamic Range Control (DRC) profiles within audio frames, allowing an audio decoder to select the appropriate profile for the specific rendering mode, ensuring high-quality playback without distortion by adjusting loudness levels based on dynamic range compression curves and metadata.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a single fixed dynamic range is used in audio encoding, then the encoding process is simple, but the audio quality deteriorates on devices with different rendering capabilities and in different listening environments
Solution Approach 1:
The audio encoding process is segmented into multiple versions with different dynamic range characteristics. Instead of encoding a single version, the system creates multiple encoded versions (e.g., different DRC profiles) that can be selected based on the playback device capabilities and listening environment, thereby resolving the contradiction between adaptability and encoding simplicity
Solution Approach 2:
The encoding system transitions from a static single-version approach to a dynamic multi-version approach. The encoder can adaptively select or generate appropriate dynamic range profiles based on metadata about the playback device and environment, enabling the system to be versatile without requiring complex real-time processing at playback
2Manufacturing precision
If dynamic range is expanded to maintain audio quality on high-end devices, then audio fidelity is improved, but distortion and inaudibility occur on devices with limited rendering capabilities or in noisy environments
Solution Approach 1:
Multiple encoded versions with different dynamic range parameters are created during encoding. Each version has optimized parameters (such as different compression ratios, threshold levels, or gain structures) suitable for specific device capabilities and environmental conditions. The playback system selects the appropriate version based on detected conditions, thereby maintaining audio fidelity across different scenarios without causing distortion
Solution Approach 2:
The system performs preliminary encoding of multiple dynamic range profiles during the encoding phase rather than attempting to adapt in real-time at playback. This preliminary preparation of multiple versions allows the playback system to simply select the appropriate pre-encoded version based on device capabilities, avoiding the need for complex real-time processing that could introduce distortion
3Adaptability or versatility
If multiple DRC profiles are transmitted for every frame, then adaptability to different rendering modes is maximized, but bandwidth consumption increases significantly
Solution Approach 1:
The transmission is segmented such that only necessary DRC profile information is sent in each frame. Instead of transmitting complete DRC profiles for every frame, the system sends profile identifiers or differential updates only when the rendering mode actually changes, significantly reducing bandwidth consumption while maintaining adaptability
Solution Approach 2:
The DRC profile data structure is designed to serve multiple rendering modes simultaneously. A single set of encoded audio data with embedded DRC profiles can be adapted to different playback scenarios (headphones, speakers, noisy environments, quiet rooms) without requiring separate encodings for each mode, thereby reducing overall bandwidth requirements
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Decoding an encoded audio signal comprising frames of encoded audio data and metadata, the metadata including different dynamic range control, DRC, gains. Decoding according to a DRC gain profile corresponding to a desired output reference level. Determining a loudness related gain based on the indication of the loudness level of the audio signal and the desired output reference level. Applying the loudness related gain to obtain loudness adjusted audio data that has the desired output reference level.