Audio Bitstream Loudness Metadata for Device-Specific Playback
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio encoding systems, such as those using AC-3 bitstreams, face challenges in accurately setting and maintaining loudness parameters, leading to inconsistent audio playback quality across different devices and environments due to reliance on user-defined dialnorm values, potential measurement errors, and metadata changes during transmission and storage.
Innovation Solution
The method involves analyzing metadata in audio bitstreams to determine available loudness parameters for specific groups of devices, with processing components adjusting audio data accordingly to ensure optimal playback, and embedding loudness processing state metadata to maintain accurate loudness and dynamic range control across various playback devices and environments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If user-defined dialnorm values are used for loudness control, then the system is simple to operate, but the loudness measurement precision and consistency across different devices deteriorates
Solution Approach 1:
The system automatically measures loudness using ITU-R BS.1770 standards and embeds the measured values in metadata, eliminating the need for manual user input while ensuring consistent, standardized measurements across all content
Solution Approach 2:
The system embeds loudness measurement results in metadata that is carried through the audio bitstream, providing feedback information that enables automatic gain adjustment at playback to achieve consistent loudness across different devices and content
2Reliability
If metadata is embedded in audio bitstreams for loudness control, then the loudness information is preserved, but the metadata may be changed or lost during transmission and storage
Solution Approach 1:
The system performs preliminary loudness measurements and embeds the results in metadata during encoding, so that the loudness information is prepared and protected before transmission or storage, reducing the risk of information loss
Solution Approach 2:
The system uses standardized metadata fields (such as dialnorm or alternative loudness metadata) as intermediaries to carry loudness information through the audio bitstream, ensuring compatibility and reducing the risk of information loss during transmission and storage
3Adaptability or versatility
If a single audio bitstream is used for multiple playback devices, then the system is versatile, but the audio quality optimization for each specific device deteriorates
Solution Approach 1:
The system embeds device-specific or content-specific loudness metadata in the audio bitstream, allowing different playback devices to apply appropriate gain adjustments based on their characteristics, achieving optimized audio quality for each device while maintaining a single source bitstream
Solution Approach 2:
The system uses dynamic metadata that can be interpreted differently by different playback devices, allowing the same audio bitstream to be adaptively optimized for various devices with different loudness characteristics and capabilities
4Ease of manufacture
If dynamic range compression is applied to fit wide dynamic range content into narrower recorded dynamic range, then the content is more easily stored and reproduced, but the dynamic range and audio quality are reduced
Solution Approach 1:
The system extracts and preserves the original loudness information in metadata while applying dynamic range compression to the audio content, allowing the compressed audio to be stored and reproduced easily while the metadata enables recovery of the original dynamic range characteristics during playback
Data Source
AI summary
Embodiments are directed to a method and system for receiving, in a bitstream, metadata associated with the audio data, and analyzing the metadata to determine whether a loudness parameter for a first group of audio playback devices are available in the bitstream. Responsive to determining that the parameters are present for the first group, the system uses the parameters and audio data to render audio. Responsive to determining that the loudness parameters are not present for the first group, the system analyzes one or more characteristics of the first group, and determines the parameter based on the one or more characteristics.


