Audio Loudness Metadata Control for Consistent Playback Levels
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The existing methods for audio signal processing struggle to maintain consistent loudness levels across different audio content, leading to inconvenience for users who need to repeatedly adjust volumes, and there is a lack of effective standards for loudness normalization.
Innovation Solution
An audio signal processing device and method that includes a receiver, processor, and outputter to generate and transmit loudness metadata, using the Quality Secure Histogram Index (QSHI) to adjust the output loudness level, ensuring it does not exceed a threshold that causes cognitive sound quality damage, and applying a loudness limiter to maintain stable output.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If content creators increase the loudness of audio signals to improve perceived sound quality, then sound quality perception is improved, but loudness difference between contents increases and users need to repeatedly adjust volume
Solution Approach 1:
The system performs preliminary loudness measurement and metadata generation during content production or pre-processing. The loudness metadata (including QSHI values) is embedded in the content before playback, enabling the playback device to automatically apply appropriate gain adjustments without requiring user intervention during operation.
Solution Approach 2:
The system uses measured loudness information and QSHI values as feedback to automatically control the output loudness level. The playback device reads the loudness metadata, compares it with target loudness levels, and applies compensatory gain adjustments to maintain consistent perceived loudness across different contents.
2Ease of operation
If loudness normalization is applied to maintain consistent output loudness, then user convenience is improved, but cognitive sound quality damage may occur if threshold is exceeded
Solution Approach 1:
The system changes the parameter representation from simple loudness level to QSHI (Quality Secure Histogram Index), which incorporates both loudness information and sound quality considerations. By using QSHI as the control parameter, the system can adjust loudness while maintaining the threshold below which cognitive sound quality damage does not occur.
Solution Approach 2:
The QSHI value acts as an intermediary parameter between the input audio signal and the output loudness control. Instead of directly controlling output loudness based on simple loudness measurements, the system uses QSHI (which represents the threshold loudness level) as a mediator to determine safe gain adjustments that prevent cognitive sound quality damage.
3Ease of operation
If loudness metadata is generated and transmitted to control output loudness, then loudness normalization is achieved, but device complexity increases
Solution Approach 1:
The system extracts only the essential loudness information (QSHI value) from the complex audio signal and represents it as compact metadata. This extracted metadata can be easily transmitted and processed, avoiding the need to handle the entire audio signal for loudness control decisions.
Solution Approach 2:
The system replaces complex mechanical/audio processing approaches with a data-driven approach using loudness metadata. Instead of relying on real-time complex signal processing during playback, the system uses pre-computed loudness metadata to guide simple gain adjustments, reducing processing complexity.
Data Source
AI summary
An audio signal processing device comprises: a receiver for receiving an input audio signal; a processor for generating loudness metadata corresponding to the input audio signal; and an outputter for transmitting the loudness metadata generated by the processor. The processor is configured to acquire loudness information analyzed from input content, acquires loudness information about the input audio signal by measuring the loudness of the input audio signal, generates the loudness metadata by converting the loudness information, and transmits, through the outputter, the generated loudness metadata to an output device for outputting the input audio signal.


