Audio Loudness Metadata for Consistent Playback Volume

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio signal processing methods struggle to efficiently adjust the output loudness level of audio content, leading to inconsistencies and the need for frequent volume adjustments by users.

Innovation Solution

An audio signal processing device that includes a receiver, a processor, and an outputter, which generates and transmits loudness metadata to adjust the output loudness level of audio content based on acquired loudness information, including the Quality Secure Histogram Index (QSHI).

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If content creators increase the sound magnitude of audio signals to improve perceived sound quality, then sound quality perception is improved, but loudness difference between contents increases and users experience inconvenience

Engineering Contradiction:
Improvesound qualityVSAvoidvolume adjustment convenience
Core Design Contradiction:
Manufacturing precisionVSEase of operation

Solution Approach 1:

The patent applies preliminary action by pre-calculating and embedding loudness metadata (including QSHI values) into audio content during the content creation process. This allows playback devices to automatically retrieve and apply the stored loudness information, normalizing volume levels before playback without requiring user intervention. The loudness histogram and QSHI thresholds are computed in advance and stored with the content, enabling automatic loudness normalization that resolves the contradiction between sound quality enhancement and operational convenience.

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If different loudness levels are set for each audio content during production, then content-specific loudness optimization is achieved, but users must repeatedly adjust volume for different contents

Engineering Contradiction:
Improvecontent-specific loudness optimizationVSAvoidtime for volume adjustment
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent implements feedback by using pre-computed loudness metadata (loudness histograms and QSHI values) stored in the audio content to automatically adjust playback volume. The playback device retrieves the loudness information, compares it with a reference level, and applies appropriate gain adjustment without user input. This closed-loop approach maintains content-specific loudness optimization while eliminating the time users would otherwise spend manually adjusting volume between different audio contents.

Inventive Principle:
Principle #23Feedback

3Ease of operation

If loudness normalization is implemented to improve user convenience, then volume consistency across contents is improved, but cognitive sound quality damage may occur if not properly controlled

Engineering Contradiction:
Improvevolume consistencyVSAvoidsound quality integrity
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent applies parameter changes by utilizing the QSHI (Quality Secure Histogram Index) threshold parameter derived from loudness histogram analysis. The system adjusts playback gain dynamically based on the relationship between the content's loudness characteristics and the QSHI threshold, ensuring that normalization operations remain within safe boundaries that prevent cognitive sound quality damage. This parameter-driven approach allows volume consistency while protecting sound quality integrity through mathematically determined safe operating limits.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12237818B2Audio signal processing method and device for controlling loudness level
Publication Date: 2025.02.25 GAUDI AUDIO LAB
  • US12237818B2 patent drawing
  • US12237818B2 patent drawing
  • US12237818B2 patent drawing

AI summary

An audio signal processing device comprises: a receiver for receiving an input audio signal; a processor for generating loudness metadata corresponding to the input audio signal; and an outputter for transmitting the loudness metadata generated by the processor. The processor is configured to acquire loudness information analyzed from input content, acquires loudness information about the input audio signal by measuring the loudness of the input audio signal, generates the loudness metadata by converting the loudness information, and transmits, through the outputter, the generated loudness metadata to an output device for outputting the input audio signal.