Audio Loudness Metadata for Group-Level Track Normalization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for generating audio loudness metadata fail to accurately predict the overall loudness of content, leading to inconsistencies and distortions in artistic intent due to differences in loudness between audio tracks, especially in multi-clip content like symphonies or videos, which requires unified loudness normalization across entire pieces.
Innovation Solution
A method and device that receive loudness information from each audio track, predict an intermediate loudness distribution, and generate integrated loudness for a group of tracks, allowing for efficient loudness normalization by using a processor to analyze and adjust the loudness distribution based on integrated loudness, loudness range, and duration of each track.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If loudness correction is performed independently for each clip, then the loudness between clips is corrected, but the context of the content in its entirety is damaged
Solution Approach 1:
The patent segments the content into multiple clips while maintaining group-level loudness metadata. Each clip retains its own integrated loudness information, and the system provides both individual clip loudness correction capability and group-level loudness normalization, allowing users to choose between clip-by-clip correction and group-wide normalization based on the content type.
Solution Approach 2:
The patent introduces a new dimension of group-level loudness metadata that encompasses multiple clips. By providing loudness information at both the individual clip level and the group level, the system enables loudness normalization across the entire content while preserving the ability to correct individual clips when needed.
2Reliability
If loudness normalization is performed for the entire content, then the artistic intent is preserved, but the processing complexity increases
Solution Approach 1:
The patent performs preliminary loudness measurement and generates loudness metadata for each clip and for the group during the content production or preprocessing stage. This preliminary action stores the loudness information that can be later used for both individual clip correction and group-wide normalization without requiring complex real-time processing.
Solution Approach 2:
The patent introduces loudness metadata as an intermediary that bridges individual clip loudness information and group-level loudness normalization. The metadata includes integrated loudness, loudness range, and other parameters that facilitate both types of normalization through simple lookup and adjustment operations rather than complex real-time analysis.
3Adaptability or versatility
If multiple loudness standards are used for different countries, then local requirements are met, but direct use of loudness information becomes difficult
Solution Approach 1:
The patent creates a universal loudness metadata structure that can serve multiple purposes and comply with different standards. The metadata includes fundamental loudness parameters (integrated loudness, loudness range, momentary loudness) that are applicable across different countries and standards (EBU R 128, ATSC A/85, ITU-R BS.1770), allowing the same metadata to be used for both local compliance and international interoperability.
Data Source
AI summary
A method of generating audio loudness performed by an audio loudness generation device may include: receiving loudness information on each of a plurality of audio tracks included in one group; predicting an intermediate loudness distribution, which is a loudness distribution for the one group, on the basis of the loudness information on each of the plurality of audio tracks; and generating an integrated loudness for the one group on the basis of the intermediate loudness distribution.


