Loudness Normalization via Volume Metadata
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing loudness normalization methods for audio content delivery across networks face challenges such as damaging dynamic range, high transcoding costs, inability to optimize for various client environments, and limited control over external content volumes, leading to user discomfort due to volume differences between content items.
Innovation Solution
A loudness normalization method and system that adjusts the volume output level of a player using volume level metadata, allowing for optimized playback on client devices without requiring server-side transcoding, enabling adaptation to different client environments and allowing volume control for external content.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If loudness normalization is performed through server-side transcoding, then volume levels can be adjusted to proper broadcasting standards, but dynamic range is damaged and original creator intention is disrupted
Solution Approach 1:
The patent extracts the volume level adjustment function from the transcoding process and implements it separately through metadata-based control. The server provides volume level metadata without modifying the actual audio content, allowing clients to adjust volume independently while preserving the original audio's dynamic range and creator intention.
Solution Approach 2:
The patent introduces volume level metadata as an intermediary between the audio content and the playback system. This metadata acts as a mediator that conveys volume information without altering the actual audio stream, enabling volume normalization through client-side processing rather than server-side transcoding.
2Manufacturing precision
If entire content is transcoded to adjust volume level, then proper broadcasting volume is achieved, but transcoding cost increases significantly
Solution Approach 1:
The patent extracts only the essential volume level information from the audio content and represents it through compact metadata structures. This eliminates the need to process and re-encode the entire audio stream, reducing transcoding requirements to minimal operations while achieving the same volume adjustment effect.
Solution Approach 2:
Instead of creating a new transcoded audio stream, the patent uses volume level metadata as a copy of the original content's acoustic characteristics. This metadata copy contains all necessary volume information without requiring actual audio reprocessing, significantly reducing computational resources and transcoding costs.
3Manufacturing precision
If server performs volume level adjustment for all content, then broadcasting standards are met, but adaptability to different client environments is lost
Solution Approach 1:
The patent implements dynamic volume level adjustment at the client end rather than static server-side processing. Each client can independently process the volume level metadata according to its specific acoustic characteristics and user preferences, enabling adaptive normalization for different devices, platforms, and listening environments while maintaining broadcasting standards.
4Manufacturing precision
If server controls volume level for all content, then broadcasting volume is standardized, but control over external content volume is lost
Solution Approach 1:
The patent enables clients to autonomously extract and process volume level metadata from external content without requiring server intervention. Clients can independently manage volume levels for external content based on their own acoustic calibration and user preferences, giving them full control over external content volume while maintaining consistency with internal content.
Data Source
AI summary
A loudness normalization method includes receiving data for playback of content from a server in response to a user's request to play back the content; normalizing the loudness of the content by adjusting the volume output level of a player using volume level metadata of the content included in the received data; and providing the content by playing audio of the content based on the adjusted volume output level of the player.


