Immersive Audio Loudness Normalization Using Anchor Speech

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies struggle to effectively adjust loudness levels in immersive audio scenes for MPEG-I presentations, leading to inconsistent and unrealistic sound experiences for users navigating and interacting with virtual or augmented reality environments.

Innovation Solution

A method for loudness adjustment in MPEG-I immersive audio streams involving the use of an anchor speech signal and a general binaural renderer with Dirac head-related transfer function to normalize sound levels, with optional generation of an adjusted speech signal based on multiple speech signals present in the scene, and signaling information for adjusting sound levels.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If loudness levels are adjusted manually or normalized by loudness measurements in existing technologies, then some level of sound control is achieved, but the sound experience becomes inconsistent and unrealistic for users in virtual or augmented reality environments

Engineering Contradiction:
Improvesound experience consistencyVSAvoidloudness adjustment complexity
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The system automatically determines reference signals from the audio scene and performs loudness normalization without manual intervention. The processor independently identifies speech signals, determines reference signals, and adjusts loudness levels, enabling the system to self-regulate audio output for consistent user experience across different devices and scenarios.

Inventive Principle:
Principle #25Self-service

2Adaptability or versatility

If different sound levels are set in listening test setups or virtual scenes, then audio variety and realism are improved, but clipping and silence issues occur that degrade audio quality

Engineering Contradiction:
Improveaudio scene varietyVSAvoidclipping and silence
Core Design Contradiction:
Adaptability or versatilityVSObject-affected harmful factors

Solution Approach 1:

The system performs preliminary loudness normalization by determining reference signals from the audio scene before final playback. By pre-processing the audio to establish appropriate reference signals and normalization factors, the system prevents clipping and silence issues from occurring during actual playback, ensuring audio quality is maintained across varied sound levels and scenes.

Inventive Principle:
Principle #10Preliminary action

3Productivity

If loudness normalization is applied without proper reference signal determination, then processing speed is maintained, but the reference signal accuracy deteriorates leading to poor loudness adjustment

Engineering Contradiction:
Improveprocessing speedVSAvoidreference signal accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system uses feedback by analyzing the audio scene to determine whether speech signals are present and using them as reference signals. The processor continuously monitors the audio content, identifies appropriate reference signals based on speech detection, and adjusts normalization accordingly, ensuring both accuracy and efficient processing through intelligent signal selection.

Inventive Principle:
Principle #23Feedback

Data Source

PatentEP4101181B1Signaling loudness adjustment for an audio scene
Publication Date: 2025.12.03 TENCENT AMERICA LLC
  • EP4101181B1 patent drawingFigure 1
  • EP4101181B1 patent drawingFigure 2
  • EP4101181B1 patent drawingFigure 3

AI summary

Aspects of the disclosure include methods, apparatuses, and non-transitory computer-readable storage mediums for loudness adjustment for an audio scene associated with an MPEG-I immersive audio stream. One apparatus includes processing circuitry that receives a first syntax element indicating a number of sound signals included in the audio scene. The processing circuitry determines whether one or more speech signals are included in the sound signals indicated by the first syntax element. The processing circuitry determines a reference speech signal from the one or more speech signals based on the one or more speech signals being included in the sound signals. The processing circuitry adjusts a loudness level of the reference speech signal of the audio scene based on an anchor speech signal. The processing circuitry adjusts loudness levels of the sound signals based on the adjusted loudness level of the reference speech signal.