Immersive Audio Loudness Normalization Using Anchor Speech
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies struggle to effectively adjust loudness levels in immersive audio scenes for MPEG-I presentations, leading to inconsistent and unrealistic sound experiences for users navigating and interacting with virtual or augmented reality environments.
Innovation Solution
A method for loudness adjustment in MPEG-I immersive audio streams involving the use of an anchor speech signal and a general binaural renderer with Dirac head-related transfer function to normalize sound levels, with optional generation of an adjusted speech signal based on multiple speech signals present in the scene, and signaling information for adjusting sound levels.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If loudness levels are adjusted manually or normalized by loudness measurements in existing technologies, then some level of sound control is achieved, but the sound experience becomes inconsistent and unrealistic for users in virtual or augmented reality environments
Solution Approach 1:
The system automatically determines reference signals from the audio scene and performs loudness normalization without manual intervention. The processor independently identifies speech signals, determines reference signals, and adjusts loudness levels, enabling the system to self-regulate audio output for consistent user experience across different devices and scenarios.
2Adaptability or versatility
If different sound levels are set in listening test setups or virtual scenes, then audio variety and realism are improved, but clipping and silence issues occur that degrade audio quality
Solution Approach 1:
The system performs preliminary loudness normalization by determining reference signals from the audio scene before final playback. By pre-processing the audio to establish appropriate reference signals and normalization factors, the system prevents clipping and silence issues from occurring during actual playback, ensuring audio quality is maintained across varied sound levels and scenes.
3Productivity
If loudness normalization is applied without proper reference signal determination, then processing speed is maintained, but the reference signal accuracy deteriorates leading to poor loudness adjustment
Solution Approach 1:
The system uses feedback by analyzing the audio scene to determine whether speech signals are present and using them as reference signals. The processor continuously monitors the audio content, identifies appropriate reference signals based on speech detection, and adjusts normalization accordingly, ensuring both accuracy and efficient processing through intelligent signal selection.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Aspects of the disclosure include methods, apparatuses, and non-transitory computer-readable storage mediums for loudness adjustment for an audio scene associated with an MPEG-I immersive audio stream. One apparatus includes processing circuitry that receives a first syntax element indicating a number of sound signals included in the audio scene. The processing circuitry determines whether one or more speech signals are included in the sound signals indicated by the first syntax element. The processing circuitry determines a reference speech signal from the one or more speech signals based on the one or more speech signals being included in the sound signals. The processing circuitry adjusts a loudness level of the reference speech signal of the audio scene based on an anchor speech signal. The processing circuitry adjusts loudness levels of the sound signals based on the adjusted loudness level of the reference speech signal.