Audio Decoder User Position Metadata Update
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio coding methods, such as MPEG-H 3D audio, fail to maintain a realistic sense of immersion when a user changes position in a virtual reality or three-dimensional audio environment, as they do not effectively update audio object positions and volumes in response to user movement.
Innovation Solution
A method and apparatus that incorporate user position information, including a user position change indicator and offset, to modify metadata and render audio signals accordingly, allowing for dynamic adjustment of audio object positions and gains based on user movement, enabling continuous immersive audio output even when the user changes location.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If existing audio coding methods are used, then audio output is provided at a fixed location, but the sense of immersion degrades when user position changes
Solution Approach 1:
The patent applies dynamics by making the audio rendering system adaptive to user position changes. The binaural renderer dynamically adjusts audio object positions and metadata based on detected user position information, allowing the audio environment to respond continuously as the user moves through arbitrary spaces, thereby maintaining immersion reliability across varying positions.
Solution Approach 2:
The patent changes parameters by modifying metadata based on user position offsets. When user position changes are detected, the system adjusts audio object positions, gains, and other rendering parameters accordingly, enabling the audio output to adapt to new spatial configurations while preserving the immersive experience.
2Reliability
If user position information is added to determine user position during audio decoding, then audio output performance is enhanced, but device complexity increases
Solution Approach 1:
The patent applies preliminary action by pre-defining user position information structures and offset metadata formats in advance. The system prepares position change indicators and offset data structures beforehand, allowing efficient processing during runtime without requiring complex real-time calculations, thus enhancing audio performance while controlling processing complexity.
Data Source
AI summary
A method and apparatus for outputting an audio signal corresponding to a user position are disclosed. The method includes receiving an audio signal and providing a decoding audio signal and decoded metadata, checking whether a user position is changed in an arbitrary space using user position information including a user position change indicator and user position change offset, when the user position is changed, providing modified metadata obtained by correcting the decoded metadata based on the user position change offset, and rendering the decoded audio signal using the modified metadata. Accordingly, it is possible to provide an audio sound image that is changed in response to change in user position in an arbitrary space, thereby providing more realistic audio output.


