Audio Decoder Scene Orientation Rendering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional MPEG-H 3D Audio encoders do not account for video scene orientation information, leading to inaccurate sound source direction during VR content playback, as audio and video are encoded independently without mutual interaction, resulting in a decreased sense of immersion.
Innovation Solution
A method and apparatus for outputting an audio signal using scene orientation information, which includes receiving an audio signal, generating decoded audio and object metadata, modifying metadata based on external control information, and rendering the audio signal according to scene orientation information, ensuring adaptive audio playback with video scene changes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If audio and video are encoded independently without mutual interaction, then encoding complexity is reduced, but sound source direction accuracy deteriorates
Solution Approach 1:
The patent merges audio encoding with video scene orientation information by having the audio encoder receive and process scene orientation data from the video encoder. This allows the audio signal to be encoded with directional information that corresponds to the video scene, resolving the contradiction by combining previously separate encoding processes while maintaining manageable complexity through standardized data exchange protocols.
Solution Approach 2:
The patent introduces scene orientation information as an intermediary element that mediates between video and audio encoding. The video encoder generates scene orientation information that is then transmitted to the audio encoder, serving as a bridge that enables accurate sound source direction without requiring direct complex interaction between the audio and video encoding processes.
2Measurement precision
If scene orientation information is integrated into audio encoding, then sound source direction accuracy is improved, but device complexity increases
Solution Approach 1:
The patent segments the encoding system into distinct functional modules: a video encoder that generates scene orientation information, a separate audio encoder that receives and processes this information, and a decoder that reconstructs the audio signal. This segmentation allows each component to handle specific tasks independently, improving sound source direction accuracy while managing overall system complexity through modular design.
Solution Approach 2:
The patent creates a multi-functional encoding framework where the audio encoder serves both traditional audio compression functions and the additional function of integrating scene orientation information. This universal approach allows the same encoding infrastructure to handle both audio signal processing and spatial orientation data, reducing the need for separate dedicated systems and thereby controlling complexity while improving directional accuracy.
Data Source
AI summary
A method for decoding a bitstream by an apparatus, includes obtaining a decoded audio signal and metadata from the bitstream, the metadata comprising scene orientation information; and rendering the decoded audio signal based on the scene orientation information, wherein the scene orientation information is information for a direction of a video scene related to the decoded audio signal.


