Parametric Spatial Audio Rendering for Immersive Communication
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing immersive audio codecs face challenges in efficiently processing and rendering spatial audio streams for immersive audio applications, particularly in virtual reality (VR), augmented reality (AR), and mixed reality (MR), while maintaining low latency and high error robustness.
Innovation Solution
The proposed solution involves an apparatus and method that receive multiple audio data streams, identify spatial audio streams, process them with parameters defining room characteristics or scene descriptions, and render the processed streams alongside other audio streams, allowing for flexible and efficient rendering of immersive audio.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If spatial audio streams are processed with parametric spatial audio processing to describe spatial properties, then the quality and flexibility of immersive audio rendering is improved, but the device complexity and processing requirements increase
Solution Approach 1:
The patent applies parametric spatial audio processing to describe the spatial properties of sound sources using parameters such as direction, distance, and spatial extent. This allows flexible manipulation of spatial audio characteristics without requiring complex full-sphere microphone array processing, thereby improving rendering flexibility while controlling processing complexity through parameter-based representation
2Adaptability or versatility
If immersive audio codecs support multiple operating points from low bit rate to transparency, then the adaptability to different applications is improved, but the device complexity increases
Solution Approach 1:
The IVAS codec is designed to handle multiple types of audio content (speech, music, generic audio) and support multiple operating points from low bit rate to transparency through a unified codec structure. This multi-functional design allows the same codec to adapt to different application requirements without requiring separate specialized codecs for each scenario
Solution Approach 2:
The codec dynamically adapts its processing based on the input signal characteristics and transmission conditions, adjusting between different operating points to optimize performance for speech, music, or generic audio content while maintaining appropriate quality levels across varying bit rate requirements
3Speed
If the codec operates with low latency to enable conversational services, then the speed of audio processing is improved, but the processing precision may be compromised
Solution Approach 1:
The codec applies different processing levels to different audio content types, using more aggressive compression and processing for non-critical content while maintaining higher precision for speech content where low latency is critical for conversational services. This selective processing approach balances speed requirements with precision needs across different audio streams
4Reliability
If the codec supports high error robustness under various transmission conditions, then the reliability is improved, but the processing complexity increases
Solution Approach 1:
The codec incorporates error robustness mechanisms in advance, using redundant encoding and error concealment techniques that are built into the encoding process. This allows the decoded audio to maintain quality even when transmission errors occur, providing beforehand protection against transmission failures without requiring complex real-time error correction during playback
Data Source
AI summary
An apparatus configured to: receive, at least, at least one first audio channel and at least one second audio channel, wherein at least one of the at least one first audio channel or the at least one second audio channel comprises spatial audio configured to enable immersive audio communication; determine a format of at least one of the at least one first or the at least one second audio channel to identify which of the received at least one first audio channel and the at least one second audio channel comprises the spatial audio; process the identified at least one audio channel with at least one parameter dependent on the determined format; and render the processed at least one audio channel and another of the at least one first audio channel or the at least one second audio channel.


