Viewport-Aware Audio Streaming for Low-Complexity VR Delivery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing VR systems face complexity and computational challenges in delivering real-time audio streams that adapt to user movement and orientation, requiring advanced client-server communication and high bitrate, which are beyond current equipment capabilities.
Innovation Solution
A system that requests and decodes audio streams based on user viewport, head orientation, and movement data, prioritizing relevant audio elements at higher bitrates, and selectively delivers streams based on proximity to scene boundaries, using a server that encodes and stores audio elements associated with specific scenes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If complete audio scenes are encoded into multiple streams for different user positions and orientations, then audio adaptability to user movement is improved, but system complexity and computational requirements increase beyond current equipment capabilities
Solution Approach 1:
The patent segments the complete audio scene into multiple discrete audio elements, each associated with specific spatial coordinates and characteristics. This allows the system to deliver only the relevant audio elements needed for the user's current viewport and orientation, rather than transmitting entire audio scenes. The segmentation enables selective delivery based on user position, reducing complexity while maintaining adaptability.
Solution Approach 2:
The patent applies local quality by delivering audio elements with different quality levels based on their relevance to the user's current viewport. Audio elements within the user's field of view are delivered at higher quality, while those outside are delivered at lower quality or not at all. This selective quality delivery reduces overall system complexity while maintaining high adaptability to user movement and orientation.
2Reliability
If all audio elements are delivered at high bitrate to maintain quality, then audio quality is improved, but bandwidth consumption increases beyond available communication capacity
Solution Approach 1:
The patent implements local quality by delivering audio elements at different bitrate levels based on their spatial relationship to the user's viewport. Audio elements directly within the user's field of view receive high bitrate delivery to maintain quality, while audio elements outside the viewport or less relevant to current user orientation are delivered at lower bitrates. This selective approach maintains audio quality where needed while reducing overall bandwidth consumption to levels compatible with available communication capacity.
3Adaptability or versatility
If real-time encoding is performed based on user feedback, then audio adaptability is improved, but computational requirements exceed current processing capabilities
Solution Approach 1:
The patent applies preliminary action by pre-encoding the complete audio scene into multiple discrete audio elements with associated spatial metadata before delivery. This preprocessing step creates a structured representation that enables efficient client-side selection based on user viewport and orientation without requiring complex real-time encoding operations. The computational work is shifted to the server side during content creation, reducing real-time processing requirements to simple selection and delivery operations.
Data Source
AI summary
There are disclosed techniques, systems, methods and instructions for a virtual reality, VR, augmented reality, AR, mixed reality, MR, or 360-degree video environment.In one example, the system includes at least one media video decoder configured to decode video signals from video streams for the representation of VR, AR, MR or 360-degree video environment scenes to a user. The system includes at least one audio decoder configured to decode audio signals from at least one audio stream. The system is configured to request at least one audio stream and/or one audio element of an audio stream and/or one adaptation set to a server on the basis of at least the user's current viewport and/or head orientation and/or movement data and/or interaction metadata and/or virtual positional data.


