Viewport Selection Using Audio Spatial Metadata
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing virtual reality (VR) systems experience lag time when upsampling video data due to reactive selection of viewports, which affects the user experience as they shift their field of view, while audio playback remains unaffected by quality fluctuations.
Innovation Solution
The use of auditory aspects, specifically the directionality and energy of audio objects in a higher-order ambisonics (HOA) representation, to predictively select and upsample video data at likely-next viewports, thereby preemptively enhancing video quality before the user's gaze shifts, without introducing additional scene analysis processes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If reactive viewport selection is used in VR systems, then video data can be upsampled based on current field of view, but lag time occurs during viewport transitions affecting user experience
Solution Approach 1:
The system performs preliminary action by predictively selecting future viewports based on audio object directionality and energy information before the user actually shifts their field of view. This allows video data for these predicted viewports to be upsampled in advance, eliminating the lag time that would otherwise occur during viewport transitions while maintaining high video quality.
2Measurement precision
If video data is upsampled for all viewports, then visual quality is improved, but processing time and computational resources increase significantly
Solution Approach 1:
The system applies local quality by selectively upsampling video data only for specific predicted viewports rather than all viewports. By using audio spatial metadata to identify which viewports are most likely to be viewed next, the system maintains high video quality for those specific regions while avoiding the computational overhead of processing the entire video dataset.
Solution Approach 2:
The system performs partial action by upsampling video data for only a subset of viewports - specifically those predicted to be next based on audio cues - rather than performing the complete action of upsampling all viewports. This partial processing maintains video quality where needed while preserving processing efficiency.
3Measurement precision
If audio spatial metadata is used for viewport prediction, then viewport selection accuracy is improved, but system complexity increases due to additional data processing
Solution Approach 1:
The system applies universality by using audio spatial metadata that already exists for audio playback purposes to simultaneously perform viewport prediction. The same audio object directionality and energy information that drives audio rendering is repurposed to predict video viewport selections, eliminating the need for separate scene analysis processes and reducing overall system complexity.
Solution Approach 2:
The audio spatial metadata serves a dual purpose: it drives audio playback and simultaneously enables viewport prediction without requiring additional processing. The audio data essentially serves itself by providing the information needed for both audio rendering and video viewport selection, reducing the need for separate analysis systems.
Data Source
Figure 1
Figure 2
Figure 3A
AI summary
An example device includes a memory device, and a processor coupled to the memory device. The memory is configured to store audio spatial metadata associated with a soundfield and video data. The processor is configured to identify one or more foreground audio objects of the soundfield using the audio spatial metadata stored to the memory device, and to select, based on the identified one or more foreground audio objects, one or more viewports associated with the video data. Display hardware coupled to the processor and the memory device is configured to output a portion of the video data being associated with the one or more viewports selected by the processor.