Ambisonic Audio Rendering with Stereo Tracks for VR
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current ambisonic audio systems in virtual reality and 360-degree video applications provide immersive but non-interactive audio experiences, requiring significant computational resources for binaural rendering.
Innovation Solution
A media delivery computer system that performs convolution operations on ambisonic sound fields using head-related transfer functions and integrates stereo tracks based on user-oriented events, allowing for interactive audio experiences with minimal computational overhead by combining ambisonic and stereo sound fields.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If head-tracked spatialized mix using ambisonic sound is used, then immersive audio experience is improved, but computational resources required increase
Solution Approach 1:
The audio system is segmented into multiple independent ambisonic sound fields, each associated with specific virtual objects or sound sources. This allows the system to process and render only the necessary sound fields based on user interaction and head position, rather than continuously processing all possible audio data, thereby reducing computational overhead while maintaining immersive quality.
Solution Approach 2:
Ambisonic sound fields are pre-rendered and prepared in advance based on predicted user interactions and head positions. Event data structures are pre-configured with audio triggers and parameters, allowing the system to quickly activate appropriate audio experiences without performing heavy real-time computations, thus reducing energy consumption during actual interaction.
2Adaptability or versatility
If multiple ambisonic sound fields are combined, then audio versatility is improved, but system complexity increases
Solution Approach 1:
A universal event data structure is implemented that can represent multiple types of audio triggers and interactions within a single standardized format. This structure handles various ambisonic sound field combinations, head position events, and audio playback scenarios through a unified interface, reducing system complexity while maintaining high audio versatility and adaptability.
3Manufacturing precision
If real-time binaural rendering is performed, then audio quality is improved, but processing time increases
Solution Approach 1:
The system performs binaural rendering only for the specific ambisonic sound fields that are currently active and relevant to the user's head position and interaction state, rather than rendering all sound fields continuously. This partial action approach maintains high audio quality for the necessary components while significantly reducing overall processing time and computational load.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
Enables interactive and immersive audio experiences in virtual reality environments with reduced computational burden, allowing users to engage with virtual objects through audio triggers without the need for additional processing, thus enhancing user interaction with minimal resource usage.
Implementation Method 1
performing, by a media delivery computer configured to produce the sound field in the ears of the human user, a convolution operation on an ambisonic portion of the sound field being output to the ears of the human user over a set of ambisonic sound channels with a head-related transfer function (HRTF) for that ear to produce a rendered ambisonic portion of the sound field in the ears of the human listener over the set of ambisonic sound channels
Data Source
Figure 1
Figure 2
Figure 3A~3B
AI summary
Techniques of performing involve providing interactive audio in addition to ambisonic audio in stereo tracks selected according to the occurrence of events in a media delivery system. For example, a user of a VR system observes a virtual environment that contains many virtual objects. The user may experience binaurally rendered audio played over N ambisonic channels from any number of virtual loudspeakers. In addition, the user may also activate another audio source by positioning his/her head at a certain angle, e.g., to look at a particular virtual object. As a specific example, when the user looks at a picture of a person, an audio track may play over a pair of stereo channels N+1 and N+2. Because they are stereo channels, there is no need to perform convolutions with HRTFs. In this way, audio may be provided for all virtual objects in the virtual environment with a small computational overhead.