Spatial Audio File Format for Dynamic SR Sound Composition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio formats struggle to effectively produce three-dimensional sound effects in simulated reality environments, such as augmented, virtual, and mixed reality, as they were designed for stationary listeners in physical spaces, lacking a uniform method to access and integrate various sound sources dynamically.
Innovation Solution
A file format for spatial audio that enables developers to compose and manage audio assets with metadata describing sound characteristics and transformation parameters, allowing for dynamic rendering of 3D sound experiences in simulated reality environments, supporting various sound sources and playback systems, including binaural and loudspeaker-based systems.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If existing spatial audio formats (MPEG-H, HOA, DOLBY ATMOS) are used to produce 3D sound in SR environments, then 3D sound effects can be generated, but the formats were designed for stationary listeners in physical spaces with fixed speaker locations, making them difficult to access and integrate dynamically in virtual environments
Solution Approach 1:
The patent segments audio into discrete audio objects that can be independently manipulated, each with its own spatial metadata. This allows individual audio elements to be accessed, retrieved, and integrated into SR environments separately, providing uniform access to various sound sources without requiring complex integration processes.
Solution Approach 2:
The patent adds a metadata dimension to traditional audio formats, storing not only audio data but also spatial information describing listener position, object position, and transformation parameters. This dimensional extension enables dynamic adaptation to SR environments while maintaining ease of access through standardized retrieval mechanisms.
2Adaptability or versatility
If audio assets are stored with comprehensive metadata describing encoding and listener experience characteristics, then dynamic transformation for non-stationary listeners is enabled, but the complexity of managing and processing the metadata increases
Solution Approach 1:
The patent performs preliminary encoding of spatial metadata during audio asset creation, storing transformation parameters that describe listener position, object position, and spatial relationships before the audio is needed in the SR environment. This advance preparation eliminates the need for complex real-time metadata processing while enabling dynamic adaptation.
Solution Approach 2:
The patent introduces standardized metadata as an intermediary layer between raw audio data and SR environment requirements. This metadata acts as a mediator that translates audio encoding characteristics into SR-specific spatial transformations, simplifying the management complexity while maintaining adaptability.
3Ease of operation
If a uniform method is implemented to access various sound sources, then ease of composition is improved, but the ability to handle diverse audio encodings and playback systems may be compromised
Solution Approach 1:
The patent creates a universal audio asset structure that can accommodate multiple audio encodings and playback systems through a common metadata framework. The standardized asset format serves multiple functions: it stores audio data, describes encoding characteristics, specifies spatial transformations, and provides retrieval mechanisms, thereby maintaining both uniform access and diverse encoding support.
Data Source
AI summary
An audio asset library containing audio assets formatted in accordance with a file format for spatial audio includes asset metadata that enables simulated reality (SR) application developers to compose sounds for use in SR applications. The audio assets are formatted to include audio data encoding a sound capable of being composed into a SR application along with asset metadata describing not only how the sound was encoded, but also how a listener in SR environment experiences the sound. A SR developer platform is configured so that developers can compose sound for SR objects using audio assets stored in the audio library, including editing the asset metadata to include transformation parameters that support dynamic transformation of the asset metadata in the SR environment to alter how the SR listener experiences the composed sound. Other embodiments are also described and claimed.


