Spatial Audio Focus Objects with Directional Priority Encoding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing spatial audio systems lack user control over audio focus capture, resulting in encoded streams that are not configurable, leading to increased computational complexity and transmission bandwidth requirements.
Innovation Solution
Incorporating an audio signal type attribute and direction information into audio focus objects, along with a priority attribute, to enable user-controlled focus adjustment and priority-based encoding and decoding of audio streams, allowing for flexible focus control without additional streams.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If an audio focused stream is captured and encoded for delivery, then the audio capture is focused on a selected direction, but the end user has no control over the audio focus configuration
Solution Approach 1:
The audio scene is segmented into multiple audio focus objects, each representing a distinct sound source with its own focus direction and priority level. This segmentation allows the encoder to process and prioritize different audio sources independently, enabling user control over which audio events receive focus without requiring a complete redesign of the encoding system.
Solution Approach 2:
The audio focus configuration is made dynamic by allowing the encoder to adjust focus directions and priority assignments based on real-time audio scene analysis and user preferences. The system can dynamically switch between different focus configurations without requiring multiple static encoded streams, thus providing user control while maintaining encoding efficiency.
2Adaptability or versatility
If multiple audio streams are captured and encoded to provide user control options, then the end user can select different focus configurations, but the transmission bandwidth increases
Solution Approach 1:
A single encoded audio stream is designed to serve multiple functions by incorporating metadata that describes multiple audio focus objects with different focus directions and priority levels. The decoder can selectively render different focus configurations from this universal stream, eliminating the need to transmit multiple separate streams for different user preferences.
Solution Approach 2:
The system encodes audio with parameters that describe focus directions, priority levels, and spatial characteristics of multiple audio events. By changing these parameters at the decoder based on user selection, the system provides configurable audio focus options without increasing transmission bandwidth, as all parameter variations are contained within a single encoded stream.
3Adaptability or versatility
If multiple audio streams are captured and encoded to provide alternative audio scene experiences, then the end user can access different rendering options, but the computational complexity increases
Solution Approach 1:
The encoder performs preliminary analysis of the audio scene to identify and characterize multiple audio events, their directions, and relative priorities before encoding. This preliminary action creates a comprehensive metadata structure that describes all possible focus configurations, allowing the decoder to efficiently render different audio scene experiences without performing complex real-time analysis, thus reducing computational complexity.
Solution Approach 2:
Metadata acts as an intermediary between the encoded audio stream and the decoder, carrying information about multiple audio focus objects and their characteristics. This intermediary structure allows the decoder to selectively render different audio scene experiences by interpreting the metadata without requiring multiple separate encoded streams, thereby reducing computational complexity while maintaining versatility.
Data Source
Figure 1a
Figure 1b
Figure 2
AI summary
There is disclosed inter alia a method for spatial audio signal encoding comprising adding an audio signal type attribute to an audio focus object, wherein the audio focus object comprises a downmixed audio difference signal and direction information relating to an audio focus direction associated with the audio focus object, and wherein the audio signal type attribute identifies that an encoded audio signal is an audio focus object.