Spatial Audio Focus Objects with Directional Priority Encoding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing spatial audio systems lack user control over audio focus capture, resulting in encoded streams that are not configurable, leading to increased computational complexity and transmission bandwidth requirements.

Innovation Solution

Incorporating an audio signal type attribute and direction information into audio focus objects, along with a priority attribute, to enable user-controlled focus adjustment and priority-based encoding and decoding of audio streams, allowing for flexible focus control without additional streams.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If an audio focused stream is captured and encoded for delivery, then the audio capture is focused on a selected direction, but the end user has no control over the audio focus configuration

Engineering Contradiction:
Improveuser control over audio focusVSAvoidaudio encoding system complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The audio scene is segmented into multiple audio focus objects, each representing a distinct sound source with its own focus direction and priority level. This segmentation allows the encoder to process and prioritize different audio sources independently, enabling user control over which audio events receive focus without requiring a complete redesign of the encoding system.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The audio focus configuration is made dynamic by allowing the encoder to adjust focus directions and priority assignments based on real-time audio scene analysis and user preferences. The system can dynamically switch between different focus configurations without requiring multiple static encoded streams, thus providing user control while maintaining encoding efficiency.

Inventive Principle:
Principle #15Dynamics

2Adaptability or versatility

If multiple audio streams are captured and encoded to provide user control options, then the end user can select different focus configurations, but the transmission bandwidth increases

Engineering Contradiction:
Improveconfigurable audio focus optionsVSAvoidtransmission bandwidth
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

A single encoded audio stream is designed to serve multiple functions by incorporating metadata that describes multiple audio focus objects with different focus directions and priority levels. The decoder can selectively render different focus configurations from this universal stream, eliminating the need to transmit multiple separate streams for different user preferences.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system encodes audio with parameters that describe focus directions, priority levels, and spatial characteristics of multiple audio events. By changing these parameters at the decoder based on user selection, the system provides configurable audio focus options without increasing transmission bandwidth, as all parameter variations are contained within a single encoded stream.

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If multiple audio streams are captured and encoded to provide alternative audio scene experiences, then the end user can access different rendering options, but the computational complexity increases

Engineering Contradiction:
Improvealternative audio scene renderingVSAvoidcomputational complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The encoder performs preliminary analysis of the audio scene to identify and characterize multiple audio events, their directions, and relative priorities before encoding. This preliminary action creates a comprehensive metadata structure that describes all possible focus configurations, allowing the decoder to efficiently render different audio scene experiences without performing complex real-time analysis, thus reducing computational complexity.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Metadata acts as an intermediary between the encoded audio stream and the decoder, carrying information about multiple audio focus objects and their characteristics. This intermediary structure allows the decoder to selectively render different audio scene experiences by interpreting the metadata without requiring multiple separate encoded streams, thereby reducing computational complexity while maintaining versatility.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentEP3824464B1Controlling audio focus for spatial audio processing
Publication Date: 2024.07.17 NOKIA TECHNOLOGIES OY
  • EP3824464B1 patent drawingFigure 1a
  • EP3824464B1 patent drawingFigure 1b
  • EP3824464B1 patent drawingFigure 2

AI summary

There is disclosed inter alia a method for spatial audio signal encoding comprising adding an audio signal type attribute to an audio focus object, wherein the audio focus object comprises a downmixed audio difference signal and direction information relating to an audio focus direction associated with the audio focus object, and wherein the audio signal type attribute identifies that an encoded audio signal is an audio focus object.